AI GovernanceAugust 16, 2026· 5 min read

AI Evaluation, Observability and Red Teaming: What CIOs Should Demand Before Agents Go Live

Three disciplines catch three different failures: evaluation catches the quality drop, observability explains it, red teaming finds what benign testing never will. The three-layer failure-detection map, the go-live gate that uses all three, and what to demand from vendors in Barcelona.

Alejo Hernandez

Alejo Hernandez

CTO


Every agent programme reaches the same meeting: the demo went well, the business wants it live, and someone — usually the CIO — has to decide what ready means. If that meeting is on your calendar before Gartner IT Symposium/Xpo 2026, this is the framework worth taking to Barcelona: three disciplines, three failure classes, one gate.

The three disciplines get conflated constantly — vendors use the words interchangeably, which is precisely why buyers should not. They catch different failures, and a gap in any one of them is a failure class you have chosen not to see.

The three-layer failure-detection map

LayerThe question it answersThe failure only it catches
EvaluationDoes it meet the bar we set?The quality regression: the task that quietly stopped completing after a model, prompt or retrieval change.
ObservabilityWhat did it actually do, and why?The unexplainable incident: a wrong answer with no trace of what was retrieved, which tool ran, what the guardrail saw.
Red teamingCan it be made to break the rules?The adversarial failure: injection, data leakage, prohibited actions — behaviour no benign test suite will ever surface.

Read the map defensively: a strength in one layer cannot cover a gap in another. Perfect evals will not explain a production incident. Perfect traces will not tell you the agent is one crafted email away from exfiltrating a customer table. A clean red-team report says nothing about whether quality regressed last Tuesday.

Evaluation: gates, not demos

The demo is a sample of one. An evaluation is a repeatable definition — a subject, a dataset, evaluators, thresholds — whose second run costs one click, which is what makes it a gate rather than a ceremony. For agents, evaluate more than answer quality: task completion, tool selection, argument correctness, trajectory efficiency, policy compliance, cost per completed task. And keep two honesty rules in view: deterministic checks wherever a rule can be written, and where an LLM judge is needed, remember a judge makes a rubric repeatable, not objective — ship its reasoning with every score and calibrate it against human reviewers.

Offline evals gate the release. Online evals — production traffic sampled and scored continuously — catch what drifts afterwards. A programme with only the first ages out in weeks.

Observability: the evidence substrate

When quality drops or a regulator asks why did the system say this, averages will not answer. You need the trace: the request, what retrieval fetched, every tool call and agent hop, the guardrail decisions on the way in and out, latency and cost per span, and the eval score attached afterwards. Two properties decide whether that trace is worth having. It must include the policy decisions — a trace that shows the timing but not what the guardrail decided is a debugging aid, not evidence. And it must live inside your boundary — traces embed prompts, retrieved documents and identities, which is exactly the data many enterprises cannot send to a multi-tenant cloud.

Wire this before launch. Observability added after the first incident explains the second one.

Red teaming: the tests that fight back

Benign test suites share an assumption: the input wants the system to succeed. Attackers do not. Red teaming — single-turn and escalating multi-turn attacks probing injection, leakage, tool abuse, policy circumvention — is the only layer that tests the system against inputs designed to make it fail. For agents, attack the whole assembly: the complete agent with its tools and retrieval, not the bare model, because the dangerous behaviours live in the connections.

The discipline that separates a programme from a pentest is what happens to findings. Every one should map to a disposition: prompt or retrieval fix, tool-permission change, gateway guardrail, approval requirement, rate or budget limit, containment boundary, or accepted residual risk with a named owner. Not every finding becomes a guardrail — but every finding becomes a decision. A finding that becomes a ticket that ages is the most expensive kind: you paid to learn the risk, then kept it.

The go-live gate, in one paragraph

An agent ships when: its offline evaluation passes versioned thresholds someone owns; it has been red-teamed as assembled, findings dispositioned; its traces — policy decisions included — flow before day one; production sampling into online evals is switched on; and a failed anything has a named path to an enforced control. That last clause is where the three layers stop being three tools: in Kosmoy the finding, the guardrail that fixes it, and the re-run that proves the fix live on one platform, and every run lands in the same evidence stream an auditor reads.

What to demand in Barcelona

Ask every vendor claiming this space three questions. Which of the three layers do you actually do — and watch for the word swap where monitoring is sold as observability, or a benchmark slide as evaluation. Show me a finding becoming a control — not a report, the enforcement. Where do my traces live? Then compare notes with the written versions: the evaluation buyer's guide, the observability guide, and the red-teaming guide — honest about where specialists beat us.

We will demonstrate all three layers as one loop on the expo floor, 9–12 November: eval fails, finding becomes guardrail, re-run proves it. Bring your hardest agent. Book a meeting before the event, or email sales@kosmoy.com with the days you are on site.

FAQ

What is the difference between AI evaluation, observability and red teaming?

Evaluation measures outputs against criteria — does the system meet the bar we set? Observability keeps the traces and context that explain what it actually did and why. Red teaming attacks the system on purpose to find what benign testing never will: injection, leakage, prohibited actions. Each catches a failure class the other two miss; a go-live gate needs all three.

What should a pre-launch gate for AI agents include?

Offline evaluation runs on a versioned dataset with pass thresholds someone owns; adversarial red teaming of the complete agent — tools and retrieval included, not the bare model; trace-level observability wired before launch, not after the first incident; and a named path from every failed test to an enforced control. A demo is not a gate.

Should red-team findings become runtime controls automatically?

Not automatically — but traceably. Every finding should map to a disposition: a prompt or retrieval fix, a tool-permission change, a gateway guardrail, an approval requirement, a rate or budget limit, a containment boundary, or an accepted residual risk with an owner. A finding that becomes a ticket that ages is the most expensive kind: you paid to learn the risk and then kept it.


Gartner and Gartner IT Symposium/Xpo are trademarks of Gartner, Inc. and/or its affiliates. Kosmoy is an exhibitor at the 2026 Barcelona conference. Gartner does not endorse Kosmoy or its products.

gartner-it-symposiumai-evaluationai-observabilityred-teaming

See how Kosmoy works

Discover how enterprises govern, secure, and optimize AI at scale.

Or email sales@kosmoy.com.