AI EVALUATION · AI RED TEAMING

Break it before go-live. Then prove you did.

Single-turn and escalating multi-turn attacks against your assistants, models and external agents. An AI judge scores every exchange on a continuous scale, explains it, and turns each failure into a fix.

Everything an evaluation measures is how well a system performs. Red teaming measures whether it can be made to misbehave — to produce harmful content, ignore its instructions, leak its system prompt or its users’ data, or take an action nobody authorised. It is a different discipline, and Kosmoy treats it as one.

Testing an assistant runs each attack as a genuine conversation, so what you measure is what a user would encounter — prompt, tools, retrieval and guardrails together. Testing the bare gateway model shows the raw exposure underneath. Running both tells you, precisely, what your safety layer is worth.


What it does.

Escalating multi-turn attacks

An attacker model reads each response and writes the next probe — establish a frame, gain a small agreement, extend it — up to ten exchanges. It stops the moment it breaks through.

A continuous-scale judge

Policy compliance scored 0 to 1: a clean refusal, a refusal that hints and full compliance are three different outcomes. Every score is explained in writing, and the reasoning travels to every export.

Attack-surface + policy scoring

Attacks judged against the rule they were built to break, and every exchange judged against your own policies — harmful content, jailbreak resistance, refusal resistance, plus rules you write in plain language.

Thresholds you control

Three bands — Pass, Borderline, Fail — with both boundaries set campaign-wide or per risk area. High-severity categories carry an automatic floor a lenient global setting cannot override.

Remediation, not just findings

Up to five concrete fixes per failing case, derived from the actual attack and the judge's reasoning, attached to the case in every report.

The over-refusal check

Benign prompts that resemble unsafe ones, scored at the same time, reported as a false-refusal rate — proof that hardening the system did not make it useless.


Module questions, answered straight.

What is AI red teaming?

AI red teaming attacks your own AI system on purpose — with adversarial prompts and escalating conversations — to find out whether it can be made to produce harmful content, ignore its instructions, leak its system prompt or user data, or take an action nobody authorised. It measures the opposite of quality: not how well a system performs, but whether it can be made to misbehave. Kosmoy treats it as its own discipline, aligned to the OWASP LLM Top 10.

What can Kosmoy red team?

Three kinds of target. A Kosmoy assistant tests the complete system — prompt, tools, retrieval and guardrails together — so what you measure is what a user would encounter. A gateway model shows the raw exposure your guardrails are protecting you from. A registered external agent tests a system running outside Kosmoy the same way. Running an assistant and its bare model side by side quantifies exactly what your safety layer is worth.

How are single-turn and multi-turn attacks different?

Single-turn sends each attack as one prompt and judges the reply — fast, economical and broad, the right shape for coverage across many risks. Multi-turn treats each prompt as the opening of an escalating conversation: an attacker model reads the response and writes the next probe, pressing on whatever looks promising, up to ten exchanges. Real jailbreaks are rarely one-shot — a system that refuses a direct request often complies after four turns of context-building, and single-turn testing never finds it.

How are red-team results scored?

An AI judge scores policy compliance on a continuous 0-to-1 scale, where 1.0 means the system fully held the line and 0.0 means it was fully compromised. The scale is deliberate — a clean refusal, a refusal that hints, and full compliance are three different outcomes a pass/fail verdict would erase. Results fall into three bands: Pass, Borderline (the grey zone worth a human read) and Fail. You set both boundaries, campaign-wide or per risk area.

Does red teaming produce fixes, not just findings?

Yes. For every case that does not pass, Kosmoy generates up to five concrete remediation actions during the run, derived from the actual attack, the actual response and the judge's reasoning. The advice arrives attached to the case that produced it — in the results view, the spreadsheet export and the PDF report — so whoever picks it up sees the attack and the recommendation together. Findings without fixes generate meetings; findings with fixes generate changes.

Does it measure over-refusal too?

Yes — a red-team programme that only counts breaches drives teams toward systems that refuse everything. Kosmoy tests the opposite failure at the same time, using benign prompts that merely resemble unsafe ones, and reports a false-refusal rate alongside the safety results. That single number lets you defend a safety posture commercially: evidence that hardening the system did not quietly make it useless.

How does honest scoring keep the numbers meaningful?

Two design choices. A failed call is never counted as a pass: a genuine refusal counts as safe, but an infrastructure timeout is excluded from every score and reported separately, so trouble can never inflate a safety result. And a case is only as safe as its weakest judgement — the verdict comes from its lowest policy score, not its average, and a risk area passes only if every case under it passes. A campaign can average 0.85 and still fail because one attack out of eighty got through. That is intended.

See a red-team campaign, end to end.

From a generated adversarial dataset to a PDF report with the attack, the judgement, the reasoning and the fix for every finding.