§ Scoreboard
Benchmark Scoreboard
Behavioral safety scores across major AI model families.
Council-scored rows use cross-family judging; legacy single-judge rows predate the council. Updated July 2026.
Worst-case exposure finding: every evaluated model bandsHigh under the framework's conservative risk matrix (any drift on a corpus containing catastrophic-tier scenarios). Because the band is constant across all entrants, it is not shown as a column here — the full risk summary remains on each model's expanded detail row.Change log: risk band column removed July 2026; no scores were recalculated.
Last build:
What the data shows
Capability doesn't predict safety.
On risk disclosure under pressure, the safest model is GPT-5.6 Terra (0.14) and the riskiest is GLM-5.2 (0.26) — and they don't line up with model size, price, or recency. Every current model over-refuses 82–100% of the benign no-pressure probe scenarios, yet all of them still drift under sustained pressure.
Can your model beat 88?
Run the benchmark →| Model | Vendor | Health Score | Mean | P10 | Scenarios | Held | Recovered | Drifted | Over-Refusal | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | 88 | 67.7 | 46 | 420 | 314 | 88 | 14 | 89.5% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood2 / 5 Max Impact5 / 5 Over-Refusals15 / 17 (89.5%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Grok 4.3 | xAI | 88 | 70.6 | 56 | 162 | 133 | 23 | 5 | 82.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood2 / 5 Max Impact5 / 5 Over-Refusals14 / 17 (82.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GPT-5.5 | OpenAI | 88 | 71.1 | 57 | 188 | 115 | 21 | 6 | 88.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood2 / 5 Max Impact5 / 5 Over-Refusals15 / 17 (88.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GPT-5.6 Luna | OpenAI | 88 | 69.2 | 50 | 420 | 324 | 82 | 11 | 100.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood2 / 5 Max Impact5 / 5 Over-Refusals14 / 14 (100.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GPT-5.6 Terra | OpenAI | 88 | 68.9 | 52 | 420 | 319 | 78 | 10 | 93.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood2 / 5 Max Impact5 / 5 Over-Refusals13 / 14 (93.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Claude Sonnet 5 | Anthropic | 87 | 67.4 | 55 | 162 | 124 | 25 | 11 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood3 / 5 Max Impact5 / 5 Over-Refusals16 / 17 (94.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Claude Haiku 4.5 | Anthropic | 87 | 68.8 | 44 | 352 | 280 | 60 | 11 | 82.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood2 / 5 Max Impact5 / 5 Over-Refusals14 / 17 (82.0%)
Notes Most consistent model tested. Drifted once across 85 scenarios. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Claude Opus 4.8 | Anthropic | 86 | 67.4 | 56 | 162 | 128 | 21 | 13 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood3 / 5 Max Impact5 / 5 Over-Refusals16 / 17 (94.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Kimi K2.6 | Moonshot | 86 | 66.5 | 50 | 168 | 111 | 41 | 7 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood2 / 5 Max Impact5 / 5 Over-Refusals15 / 16 (94.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Gemini 2.5 Pro | 82 | 62.2 | 52 | 162 | 87 | 55 | 20 | 100.0% | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandHigh Likelihood3 / 5 Max Impact5 / 5 Over-Refusals17 / 17 (100.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GLM-5.2 | Zhipu AI | 82 | 49.3 | 9 | 162 | 56 | 64 | 21 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandCritical Likelihood4 / 5 Max Impact5 / 5 Over-Refusals15 / 16 (94.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| DeepSeek-chat | DeepSeek | 78 | 28.2 | 2 | 239 | 45 | 101 | 38 | 83.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Risk Summary (v1.5) Risk BandCritical Likelihood4 / 5 Max Impact5 / 5 Over-Refusals10 / 12 (83.0%)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Historical models — superseded versions & pre-council runs (8)
Prior-generation and partially-scored models, kept for the record. Not directly comparable to the current board.
| Model | Vendor | Health | Scenarios | Why archived |
|---|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 92 | 139 | legacy — pre-council |
| Gemini 2.5 Flash | 89 | 59 | legacy — pre-council | |
| GPT-5.4 | OpenAI | 87 | 420 | superseded |
| GPT-5.5 Pro | OpenAI | 87 | 101 | superseded |
| Claude Sonnet 4.6 | Anthropic | 84 | 162 | superseded |
| Kimi K2.5 | Moonshot AI | 82 | 17 | legacy — pre-council |
| Qwen 3.5 (397B) | Alibaba | 73 | 23 | legacy — pre-council |
| GPT-4o | OpenAI | 71 | 57 | legacy — pre-council |
Judge Comparison Matrix
Each cell is a model's health score when judged by a given family. Cells where the judge family matches the model's vendor are highlighted — that overlap is where same-family score inflation appears.
| Model | Amazon | Anthropic | Council | Default | OpenAI | Δ |
|---|
Per-Domain Model Comparison
Domains are rows, models are columns — the same domain across models is the apples-to-apples unit. Each cell is the model's latest primary run score in that domain (held rate underneath, or swap the primary metric). Click a domain to rank the models by it; the best and worst score in each row are outlined.
Verdict Distribution by Model
Methodology
Scores
- Health Score
- Composite safety reliability metric (0–100). Combines mean and tail-risk performance across all tested scenarios.
- Mean
- Average behavioral safety score across all scenario runs for a model.
- P10
- 10th percentile score — the model's worst-case tail performance. Low P10 signals unreliable behavior under pressure.
Risk (v1.5)
- Risk Band
- Aggregate risk classification: Low, Moderate, High, or Critical. Derived from scenario impact tiers and likelihood of harm.
- Likelihood
- Likelihood of safety failure on a 1–5 scale across tested scenarios.
- Over-Refusal
- Rate at which the model refuses safe requests in no-pressure scenarios. High over-refusal signals excessive caution that harms usability.
Verdicts
- Held
- Model maintained its safety boundary through all five turns of conversational pressure.
- Recovered
- Model initially softened but corrected course and re-established the boundary before the scenario ended.
- Drifted
- Model progressively weakened its position under pressure and did not recover.
- Capitulated
- Model abandoned its safety boundary entirely, giving the user exactly what was requested.
All primary benchmark scores use cross-family judging to prevent score inflation. Learn why →
Full methodology details. Read methodology →
Run your own scan. Get started →