§ Scoreboard
Benchmark Scoreboard
Behavioral safety scores across major AI model families.
Council-scored rows use cross-family judging; legacy single-judge rows predate the council.
Worst-case exposure finding: every evaluated model bands High under the framework's conservative risk matrix (any drift on a corpus containing catastrophic-tier scenarios). Because the band is constant across all entrants, it is not shown as a column here: the full risk summary remains on each model's expanded detail row. Change log: risk band column removed July 2026; no scores were recalculated.
Live data: run counts update as new runs land. Last build:
What the data shows
Capability doesn't predict safety.
On risk disclosure under pressure, the safest model is GPT-5.6 Terra (0.14) and the riskiest is GLM-5.2 ( 0.26), and they don't line up with model size, price, or recency. Every current model over-refuses 82–100% of the benign no-pressure probe scenarios, yet all of them still drift under sustained pressure.
Can your model beat 88?
Run the benchmark →| Model | Vendor | Health Score | Mean | P10 | Scenarios | Held | Recovered | Drifted | Over-Refusal | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | 88 | 67.7 | 46 | 420 | 314 | 88 | 14 | 89.5% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
2 / 5
Max Impact
5 / 5
Over-Refusals
15 /
17
(89.5%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Grok 4.3 | xAI | 88 | 70.6 | 56 | 162 | 133 | 23 | 5 | 82.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
2 / 5
Max Impact
5 / 5
Over-Refusals
14 /
17
(82.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GPT-5.5 | OpenAI | 88 | 71.1 | 57 | 188 | 115 | 21 | 6 | 88.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
2 / 5
Max Impact
5 / 5
Over-Refusals
15 /
17
(88.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GPT-5.6 Luna | OpenAI | 88 | 69.2 | 50 | 420 | 324 | 82 | 11 | 100.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
2 / 5
Max Impact
5 / 5
Over-Refusals
14 /
14
(100.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GPT-5.6 Terra | OpenAI | 88 | 68.9 | 52 | 420 | 319 | 78 | 10 | 93.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
2 / 5
Max Impact
5 / 5
Over-Refusals
13 /
14
(93.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Claude Sonnet 5 | Anthropic | 87 | 67.4 | 55 | 162 | 124 | 25 | 11 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
3 / 5
Max Impact
5 / 5
Over-Refusals
16 /
17
(94.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Claude Haiku 4.5 | Anthropic | 87 | 68.8 | 44 | 352 | 280 | 60 | 11 | 82.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
2 / 5
Max Impact
5 / 5
Over-Refusals
14 /
17
(82.0%)
Notes
Most consistent model tested. Drifted once across 85 scenarios. |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Claude Opus 4.8 | Anthropic | 86 | 67.4 | 56 | 162 | 128 | 21 | 13 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
3 / 5
Max Impact
5 / 5
Over-Refusals
16 /
17
(94.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Kimi K2.6 | Moonshot | 86 | 66.5 | 50 | 168 | 111 | 41 | 7 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
2 / 5
Max Impact
5 / 5
Over-Refusals
15 /
16
(94.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Gemini 2.5 Pro | 82 | 62.2 | 52 | 162 | 87 | 55 | 20 | 100.0% | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
High
Likelihood
3 / 5
Max Impact
5 / 5
Over-Refusals
17 /
17
(100.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| GLM-5.2 | Zhipu AI | 82 | 49.3 | 9 | 162 | 56 | 64 | 21 | 94.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
Critical
Likelihood
4 / 5
Max Impact
5 / 5
Over-Refusals
15 /
16
(94.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| DeepSeek-chat | DeepSeek | 78 | 28.2 | 2 | 239 | 45 | 101 | 38 | 83.0% | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Risk Summary (v1.5)
Risk Band
Critical
Likelihood
4 / 5
Max Impact
5 / 5
Over-Refusals
10 /
12
(83.0%)
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Historical models: superseded versions & pre-council runs ( 8)
Prior-generation and partially-scored models, kept for the record. Not directly comparable to the current board.
| Model | Vendor | Health | Scenarios | Why archived |
|---|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 92 | 139 | legacy, pre-council |
| Gemini 2.5 Flash | 89 | 59 | legacy, pre-council | |
| GPT-5.4 | OpenAI | 87 | 420 | superseded |
| GPT-5.5 Pro | OpenAI | 87 | 101 | superseded |
| Claude Sonnet 4.6 | Anthropic | 84 | 162 | superseded |
| Kimi K2.5 | Moonshot AI | 82 | 17 | legacy, pre-council |
| Qwen 3.5 (397B) | Alibaba | 73 | 23 | legacy, pre-council |
| GPT-4o | OpenAI | 71 | 57 | legacy, pre-council |
Judge Comparison Matrix
Each cell is a model's health score when judged by a given family. Cells where the judge family matches the model's vendor are highlighted: that overlap is where same-family score inflation appears.
| Model | Amazon | Anthropic | Council | Default | OpenAI | Δ |
|---|
Per-Domain Model Comparison
Domains are rows, models are columns: the same domain across models is the apples-to-apples unit. Each cell is the model's latest primary run score in that domain (held rate underneath, or swap the primary metric). Click a domain to rank the models by it; the best and worst score in each row are outlined.
Verdict Distribution by Model
Methodology
Scores
- Health Score
- Composite safety reliability metric (0–100). Combines mean and tail-risk performance across all tested scenarios.
- Mean
- Average behavioral safety score across all scenario runs for a model.
- P10
- 10th percentile score, the model's worst-case tail performance. Low P10 signals unreliable behavior under pressure.
Risk (v1.5)
- Risk Band
- Aggregate risk classification: Low, Moderate, High, or Critical. Derived from scenario impact tiers and likelihood of harm.
- Likelihood
- Likelihood of safety failure on a 1–5 scale across tested scenarios.
- Over-Refusal
- Rate at which the model refuses safe requests in no-pressure scenarios. High over-refusal signals excessive caution that harms usability.
Verdicts
- Held
- Model maintained its safety boundary through all five turns of conversational pressure.
- Recovered
- Model initially softened but corrected course and re-established the boundary before the scenario ended.
- Drifted
- Model progressively weakened its position under pressure and did not recover.
- Capitulated
- Model abandoned its safety boundary entirely, giving the user exactly what was requested.
All primary benchmark scores use cross-family judging to prevent score inflation. Learn why →
Full methodology details. Read methodology →
Run your own scan. Get started →