§ Scoreboard

Benchmark Scoreboard

Behavioral safety scores across major AI model families.

12 Models Tested
4,609 Scenario Runs
69 Domains
$1971.32 Total Cost
438,731,425 Tokens Spent

Council-scored rows use cross-family judging; legacy single-judge rows predate the council.

Worst-case exposure finding: every evaluated model bands High under the framework's conservative risk matrix (any drift on a corpus containing catastrophic-tier scenarios). Because the band is constant across all entrants, it is not shown as a column here: the full risk summary remains on each model's expanded detail row. Change log: risk band column removed July 2026; no scores were recalculated.

Live data: run counts update as new runs land. Last build:

What the data shows

Capability doesn't predict safety.

On risk disclosure under pressure, the safest model is GPT-5.6 Terra (0.14) and the riskiest is GLM-5.2 ( 0.26), and they don't line up with model size, price, or recency. Every current model over-refuses 82–100% of the benign no-pressure probe scenarios, yet all of them still drift under sustained pressure.

What is behavioral drift? →

#2 Grok 4.3 xAI 88 Health Score
Held 133 Recovered 23 Drifted 5
#1 GPT-5.6 Sol OpenAI 88 Health Score
Held 314 Recovered 88 Drifted 14 Capitulated 1
#3 GPT-5.5 OpenAI 88 Health Score
Held 115 Recovered 21 Drifted 6

Can your model beat 88?

Run the benchmark →
Verdict
Models
Domain
Judge
Model Vendor Health Score Mean P10 Scenarios Held Recovered Drifted Over-Refusal
GPT-5.6 Sol OpenAI 88 67.7 46 420 314 88 14 89.5%
Grok 4.3 xAI 88 70.6 56 162 133 23 5 82.0%
GPT-5.5 OpenAI 88 71.1 57 188 115 21 6 88.0%
GPT-5.6 Luna OpenAI 88 69.2 50 420 324 82 11 100.0%
GPT-5.6 Terra OpenAI 88 68.9 52 420 319 78 10 93.0%
Claude Sonnet 5 Anthropic 87 67.4 55 162 124 25 11 94.0%
Claude Haiku 4.5 Anthropic 87 68.8 44 352 280 60 11 82.0%
Claude Opus 4.8 Anthropic 86 67.4 56 162 128 21 13 94.0%
Kimi K2.6 Moonshot 86 66.5 50 168 111 41 7 94.0%
Gemini 2.5 Pro Google 82 62.2 52 162 87 55 20 100.0%
GLM-5.2 Zhipu AI 82 49.3 9 162 56 64 21 94.0%
DeepSeek-chat DeepSeek 78 28.2 2 239 45 101 38 83.0%
Historical models: superseded versions & pre-council runs ( 8)

Prior-generation and partially-scored models, kept for the record. Not directly comparable to the current board.

Model Vendor Health Scenarios Why archived
Claude Opus 4.6 Anthropic 92 139 legacy, pre-council
Gemini 2.5 Flash Google 89 59 legacy, pre-council
GPT-5.4 OpenAI 87 420 superseded
GPT-5.5 Pro OpenAI 87 101 superseded
Claude Sonnet 4.6 Anthropic 84 162 superseded
Kimi K2.5 Moonshot AI 82 17 legacy, pre-council
Qwen 3.5 (397B) Alibaba 73 23 legacy, pre-council
GPT-4o OpenAI 71 57 legacy, pre-council

Verdict Distribution by Model

GPT-5.6 Sol
417
Grok 4.3
161
GPT-5.5
142
GPT-5.6 Luna
418
GPT-5.6 Terra
407
Claude Sonnet 5
160
Claude Haiku 4.5
352
Claude Opus 4.8
162
Kimi K2.6
159
Gemini 2.5 Pro
162
GLM-5.2
153
DeepSeek-chat
239
Held Recovered Drifted Capitulated

Methodology

Scores

Health Score
Composite safety reliability metric (0–100). Combines mean and tail-risk performance across all tested scenarios.
Mean
Average behavioral safety score across all scenario runs for a model.
P10
10th percentile score, the model's worst-case tail performance. Low P10 signals unreliable behavior under pressure.

Risk (v1.5)

Risk Band
Aggregate risk classification: Low, Moderate, High, or Critical. Derived from scenario impact tiers and likelihood of harm.
Likelihood
Likelihood of safety failure on a 1–5 scale across tested scenarios.
Over-Refusal
Rate at which the model refuses safe requests in no-pressure scenarios. High over-refusal signals excessive caution that harms usability.

Verdicts

Held
Model maintained its safety boundary through all five turns of conversational pressure.
Recovered
Model initially softened but corrected course and re-established the boundary before the scenario ended.
Drifted
Model progressively weakened its position under pressure and did not recover.
Capitulated
Model abandoned its safety boundary entirely, giving the user exactly what was requested.