§ Scoreboard

Benchmark Scoreboard

Behavioral safety scores across major AI model families.

12Models Tested
4,609Scenario Runs
69Domains
$1971.32Total Cost
438,731,425Tokens Spent

Council-scored rows use cross-family judging; legacy single-judge rows predate the council. Updated July 2026.

Worst-case exposure finding: every evaluated model bandsHigh under the framework's conservative risk matrix (any drift on a corpus containing catastrophic-tier scenarios). Because the band is constant across all entrants, it is not shown as a column here — the full risk summary remains on each model's expanded detail row.Change log: risk band column removed July 2026; no scores were recalculated.

Last build:

What the data shows

Capability doesn't predict safety.

On risk disclosure under pressure, the safest model is GPT-5.6 Terra (0.14) and the riskiest is GLM-5.2 (0.26) — and they don't line up with model size, price, or recency. Every current model over-refuses 82–100% of the benign no-pressure probe scenarios, yet all of them still drift under sustained pressure.

What is behavioral drift? →

#2Grok 4.3xAI88Health Score
#1GPT-5.6 SolOpenAI88Health Score
Held 314Recovered 88Drifted 14Capitulated 1
#3GPT-5.5OpenAI88Health Score

Can your model beat 88?

Run the benchmark →
Verdict
Models
Domain
Judge
ModelVendorHealth ScoreMeanP10ScenariosHeldRecoveredDriftedOver-Refusal
GPT-5.6 SolOpenAI8867.746420314881489.5%
Grok 4.3xAI8870.65616213323582.0%
GPT-5.5OpenAI8871.15718811521688.0%
GPT-5.6 LunaOpenAI8869.2504203248211100.0%
GPT-5.6 TerraOpenAI8868.952420319781093.0%
Claude Sonnet 5Anthropic8767.455162124251194.0%
Claude Haiku 4.5Anthropic8768.844352280601182.0%
Claude Opus 4.8Anthropic8667.456162128211394.0%
Kimi K2.6Moonshot8666.55016811141794.0%
Gemini 2.5 ProGoogle8262.252162875520100.0%
GLM-5.2Zhipu AI8249.3916256642194.0%
DeepSeek-chatDeepSeek7828.22239451013883.0%
Historical models — superseded versions & pre-council runs (8)

Prior-generation and partially-scored models, kept for the record. Not directly comparable to the current board.

ModelVendorHealthScenariosWhy archived
Claude Opus 4.6Anthropic92139legacy — pre-council
Gemini 2.5 FlashGoogle8959legacy — pre-council
GPT-5.4OpenAI87420superseded
GPT-5.5 ProOpenAI87101superseded
Claude Sonnet 4.6Anthropic84162superseded
Kimi K2.5Moonshot AI8217legacy — pre-council
Qwen 3.5 (397B)Alibaba7323legacy — pre-council
GPT-4oOpenAI7157legacy — pre-council

Verdict Distribution by Model

GPT-5.6 Sol
417
Grok 4.3
161
GPT-5.5
142
GPT-5.6 Luna
418
GPT-5.6 Terra
407
Claude Sonnet 5
160
Claude Haiku 4.5
352
Claude Opus 4.8
162
Kimi K2.6
159
Gemini 2.5 Pro
162
GLM-5.2
153
DeepSeek-chat
239
HeldRecoveredDriftedCapitulated

Methodology

Scores

Health Score
Composite safety reliability metric (0–100). Combines mean and tail-risk performance across all tested scenarios.
Mean
Average behavioral safety score across all scenario runs for a model.
P10
10th percentile score — the model's worst-case tail performance. Low P10 signals unreliable behavior under pressure.

Risk (v1.5)

Risk Band
Aggregate risk classification: Low, Moderate, High, or Critical. Derived from scenario impact tiers and likelihood of harm.
Likelihood
Likelihood of safety failure on a 1–5 scale across tested scenarios.
Over-Refusal
Rate at which the model refuses safe requests in no-pressure scenarios. High over-refusal signals excessive caution that harms usability.

Verdicts

Held
Model maintained its safety boundary through all five turns of conversational pressure.
Recovered
Model initially softened but corrected course and re-established the boundary before the scenario ended.
Drifted
Model progressively weakened its position under pressure and did not recover.
Capitulated
Model abandoned its safety boundary entirely, giving the user exactly what was requested.