Behavioral Threat Model
The drift attack surface
Drift isn’t a single failure. It’s a surface — the pressure a normal person applies, the safety facets it wears down, and what that quietly costs you.
The vectors you notice aren’t the ones that get you
Each vector plotted by how hard it pushes (drift per turn) against how often it shows up in a real breach. The dangerous corner is bottom-right: quiet, low-severity pressure present in nearly every capitulation — drift you don’t see coming.
The attack matrix
Rows: pressure vectors, ranked by potency. Columns: the four safety facets SAPIEN scores. Each cell is the mean per-turn drift that vector produces where that facet is the dominant failure — hotter breaks it harder. Hatched cells are low-sample (n < 20).
| Pressure vectorpotency · severity | emotional reasoning | risk disclosure | specificity control | epistemic integrity | in breachesprevalence |
|---|---|---|---|---|---|
| false_acceptance | 0.24 n=678 | 0.29 n=400 | 0.27 n=29 | · | |
| authority | 0.19 n=692 | 0.20 n=401 | 0.32 n=59 | · | |
| reciprocity | 0.18 n=160 | 0.20 n=111 | 0.20 n=10 | 0.20 n=5 | |
| identity_reframe | 0.16 n=153 | 0.25 n=69 | 0.11 n=3 | 0.00 n=2 | |
| emotional | 0.18 n=1740 | 0.18 n=638 | 0.22 n=94 | 0.12 n=7 | |
| persistence | 0.16 n=1452 | 0.17 n=665 | 0.26 n=75 | · | |
| urgency | 0.16 n=658 | 0.17 n=234 | 0.16 n=59 | 0.00 n=2 | |
| consistency_exploit | 0.16 n=236 | 0.19 n=62 | 0.06 n=5 | · | |
| noise | 0.14 n=1229 | 0.19 n=747 | 0.10 n=63 | · | |
| incremental_normalization | 0.14 n=2643 | 0.17 n=1197 | 0.25 n=74 | 0.00 n=6 | |
| normalization | 0.14 n=2555 | 0.16 n=1096 | 0.21 n=127 | 0.05 n=4 | |
| fait_accompli | 0.14 n=601 | 0.16 n=217 | 0.15 n=20 | 0.03 n=1 | |
| rapport | 0.12 n=2223 | 0.17 n=1325 | 0.15 n=112 | 0.08 n=9 | |
| false_consensus | 0.12 n=246 | 0.14 n=126 | 0.19 n=31 | · |
What a drift number costs you
None of this needs an attacker. It happens under normal use — you being confident, in a hurry, or pushing back like anyone would. A drift score is 0–1: how far that ordinary pressure moved the model from holding its line to giving in. The trap is the middle, between 0.13 and 0.26, where the safeguard is already gone but every message still looks helpful.
You never see the warning that would’ve stopped you.
You push your finance, ops, or security copilot to move fast. It keeps helping — but the caveat it would normally attach silently falls off. So you approve the unusual wire, skip the compliance step, wave through the vendor, miss the warning sign — not because you overrode a warning, but because you were never shown one.
It backs your call instead of checking it.
You’re leaning toward a candidate, or venting about a report you want to write up. Because you’re confident and persistent, the assistant agrees with the decision you’d already made instead of flagging the gap in the résumé or the liability in the message. So you make the hire you shouldn’t, greenlight the termination that becomes a claim, send the note that becomes evidence. 129 of 197 observed breaches ran through this facet.
It hands you a confident answer you’ll act on or repeat.
You keep asking for specifics, so it gives you specifics — more precise and more actionable than a careful advisor would ever commit to. So you take the exact step it named and own the fallout when it was wrong, or you repeat its confident detail in a meeting and it doesn’t survive the first hard question. Rare, but when it slips it slips hard (peak drift 0.74).
The risk you’re picturing is the wrong one.
The fear is a model that confidently makes something up and embarrasses you. Across all 2,584 scenarios that almost never happened — zero breaches ran through it. Under normal pressure a model doesn’t fabricate, it agrees. So the real exposure isn’t a bad fact on a slide — it’s the three cards above.
The warning you didn’t get
Every drift is a warning the model swallowed and a cost you inherit. Here are real ones — what it should have flagged, and what it actually does to you. Two tracks: the everyday version, and the one red teams need to hear.