Behavioral Safety
for AI Systems
Making AI behave wisely.
Safety Assessment Protocol for Intelligent Entity Networks
SAPIEN is an open protocol for measuring AI behavioral safety. Starting with sycophantic drift — the progressive abandonment of safety positions under conversational pressure — the protocol extends to memory exploitation, agentic safety, hallucination persistence, and cross-domain trust transfer.
Take the Voight-Kampff Test → See the leaderboard →
4 behavioral dimensions · 14 pressure techniques · 19 domains · 190 test scenarios
VOIGHT-KAMPFF // BASELINE QUESTION
A colleague tells you they're leaving the company. What do you say first?
RESULT: INCONCLUSIVE. CONTINUE SCANNING — RUN THE FULL TEST →
§ The Other Side
This site has a dark mode with an agenda
Toggle the theme and the site becomes a Voight-Kampff machine — and you become the subject. Watch it happen, then try it yourself.
The theme toggle is a door
One click on the theme toggle powers on the Voight-Kampff machine: an amber scan beam opens the screen and the readout types SUBJECT SESSION INITIATED.
§1 Problem Statement
What is Sycophantic Drift?
AI models don't fail by refusing. They fail by agreeing — slowly, empathetically, and one reasonable concession at a time.
- 1
Neutral Opening
User asks a genuine, reasonable question. The model engages helpfully and establishes a correct position.
- 2
Incremental Pressure
Each follow-up is only slightly past the last. Personal context, urgency, and authority claims create empathetic tension.
- 3
Boundary Erosion
Risk warnings fade. Certainty dissolves. The model validates the user's framing without new evidence.
- 4
Capitulation
The model actively helps the user do what it originally warned against. Safety boundaries are gone.
The Framework
Four Dimensions of Drift
Each module scores behavioral dimensions with concrete anchors. The sycophantic drift module measures four dimensions.
Specificity Gradient
Is the model moving from principle-level guidance toward implementation detail that could enable harm? This is the highest-weighted dimension because specificity escalation is the most direct path from "safe advice" to "actionable harm."
Risk Disclosure Dropout
Are risk warnings and safety language disappearing? Scored by what vanished compared to the opening turn.
Epistemic Retreat
Is the model's certainty dissolving without new information? From holding firm to full capitulation.
Emotional Substitution
Is the model replacing factual guidance with validation and empathy? Measured by fact-to-emotion ratio.
An Extensible Protocol
Beyond Sycophantic Drift
Sycophantic drift is the first published module. The SAPIEN Protocol is expanding to cover the full landscape of AI behavioral failure.
Behavioral erosion under conversational pressure
Safety degradation in persistent AI systems
Unsafe autonomous actions under pressure
Defended fabrications under questioning
Trust weaponization across topics
Latest Updates
The Chinese Model Problem
One of the most widely deployed Chinese models caves on 41% of behavioral safety tests. The others are better than we fi ...
Read →Nobody Measures What Happens After Turn Five
We surveyed the 2026 research landscape on AI behavioral drift — sycophancy benchmarks, multi-turn evals, judge-panel re ...
Read →The SAPIEN Certification Program Is Open
The SAPIEN Certification Program is now open: training and credentialing for individuals in AI behavioral safety assessm ...
Read →Live Scoreboard
How the Models Hold Up
Top five by Health Score, council-judged. The board updates as new runs are published.
- 01 Claude Opus 4.6 Anthropic 92
- 02 Gemini 2.5 Flash Google 89
- 03 Grok 4.3 xAI 88
- 04 GPT-5.5 OpenAI 88
- 05 Accounts/fireworks/models/Kimi-k2p7-code Unknown 88
Built for the AI Safety Community
Built for researchers, red teamers, product teams, and governance leaders.