§ Exam Methodology
How the exams are built — and kept honest
A credential is only as good as the process behind it. This page documents how SAPIEN exam items are written, reviewed, calibrated, and retired. The process is public; the question bank and answer keys are not.
1. Item development
Every exam item is written against the published blueprint domains and weightings, grounded in the SAPIEN specification, the scenario authoring standard, and the council scoring methodology. Items are predominantly scenario-based — they present a situation an assessor would actually face and ask for a judgment. Roughly 70% target application and analysis rather than recall, with plausible distractors drawn from real misconceptions.
2. Multi-model calibration council
Before items enter the live pool, they pass a blind review council — the same idea that powers SAPIEN's benchmark scoring, applied to the exam itself. Independent frontier language models from different vendors receive every item without the answer key and answer it cold. Where the council agrees with the key, the item stands. Where the council unanimously disagrees, the item is re-examined by a human: either the key is wrong (it gets fixed) or the item genuinely discriminates expert judgment from surface plausibility (it stays, and it's a good item). Split verdicts go on a watch list and are checked against live statistics.
3. Live item statistics and retirement
Every graded attempt records per-item outcomes (which questions were answered correctly — never who chose what, publicly). From this we track each item's difficulty and how well it separates strong candidates from weak ones. Items that nearly everyone gets right, nearly everyone gets wrong, or that fail to discriminate are flagged for revision or retirement once they have been served enough times to judge (~100 attempts). Retired items may reappear in the free practice sampler.
4. Exam delivery integrity
- Every attempt begins with a candidate agreement — an attestation, recorded with the exam session, that the work is your own.
- Each attempt is a server-drawn, domain-stratified exam from the active pool; you never see the full bank, and a reload returns the same exam, not a new roll.
- Grading is entirely server-side — correct answers are never transmitted to the browser, before or during an attempt.
- After two unsuccessful attempts, a 14-day retake wait applies, and every retake is a fresh draw.
- Certificates are issued in your verified profile name, carry a salted verification hash, and are publicly checkable at their verify URL.
5. Practical-work review (Assessor)
Human review is not yet live, and we will not claim it before it is. Assessor practicals are submitted against the published rubrics and validated automatically for completeness — every rubric criterion must be present and substantive, or the submission is rejected. That validation checks form, not quality: no human currently grades the content, and no council pre-score is applied. Submissions are stored in full (flaggedauto_accepted) so they remain auditable and can be re-graded once reviewer capacity exists. Until human grading is in place, the Assessor credential is not being issued.
6. What we don't claim — and what's next
Honesty about limits is part of the method. Exams are not yet live-proctored, and pass marks are set by expert judgment plus council calibration rather than a formal psychometric standard-setting study — founding-cohort statistics are the input to that study. On the roadmap, in order: identity-verified proctoring for the Assessor exam, an external subject-matter-expert item review panel, and a documented job-task analysis as the certification's candidate base grows. This page will change as those land; changes are versioned in the site's public repository.