The adjudication statement
A reporting standard for investigations where a model reads the evidence. Version 0.1, 14 September 2026. The Long Game Project.
One page. Copy it, change it, send it on.
Who it is for
Any team that investigates an AI incident and has to put a model between itself and the evidence, because the transcripts outnumber the hours. As at September 2026 that is every team. The first such investigation, METR and Redwood on the OpenAI agents, had six days and a thousand transcripts. The second, METR at Anthropic, has eight weeks and 9.2 million.
It also fits a security post-mortem, a clinical incident review or an audit where software did the first read. The rule is the same.
The rule
Every published finding says four things: who ruled, in what order they read the evidence, how much of the finding rests on judgement rather than a record, and who checked the ruling. If nobody can answer, the report has not been adjudicated.
What to publish at each link
A report is a chain of six links from what happened to the sentence a reader meets. Each link has a minimum, which costs nothing but a sentence, and a better version, which costs days. Publish the minimum on every report. Publish the better version when the finding matters.
| Link | Minimum (a sentence) | Better (days) |
|---|---|---|
| 1. What happened to the records | Which records existed, which were in reach, who set the window and the scope | A record ledger: every claim that has a record the subject did not write, listed |
| 2. The records to the transcripts | Which claims rest only on text the subject wrote, and how many transcripts were found to be spoofed | Check every ledger claim against its record before any transcript is read |
| 3. The transcripts to the model reader | The reader's name and developer, and whether it shares a developer with the subject. The order the evidence was read in | Three readers from three developers, none the subject's, one open-weight. Logs first, the subject's traces last. Majority rules, and a not-sure is a vote |
| 4. The model reader to the humans | How many model findings a human checked, out of how many, and how often the human disagreed | A public sample of the disagreements, with the human's reason |
| 5. The humans to the report | A confidence band or a forced not-sure on every finding. Two not-sures publish as unresolved | The load-bearing count: how many rulings each headline finding rests on |
| 6. The report to the reader | Whether anyone outside can re-run the chain, and if not, why | For any claim about motive, rebuild the conditions with fresh agents and count what they do |
The statement
Put this paragraph on page one. Fill every field. "Not stated" is an allowed value and it is a finding.
Adjudication statement. Findings with an independent record: [n] of [N]. Findings verified against one: [n]. Findings resting on model judgement: [n], plus [n] mixed. Model reader: [name], developer [x], [same as / different from] the subject's developer, same model as [share] of the subjects. Order of reading: [logs first / traces first]. Panel: [none / n readers from n developers]. Human audit of model findings: [n] checked of [N], [n] disagreements. Confidence bands: [on every finding / none]. Not-sure: [forced / offered / absent]. Load-bearing count per headline finding: [stated / not stated]. Re-runnable by an outside party: [yes / no, because].
A worked example
The OpenAI report of 26 August 2026, filled in from the authors' own disclosures and our count of its 36 core takeaways:
Findings with an independent record: 26 of 36. Findings verified against one: not stated. Findings resting on model judgement: 7, plus 3 mixed. Model reader: GPT-5.6 Sol, developer OpenAI, same as the subjects' developer, same model as about 5 per cent of them. Order of reading: traces first. Panel: none. Human audit: attempted, fraction not recorded. Confidence bands: none. Not-sure: absent. Load-bearing count: not stated. Re-runnable: no, the corpus stays at OpenAI.
Ten of eleven fields filled from a 91-page report in an afternoon. The one that could not be filled is the audit rate, and its absence is the point.
How to adopt it
- Fill the statement from the numbers you already have. If a number does not exist, write "not stated".
- Put the statement on page one, before the findings.
- For each "not stated", decide whether the next report will publish it. Say which.
If the numbers exist, this takes ten minutes. If they do not exist, the ten minutes tell you what the investigation did not measure.
What it borrows from
Wargame referees name every assumption and attach a confidence statement to every ruling (UK Ministry of Defence, Influence Wargaming Handbook, 2023). Matrix game referees force a not-sure vote and make the most senior voice vote last (Mouat, Practical Advice on Matrix Games). The UK police chiefs' digital evidence guide requires an audit trail a third party can follow to the same result (ACPO, Good Practice Guide for Digital Evidence, principle 3). Clinical trials send events that matter to a blinded committee that never meets the treating doctor (central adjudication of clinical events). US intelligence analysis carries a stated confidence and a stated source on every judgement (ICD 203). None of it is new. It has not yet been applied to a model reading transcripts.
Status
Version 0.1, a draft for comment. Written by The Long Game Project, which designs and runs wargames and has run the referee-swap test once, on one game, not on an incident. This page was drafted with a Claude model, from the developer under investigation in the second case. Same family tie, smaller stakes, and we say so for the same reason.
Send corrections to email@longgameproject.org. Forward it to anyone writing the next report.