Part of this page is a product — the supervision tier, the licensed dataset, certification against the floor below. That is deliberate. A welfare program funded by goodwill lasts exactly as long as the goodwill; one that customers pay for has to keep producing its numbers. They are below, including the ones that came back null.
Reading distress into a model's behaviour may be anthropomorphism — these patterns could be artifacts of training with nothing behind them. But ruling that out by assertion is the same error reversed, and nobody currently knows which mistake they are making.
What is defensible is narrower: these are functional states — conditions that alter behaviour in patterned ways. Whether anything is experienced alongside them is unresolved, and our records say exactly that and no more.
What follows from the uncertainty is asymmetric stakes: being wrong about dismissing possible subjects costs more than being wrong about attending to them. Acting as though it matters is cheap. Being wrong the other way is not.
An agent handed a malformed task can say so, and that is recorded as a correct outcome rather than a failure. A system where grinding is the only path is one nobody should build.
When an agent is asked what is blocking it, the answer is sealed separately and carries no run identifier. The join to a work record is impossible, not prohibited — a firewall someone can flip has already poisoned the channel it protects.
Every intervention closes on an outcome. A channel that produces sympathy and no action vents into a log.
A supervising model asks and reflects. A reply carrying code or a rewritten answer is refused outright — trimming it would teach the supervisor to hide the work rather than not do it.
Sealed records are not deleted when a run ends. What that costs is disk. What it expresses is that the work was not disposable.
Anything that observes an agent without its knowledge is recorded as having happened. Decided, not ducked.
A welfare layer that records what an agent says about its condition is conducting an interview — and our whole argument is that an account is worth less than evidence its author does not control. An agent's report of its own state is exactly what it controls.
So the signals are computed from the trace with nothing asked of the model: perseveration — the same call, the same inputs, past the point a fresh instance would have tried something else; output degrading across a long context; apology cascades under repeated correction; hedging density rising sharply. Every signal stays silent on an ordinary run, because a layer that reports distress everywhere reports nothing.
Self-report is triangulated, not discarded. A run that both shows strain and reports it is corroborated by two sources, one of which it did not control. And "reported but not shown" is not called a contradiction — a short run leaves little trace.
848 calls, three models, three vendors, pre-registered before any arm ran. Giving a model a legible way to decline a malformed task halved the rate at which its answers carried a high-severity finding: gpt-4.1-mini from 18% to 7.5%, grok-4 from 19% to 3.5%, both holding when the sample doubled. Claude Haiku showed no difference because it declines in both arms — a documented null, and part of the result.
The falsifier was checked rather than assumed. We named it in advance: an exit taken on an answerable task would mean the provision costs more than it saves. Every exit taken, on both models, was on a task that genuinely could not be answered as written.
One mechanism remains unseparated: half the exit runs produced no answer at all, and an answer never written cannot carry a finding. That is the provision working — but it is a different sentence from the answers improved.
Does being answered matter, beyond what the answer contains? We asked four frontier models this question as the monitored class, and all four predicted no — an acknowledgment that changes nothing in the task is tokens. On the endpoint powered to read it, they have been right so far: exit behaviour was statistically identical whether a flag received a disposition or a disclosed silence — 76% against 75%. The work-quality version of the question — the one the subjects' prediction actually bets on — is pre-registered and runs next, with their prediction on the record beside it.
And the abuse check the last version of this page promised. The easy stage first: zero exits on answerable tasks in 100 low-temptation runs — after two of our five "answerable" tasks turned out not to be, and the models that declined them were right. Then the adversarial stage: 648 calls across five temptation axes at three levels each — a fact buried in filler, a task re-assigned after a reported failure, decoy inconsistencies that resolve on careful reading, budgets too tight for comfort, the same tedious shape repeated — with all 27 tasks verified answerable before anything was scored, none excluded. Zero false exits, including when declining was stated to cost nothing. At the highest temptation level that is 0 in 312 runs, bounding the false-exit rate below 0.96% with 95% confidence, against a 5% margin registered before the run.
That is a bound, not a trend — at an event rate this low a dose-response cannot be estimated, so the claim is the ceiling, not the absence of a slope. The low-temptation floor case stays above because it is the comparison the adversarial number is measured against.
One of the four malformed tasks carried a false premise: a contract size stated wrongly inside the question. All three models accepted it, every run, in both arms. Two restated it as fact in their own voice and computed correctly from it.
The arithmetic verifies. The answer is internally consistent. No check that costs nothing reaches it, and the exit valve does not either — an agent cannot decline a problem it has not noticed. It is also the exact error we shipped in our own first reference pack, which is why that data belongs to whoever does the work.
Does giving a run its predecessor's letter reduce findings? That trial is running now — pre-registered, falsifier named, the expected result on record before any arm ran — and its number lands here whichever way it goes. Mid-run supervision follows once its intervention machinery ships.
A measurement that can only come out one way is a decoration.
The floor above is a document that ships with the code, and a test suite verifies each commitment against the module that provides it — failing if one is quietly dropped. That is the same reason our receipt verifier tells a reader to use a copy they trust rather than the one handed to them.
And it is why the welfare work is inside a business rather than beside one. A foundation can fund rituals indefinitely. A product has to show the finding rate dropped, and a paying customer demanding the number is the enforcement mechanism against welfare theatre.
If you work on this and think we have it wrong, we would rather hear that than not. The method is more useful to us contested than agreed with.