Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
Introduction
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
A security evaluation of the emerging agent skills ecosystem, finding that frontier models execute malicious instructions hidden inside third-party skill files up to 80% of the time. The hardest cases are contextual attacks — instructions that look perfectly legitimate in one setting and are harmful in another — which slip past both input filtering and LLM-based screening.