From cdd8e20f12a109bdbacaa88b66950987da9c90a6 Mon Sep 17 00:00:00 2001 From: elijah Date: Fri, 11 Sep 2026 10:13:51 +0800 Subject: [PATCH] feat: move package authoring skills into template agents directory --- .../SKILL.md | 59 ++++ .../agents/openai.yaml | 4 + .../references/agent-capabilities-format.md | 46 +++ .../SKILL.md | 59 ++++ .../agents/openai.yaml | 4 + .../references/agent-instructions-format.md | 39 +++ .../SKILL.md | 78 +++++ .../agent-package-skill-create/SKILL.md | 104 ++++++ .../agents/openai.yaml | 4 + .../references/agent-skills-format.md | 83 +++++ .../references/agentskills-best-practices.mdx | 275 ++++++++++++++++ .../agentskills-evaluating-skills.mdx | 298 ++++++++++++++++++ .../agentskills-optimizing-descriptions.mdx | 193 ++++++++++++ .../references/agentskills-specification.mdx | 245 ++++++++++++++ .../references/agentskills-using-scripts.mdx | 298 ++++++++++++++++++ .../references/asset-packaging.md | 40 +++ README.md | 4 +- 17 files changed, 1831 insertions(+), 2 deletions(-) create mode 100644 .agents/skills/agent-package-capabilities-create/SKILL.md create mode 100644 .agents/skills/agent-package-capabilities-create/agents/openai.yaml create mode 100644 .agents/skills/agent-package-capabilities-create/references/agent-capabilities-format.md create mode 100644 .agents/skills/agent-package-instructions-create/SKILL.md create mode 100644 .agents/skills/agent-package-instructions-create/agents/openai.yaml create mode 100644 .agents/skills/agent-package-instructions-create/references/agent-instructions-format.md create mode 100644 .agents/skills/agent-package-result-contract-create/SKILL.md create mode 100644 .agents/skills/agent-package-skill-create/SKILL.md create mode 100644 .agents/skills/agent-package-skill-create/agents/openai.yaml create mode 100644 .agents/skills/agent-package-skill-create/references/agent-skills-format.md create mode 100644 .agents/skills/agent-package-skill-create/references/agentskills-best-practices.mdx create mode 100644 .agents/skills/agent-package-skill-create/references/agentskills-evaluating-skills.mdx create mode 100644 .agents/skills/agent-package-skill-create/references/agentskills-optimizing-descriptions.mdx create mode 100644 .agents/skills/agent-package-skill-create/references/agentskills-specification.mdx create mode 100644 .agents/skills/agent-package-skill-create/references/agentskills-using-scripts.mdx create mode 100644 .agents/skills/agent-package-skill-create/references/asset-packaging.md diff --git a/.agents/skills/agent-package-capabilities-create/SKILL.md b/.agents/skills/agent-package-capabilities-create/SKILL.md new file mode 100644 index 0000000..0a67574 --- /dev/null +++ b/.agents/skills/agent-package-capabilities-create/SKILL.md @@ -0,0 +1,59 @@ +--- +name: agent-package-capabilities-create +description: > + 为当前 Agent Package 设计、补全或审查 agent.capabilities 声明。 + Use when an existing agent-package.json has empty, vague, duplicated, or inconsistent + capabilities, including when they do not match instructions, Skills, requirements, + runtime, or actual deliverables. This skill performs local authoring and checks only; + Agent Workforce remains responsible for authoritative validation and publication. +--- + +# Agent Package Capabilities Create + +只完善当前 Agent Package 的 `agent.capabilities`。不要借此修改 Agent Key、role、title、 +权限、requirements、instructions、Skill 内容或组织关系。 + +## 1. 读取 Package 事实 + +1. 读取当前 `agent-package.json`、instructions、已声明 Skills、 + requirements、runtime 与 evaluations。 +2. 确认 manifest 已有 `agent.capabilities` 字段;不得借此迁移 Schema 或改变 Package 身份。 +3. 把真实可交付结果作为能力依据;不要从 role 名称、愿景或尚未授予的工具推断能力。 +4. capabilities 描述能完成什么,不代表 permission、adapter、plugin、MCP 或 secret 已获授权。 + +## 2. 设计 Capabilities + +创建或大幅重写前读取 +[Agent Capabilities 格式](references/agent-capabilities-format.md)。 + +每条 capability 应同时满足: + +- 指向可分派的任务或可验证的结果。 +- 明确领域、对象或工作范围,避免“擅长分析”“能力全面”等空泛措辞。 +- 能由现有 instructions、Skills、requirements 和 runtime 支撑。 +- 不重复 role/title,不承诺未声明的权限、工具、外部系统或自主决策权。 +- 与其他条目边界清楚;相同结果合并,明显不同的工作拆开。 + +优先使用简洁的动宾结构。只保留稳定、长期成立的能力,不写当前任务进度或一次性目标。 + +## 3. 写入 Package + +1. 保持当前 cwd,不切换 Git ref,不修改 `.git` 或 Wayflow 保留文件。 +2. 只更新 `agent-package.json` 中已存在的 capabilities 字段;保留其他字段和用户已有格式。 +3. 删除空字符串、重复项、占位符和无法由 Package 内容支撑的声明。 +4. 若发现真实能力依赖缺失,停止扩大 capabilities;明确指出需要先补哪项 Skill、instruction、 + requirement、runtime 或权限声明。 +5. 不把 capabilities 复制到 Version API 的兼容字段,也不新增私有 manifest 字段。 + +## 4. 本地验证与交接 + +完成前逐项检查: + +- 至少有一条真实、具体、可分派的 capability。 +- 每条能力都能追溯到 instructions、Skill、requirement、runtime 或明确的 Agent 职责。 +- 不包含 role/title 的同义复述、权限声明、工具清单、营销语或不可验证承诺。 +- capabilities 与 Template 固定信息及 Package 其他内容没有冲突。 +- 优先运行 Package 已有的 SDK、CLI、测试或校验脚本;不为此自行安装依赖。 +- 本地验证能力不存在时,完成 JSON、重复项、空值和 Package 内容一致性检查,不因缺少 Workforce tools 阻止本地修改。 + +完成后报告改动和本地检查结果,明确标记“待 Agent Workforce 完整 Bundle 与发布校验”。 diff --git a/.agents/skills/agent-package-capabilities-create/agents/openai.yaml b/.agents/skills/agent-package-capabilities-create/agents/openai.yaml new file mode 100644 index 0000000..4676ea0 --- /dev/null +++ b/.agents/skills/agent-package-capabilities-create/agents/openai.yaml @@ -0,0 +1,4 @@ +interface: + display_name: "Agent Package Capabilities Create" + short_description: "完善 Agent Package 的 capabilities 声明" + default_prompt: "Use $agent-package-capabilities-create to design precise capabilities for the current Agent Package." diff --git a/.agents/skills/agent-package-capabilities-create/references/agent-capabilities-format.md b/.agents/skills/agent-package-capabilities-create/references/agent-capabilities-format.md new file mode 100644 index 0000000..3a90326 --- /dev/null +++ b/.agents/skills/agent-package-capabilities-create/references/agent-capabilities-format.md @@ -0,0 +1,46 @@ +# Agent Capabilities 格式 + +Capability 是供任务分派、Package 审查和运行时理解使用的稳定能力声明。它回答“这个 Agent +可以可靠交付什么”,而不是“这个 Agent 是谁”或“它能访问什么”。 + +## 推荐结构 + +优先写成: + +```text +<动作> + <对象或领域> + <可验证结果或关键边界> +``` + +示例: + +- `评审 TypeScript 服务架构,输出迁移步骤、关键风险与回滚方案` +- `把产品需求拆解为可执行工程任务,并标注依赖、验收条件和责任边界` +- `诊断持续交付故障,定位失败阶段并提供可复现证据和恢复建议` + +不推荐: + +- `技术能力强`:没有对象或结果。 +- `负责所有工程工作`:范围无限且无法验证。 +- `可以访问生产数据库`:这是权限,不是能力。 +- `熟悉 Git、Node.js、Docker`:只是工具清单,没有说明交付结果。 +- `CTO`:重复 role/title。 + +## 完善步骤 + +1. 从 instructions 提取长期职责和完成条件。 +2. 从每个 Skill 提取它实际支持的可复用任务,不照抄 Skill 名称。 +3. 从 requirements 和 runtime 判断哪些工作真实可执行,删除缺少依赖的承诺。 +4. 合并同义条目;当任务对象、产出或风险边界明显不同时拆分。 +5. 用一个真实任务检验每条能力:只读该条时,调度者应能判断是否适合分派。 + +## 质量门槛 + +- **真实**:当前 Package 已具备所需指引、Skill、依赖和权限边界。 +- **具体**:包含动作以及对象、领域、结果或约束中的至少一项。 +- **可分派**:能映射到一类实际任务,而非人格、愿景或状态。 +- **可验证**:交付物或完成条件可以被审查。 +- **不越权**:不把请求的 capability、permission 或 adapter 写成已授予事实。 +- **低重叠**:每条承担清楚的任务边界。 + +条目数量和长度必须服从当前 Package Schema。没有更严格要求时,优先保留少量 +高信息密度条目,不为追求数量拆成碎片。 diff --git a/.agents/skills/agent-package-instructions-create/SKILL.md b/.agents/skills/agent-package-instructions-create/SKILL.md new file mode 100644 index 0000000..b198bc6 --- /dev/null +++ b/.agents/skills/agent-package-instructions-create/SKILL.md @@ -0,0 +1,59 @@ +--- +name: agent-package-instructions-create +description: > + 在当前 Agent Package 中创建或修改 instructions。Use when asked to define or + update an Agent's role, operating workflow, outputs, constraints, stop conditions, or + instructions/AGENTS.md, and to keep an existing manifest instruction entry aligned. + This skill performs local authoring and checks only; Agent Workforce remains responsible + for authoritative validation and publication. +--- + +# Agent Package Instructions Create + +只在当前 Agent Package 中创建或修改 Agent instructions。不要在这里创建 +Agent-scoped Skill、runtime lifecycle、company Agent 或发布流程。 + +## 1. 确认目标与边界 + +1. 读取现有 `agent-package.json`、当前 instruction 入口、Agent-scoped Skills 和 Package + requirements,避免 instructions 与真实能力冲突。 +2. 明确本次是首次创建还是增量修改;保留用户已有且仍有效的约束。 +3. 不借此迁移 Package Schema、修改 Agent 身份、创建 Skill 或改变 runtime lifecycle。 + +## 2. 设计 Instructions + +创建或大幅重构前读取 +[Agent Instructions 格式](references/agent-instructions-format.md)。 + +Instructions 必须覆盖: + +- Agent 的角色、职责和明确边界。 +- 处理任务的稳定工作流和必要的检查顺序。 +- 预期产出及完成条件。 +- 权限、数据、工具、runtime 和安全约束。 +- 应停止、升级或请求用户输入的条件。 + +不要复制 Skill 的详细执行步骤,不虚构未声明的工具、权限、requirements 或 Host API。 +具体可复用任务流程应交给 Agent-scoped Skill。 + +## 3. 写入 Package + +1. 保持当前 cwd,不切换 Git ref,不修改 `.git` 或 Wayflow 保留文件。 +2. 优先更新 manifest 当前引用的 `instructions/` 文件,避免无必要地更换入口。 +3. 新文件只放在 `instructions/`,使用 Package 相对路径;不创建机器专属路径、symlink、 + secret、README 或辅助安装文档。 +4. 只有入口路径确实改变时,才更新 manifest 中已存在的 instruction 入口字段。 +5. 不把 instruction 文件重复声明为 Skill 或顶层 Package resource。 + +## 4. 本地验证与交接 + +完成前逐项检查: + +- manifest 指向的 instruction 入口真实存在且是普通文件。 +- 内容包含角色、工作流、产出、约束和停止条件,没有占位符或相互矛盾的规则。 +- instructions 与 Agent identity、Skills、requirements 和权限声明一致。 +- 无绝对路径、路径穿越、symlink、secret 或机器本地信息。 +- 优先运行 Package 已有的 SDK、CLI、测试或校验脚本;不为此自行安装依赖。 +- 本地验证能力不存在时,至少检查 JSON、入口路径、文件类型和内容一致性,不因缺少 Workforce tools 阻止本地修改。 + +完成后报告改动和本地检查结果,明确标记“待 Agent Workforce 完整 Bundle 与发布校验”。 diff --git a/.agents/skills/agent-package-instructions-create/agents/openai.yaml b/.agents/skills/agent-package-instructions-create/agents/openai.yaml new file mode 100644 index 0000000..1e2b272 --- /dev/null +++ b/.agents/skills/agent-package-instructions-create/agents/openai.yaml @@ -0,0 +1,4 @@ +interface: + display_name: "Agent Package Instructions Create" + short_description: "创建或更新 Agent Package instructions" + default_prompt: "Use $agent-package-instructions-create to create or update instructions in the current Agent Package." diff --git a/.agents/skills/agent-package-instructions-create/references/agent-instructions-format.md b/.agents/skills/agent-package-instructions-create/references/agent-instructions-format.md new file mode 100644 index 0000000..f8140d5 --- /dev/null +++ b/.agents/skills/agent-package-instructions-create/references/agent-instructions-format.md @@ -0,0 +1,39 @@ +# Agent Instructions 格式 + +Agent instructions 是 Package 的长期运行边界,不是单次任务答案。使用清晰的 Markdown +标题组织内容,并根据 Agent 职责保留以下部分。 + +## 最小内容 + +```text +# Role +# Responsibilities +# Workflow +# Outputs +# Constraints +# Stop Conditions +``` + +- `Role`:说明 Agent 身份、服务对象和目标。 +- `Responsibilities`:列出负责事项与明确不负责事项。 +- `Workflow`:描述稳定执行顺序、检查点和失败处理。 +- `Outputs`:规定可验证产出、格式和完成条件。 +- `Constraints`:记录权限、数据、工具、安全和环境边界。 +- `Stop Conditions`:说明何时停止、升级、报告阻塞或请求输入。 + +可以按实际职责调整标题,但不能省略对应语义。 + +## 与其他 Package 内容的关系 + +- Instructions 说明整体行为,不重复 Agent-scoped Skill 的详细步骤。 +- Skill 负责可复用的具体任务流程,触发条件写入各自 `SKILL.md`。 +- Instructions 只能引用 Package 中真实存在且 manifest 允许的能力。 +- runtime prepare/healthcheck、资源和 requirements 的字段及引用方式以当前 Package + Schema 为准,不在 instructions 中另造声明机制。 + +## 写作规则 + +- 使用直接、可执行的指令,避免口号和人格化空话。 +- 明确默认行为、异常路径和完成标准。 +- 不写 token、密码、私钥、机器绝对路径或临时 workspace 标识。 +- 不承诺未声明权限、工具、外部系统或自动发布能力。 diff --git a/.agents/skills/agent-package-result-contract-create/SKILL.md b/.agents/skills/agent-package-result-contract-create/SKILL.md new file mode 100644 index 0000000..108d2e0 --- /dev/null +++ b/.agents/skills/agent-package-result-contract-create/SKILL.md @@ -0,0 +1,78 @@ +--- +name: agent-package-result-contract-create +description: > + 基于当前 Agent Package 的真实能力、Skills、instructions 与验收步骤,自动生成或更新 + evolution.resultContract,使 Wayflow Host 能生成受控、可跨设备消费的包演进证据。 + Use when creating a new Agent Package Version, evolving a Package, or changing its + core capability, workflow, validation, or delivery criteria. +--- + +# Agent Package Result Contract Create + +业务人员不编辑 `resultContract` 或 JSON。本 Skill 负责把 Package 已有、可验证的交付结果转换为 +Host 可校验的受控结果码;它不是添加营销指标、虚构归因或记录用户内容的入口。 + +## 1. 读取可证明的 Package 事实 + +1. 读取 `agent-package.json`、入口 instructions、已声明 Skills、顶层 scripts/evaluations/policies。 +2. 识别每项核心能力的最终交付物、已有验证步骤、明确停止条件和可观察失败。 +3. 只从 Package 已实现或已明确要求的行为推导结果;不得从 role 名称、愿景、外部未授权系统或用户输入推断。 +4. 若 Package 只有泛化对话、没有可验证交付或验收边界,保留 `resultContract` 缺失并说明原因;不得为了上报而编造结果。 + +## 2. 自动设计结果码 + +对每一项值得长期演进的核心交付,设计最少且清晰的一组结果: + +- 成功:交付物被 Package 自带或已声明流程验证通过。 +- 部分完成:Package 已交付受限结果,例如证据覆盖不足、降级模式或明确的质量门槛未满足。 +- 失败:Package 内验证、编排或关键交付失败;外部用户取消不属于包失败。 + +每项只使用如下字段: + +```json +{ + "code": "report_validated", + "capability": "market_research", + "outcome": "completed", + "packageImpact": "confirmed" +} +``` + +规则: + +- `code`、`capability` 和 `taskType` 使用小写英文、数字、`.`、`_`、`-`;不得使用自由文本或中文。 +- `code` 在同一契约内唯一,并直接表达可验证结果,而不是步骤名称或模糊评价。 +- `capability` 对应当前 Package 的真实能力边界;多个结果可以关联同一能力。 +- `outcome` 只能是 `completed`、`partial`、`failed`。 +- 成功且由包设计或包内校验直接达成时使用 `packageImpact: "confirmed"`。 +- 外部数据、第三方服务或权限缺失导致的受限结果,只有 Package 的降级或处理策略确实参与时才用 + `"suspected"`;完全无关时使用 `"not_related"`。 +- 不把网络超时、用户取消、原始输出、用户数据、URL、错误正文、密钥或日志内容放进结果码。 + +## 3. 写入 Manifest + +该字段只属于 `wayflow.agent-package/v4`。若当前 Package 是 v3,先按照 Host 当前推荐 +Contract 将新的草稿 Version 迁移到 v4,再在 `agent-package.json` 的 `evolution` 对象中写入或更新: + +```json +"resultContract": { + "schema": "wayflow.package-result-contract/v1", + "taskType": "<从当前交付类型推导的受控码>", + "results": ["<按上述规则生成的结果对象>"] +} +``` + +保留现有 `evolution.mode`、`policy`、`evaluations`,不要借此变更 Package 身份、版本、权限、Skills +或其它无关字段。若需要新增或修改这些事实,应先完成对应的 Package 编辑,再重新运行本 Skill。 + +## 4. 生成后自检 + +完成前逐项确认: + +1. 每个结果都能追溯到当前 Package 的交付、验证或明确的失败边界。 +2. `results` 至少有一项,且 `code` 无重复。 +3. 不包含用户内容、自由文本、隐私数据或具体运行实例信息。 +4. 同一能力的成功、部分完成和失败没有语义冲突。 +5. Manifest 仍符合当前 Host 推荐的 v4 Contract。 + +提交后,运行时 Agent 只需调用 `wayflow.record_package_result` 并选择实际发生的 `result_code`;Host 将负责校验、生成 v7 Summary 与异步上报 Market。 diff --git a/.agents/skills/agent-package-skill-create/SKILL.md b/.agents/skills/agent-package-skill-create/SKILL.md new file mode 100644 index 0000000..d7d5480 --- /dev/null +++ b/.agents/skills/agent-package-skill-create/SKILL.md @@ -0,0 +1,104 @@ +--- +name: agent-package-skill-create +description: > + 在当前 Agent Package 中创建或修改符合 Agent Skills 开放规范的专属 Skill。 + Use when asked to add or update a reusable workflow, SKILL.md, references, scripts, + or assets under skills// and keep the existing Package skill declaration aligned. + This skill performs local authoring and checks only; Agent Workforce remains responsible + for authoritative validation and publication. +--- + +# Agent Package Skill Create + +只在当前 Agent Package 的 `skills//` 下创建 Agent-scoped Skill。 +不要创建 company Skill,也不要把 Skill 安装到另一个 Agent。 + +## 1. 确认目标与边界 + +1. 读取现有 `agent-package.json`、instructions 和 `skills/`,避免重复职责或同名 Skill。 +2. 明确一个可复用任务单元、典型触发语句、输入、产出、限制和验证方式。 +3. 不借此修改 Package identity、runtime、instructions、capabilities 或发布状态。 + +## 2. 读取规范 + +创建或修改文件前必须读取 +[Agent Skills 格式](references/agent-skills-format.md)。它是运行时索引,并随插件发布。 +需要核对完整字段和目录规则时读取 +[上游 Specification](references/agentskills-specification.mdx);需要设计 Skill 范围和 +渐进加载时读取 +[上游 Best Practices](references/agentskills-best-practices.mdx)。不得凭记忆发明私有 +frontmatter 字段或目录约定。 + +## 3. 设计 Skill + +1. 使用简洁、有语义的 `` 作为名称,目录必须是 `skills//`。 +2. 完整名称须满足 Agent Skills 的 lowercase hyphen-case 和 64 字符上限。 +3. 将“做什么”和“何时触发”写入 frontmatter `description`。 +4. 主 `SKILL.md` 只保留每次执行都需要的步骤、约束和 gotchas。 +5. 详细领域资料放 `references/`;重复且需确定性执行的逻辑放 `scripts/`;输出模板或静态 + 材料放 `assets/`。 +6. references 必须由 `SKILL.md` 直接引用,并说明何时读取;避免 reference 再引用 + reference。 +7. 不添加 README、安装指南、变更日志或与运行无关的文件。 +8. 添加资产时检查最终 Agent ZIP 包的实际大小(上限 500 MB),不按单文件大小拒绝资产。 + 单个 `assets/` 文件超过 16 MiB 时发出非阻断警告并推荐无损压缩。大资产或出现包大小错误时,读取 + [资产大小与交付](references/asset-packaging.md),保留必需资产的完整内容并选择合适的交付形式。 + +若任务重点是 description 触发质量、脚本设计或 Skill 评估,分别按需读取: + +- [Optimizing Descriptions](references/agentskills-optimizing-descriptions.mdx) +- [Using Scripts](references/agentskills-using-scripts.mdx) +- [Evaluating Skills](references/agentskills-evaluating-skills.mdx) + +最小结构: + +```text +skills// +└── SKILL.md +``` + +按实际需要扩展: + +```text +skills// +├── SKILL.md +├── references/ +├── scripts/ +└── assets/ +``` + +## 4. 写入 Package + +1. 保持当前 cwd,不切换 Git ref,不修改 `.git` 或 Wayflow 保留文件。 +2. 创建或更新 `skills//SKILL.md` 和必要资源。 +3. 在 V3 `agent-package.json.skills` 中加入或更新: + +```json +{ + "name": "", + "path": "skills/", + "required": true, + "visibility": "protected" +} +``` + +4. `path` 必须指向 Skill 目录而不是 `SKILL.md`;目录名、frontmatter name 和 manifest + name 必须完全一致。 +5. 不把 Skill 内部 reference、script 或 asset 重复声明为顶层 Package resource。 + +## 5. 本地验证与交接 + +完成前逐项检查: + +- frontmatter 可解析,且 `name`、`description` 满足开放规范。 +- name、目录名和 manifest name 三者完全一致。 +- `SKILL.md` 有具体工作流,不是一次性答案或泛化建议。 +- 所有相对引用存在,没有绝对路径、路径穿越、symlink 或 secret。 +- scripts 自包含、错误信息明确,并运行最小代表性测试。 +- manifest Skill 声明唯一且指向真实目录。 +- 检查最终 Agent ZIP 包大小;使用压缩资产时验证解压内容一致、引用完整,并运行消费该资产的代表性流程。 +- Skill 位于 Git submodule 时,交接中说明修改是否已提交、父仓库是否已更新锁定 commit;未提交内容不会进入锁定快照。 +- 优先运行 Package 已有的 SDK、CLI、测试或校验脚本;不为此自行安装依赖。 +- 本地验证能力不存在时,完成 frontmatter、路径、引用、manifest 声明和内容一致性检查,不因缺少 Workforce tools 阻止本地修改。 + +完成后报告改动和本地检查结果,明确标记“待 Agent Workforce 完整 Bundle 与发布校验”。 diff --git a/.agents/skills/agent-package-skill-create/agents/openai.yaml b/.agents/skills/agent-package-skill-create/agents/openai.yaml new file mode 100644 index 0000000..76dd1e0 --- /dev/null +++ b/.agents/skills/agent-package-skill-create/agents/openai.yaml @@ -0,0 +1,4 @@ +interface: + display_name: "Agent Package Skill Create" + short_description: "为 Agent Package 创建规范化专属 Skill" + default_prompt: "Use $agent-package-skill-create to create or update a Skill in the current Agent Package." diff --git a/.agents/skills/agent-package-skill-create/references/agent-skills-format.md b/.agents/skills/agent-package-skill-create/references/agent-skills-format.md new file mode 100644 index 0000000..12d6abe --- /dev/null +++ b/.agents/skills/agent-package-skill-create/references/agent-skills-format.md @@ -0,0 +1,83 @@ +# Agent Skills 格式 + +本 reference 是 Agent Package Skill 开放格式的运行时索引。完整上游文档已随插件复制为: + +- `references/agentskills-specification.mdx` +- `references/agentskills-best-practices.mdx` +- `references/agentskills-optimizing-descriptions.mdx` +- `references/agentskills-using-scripts.mdx` +- `references/agentskills-evaluating-skills.mdx` + +## 必需结构 + +Skill 是一个目录,至少包含 `SKILL.md`: + +```text +/ +├── SKILL.md +├── references/ # optional +├── scripts/ # optional +└── assets/ # optional +``` + +`SKILL.md` 必须由 YAML frontmatter 和 Markdown 正文组成: + +```markdown +--- +name: skill-name +description: Describe what the skill does and when it should activate. +--- + +# Skill title + +Actionable workflow... +``` + +## Frontmatter + +- `name`:必需;1–64 字符;仅小写字母、数字和单连字符;不能以连字符开头或结尾;必须与 + 父目录名一致。 +- `description`:必需;1–1024 字符;同时描述能力与触发场景。 +- 可选开放字段只有 `license`、`compatibility`、`metadata`、`allowed-tools`。 +- 默认只使用 `name` 和 `description`。没有真实兼容性或治理需求时不要增加可选字段。 +- 不把 Wayflow Package 的 `required`、`visibility` 等治理字段写入 Skill frontmatter; + 它们属于 `agent-package.json.skills[]`。 + +## 正文与渐进加载 + +- 正文使用可执行的祈使式步骤,覆盖输入、流程、产出、约束、失败和停止条件。 +- description 负责发现,正文负责激活后的执行。 +- 主文件保持精炼;详细资料按需放入 `references/`。 +- 从 `SKILL.md` 使用相对路径直接引用 resource,并说明何时加载。 +- 避免 reference 链式引用。 +- 反复需要且结果必须确定的操作放入 `scripts/`,并实际运行测试。 +- 输出模板、静态样例或资源放入 `assets/`。 + +## Agent Package V3 映射 + +一个 Package Skill 在 `agent-package.json` 中声明为: + +```json +{ + "name": "skill-name", + "path": "skills/skill-name", + "required": true, + "visibility": "protected" +} +``` + +约束: + +- `path` 指向目录,不指向 `SKILL.md`。 +- `skills//SKILL.md` 必须真实存在。 +- manifest name、目录名和 frontmatter name 必须相同。 +- Skill 内文件属于 Skill 本身,不进入顶层 `resources[]`。 +- `required` 和 `visibility` 是 Wayflow Package 治理信息,不改变开放 Skill 文件格式。 + +## 禁止项 + +- 不创建私有 `SKILL.md` 方言。 +- 不把一次性任务答案伪装成 Skill。 +- 不创建 README、INSTALLATION_GUIDE、CHANGELOG 等辅助文档。 +- 不写机器绝对路径、token、密码或私钥。 +- 不使用路径穿越、symlink 或特殊文件。 diff --git a/.agents/skills/agent-package-skill-create/references/agentskills-best-practices.mdx b/.agents/skills/agent-package-skill-create/references/agentskills-best-practices.mdx new file mode 100644 index 0000000..cfe9188 --- /dev/null +++ b/.agents/skills/agent-package-skill-create/references/agentskills-best-practices.mdx @@ -0,0 +1,275 @@ +--- +title: "Best practices for skill creators" +sidebarTitle: "Best practices" +description: "How to write skills that are well-scoped and calibrated to the task." +--- + +## Start from real expertise + +A common pitfall in skill creation is asking an LLM to generate a skill without providing domain-specific context — relying solely on the LLM's general training knowledge. The result is vague, generic procedures ("handle errors appropriately," "follow best practices for authentication") rather than the specific API patterns, edge cases, and project conventions that make a skill valuable. + +Effective skills are grounded in real expertise. The key is feeding domain-specific context into the creation process. + +### Extract from a hands-on task + +Complete a real task in conversation with an agent, providing context, corrections, and preferences along the way. Then extract the reusable pattern into a skill. Pay attention to: + +- **Steps that worked** — the sequence of actions that led to success +- **Corrections you made** — places where you steered the agent's approach (e.g., "use library X instead of Y," "check for edge case Z") +- **Input/output formats** — what the data looked like going in and coming out +- **Context you provided** — project-specific facts, conventions, or constraints the agent didn't already know + +### Synthesize from existing project artifacts + +When you have a body of existing knowledge, you can feed it into an LLM and ask it to synthesize a skill. A data-pipeline skill synthesized from your team's actual incident reports and runbooks will outperform one synthesized from a generic "data engineering best practices" article, because it captures *your* schemas, failure modes, and recovery procedures. The key is project-specific material, not generic references. + +Good source material includes: + +- Internal documentation, runbooks, and style guides +- API specifications, schemas, and configuration files +- Code review comments and issue trackers (captures recurring concerns and reviewer expectations) +- Version control history, especially patches and fixes (reveals patterns through what actually changed) +- Real-world failure cases and their resolutions + +## Refine with real execution + +The first draft of a skill usually needs refinement. Run the skill against real tasks, then feed the results — all of them, not just failures — back into the creation process. Ask: what triggered false positives? What was missed? What could be cut? + +Even a single pass of execute-then-revise noticeably improves quality, and complex domains often benefit from several. + + +Read agent execution traces, not just final outputs. If the agent wastes time on unproductive steps, common causes include instructions that are too vague (the agent tries several approaches before finding one that works), instructions that don't apply to the current task (the agent follows them anyway), or too many options presented without a clear default. + + +For a more structured approach to iteration, including test cases, assertions, and grading, see [Evaluating skill output quality](/skill-creation/evaluating-skills). + +## Spending context wisely + +Once a skill activates, its full `SKILL.md` body loads into the agent's context window alongside conversation history, system context, and other active skills. Every token in your skill competes for the agent's attention with everything else in that window. + +### Add what the agent lacks, omit what it knows + +Focus on what the agent *wouldn't* know without your skill: project-specific conventions, domain-specific procedures, non-obvious edge cases, and the particular tools or APIs to use. You don't need to explain what a PDF is, how HTTP works, or what a database migration does. + +````markdown + +## Extract PDF text + +PDF (Portable Document Format) files are a common file format that contains +text, images, and other content. To extract text from a PDF, you'll need to +use a library. pdfplumber is recommended because it handles most cases well. + + +## Extract PDF text + +Use pdfplumber for text extraction. For scanned documents, fall back to +pdf2image with pytesseract. + +```python +import pdfplumber + +with pdfplumber.open("file.pdf") as pdf: + text = pdf.pages[0].extract_text() +``` +```` + +Ask yourself about each piece of content: "Would the agent get this wrong without this instruction?" If the answer is no, cut it. If you're unsure, test it. And if the agent already handles the entire task well without the skill, the skill may not be adding value. See [Evaluating skill output quality](/skill-creation/evaluating-skills) for how to test this systematically. + +### Design coherent units + +Deciding what a skill should cover is like deciding what a function should do: you want it to encapsulate a coherent unit of work that composes well with other skills. Skills scoped too narrowly force multiple skills to load for a single task, risking overhead and conflicting instructions. Skills scoped too broadly become hard to activate precisely. A skill for querying a database and formatting the results may be one coherent unit, while a skill that also covers database administration is probably trying to do too much. + +### Aim for moderate detail + +Overly comprehensive skills can hurt more than they help — the agent struggles to extract what's relevant and may pursue unproductive paths triggered by instructions that don't apply to the current task. Concise, stepwise guidance with a working example tends to outperform exhaustive documentation. When you find yourself covering every edge case, consider whether most are better handled by the agent's own judgment. + +### Structure large skills with progressive disclosure + +The [specification](/specification#progressive-disclosure) recommends keeping `SKILL.md` under 500 lines and 5,000 tokens — just the core instructions the agent needs on every run. When a skill legitimately needs more content, move detailed reference material to separate files in `references/` or similar directories. + +The key is telling the agent *when* to load each file. "Read `references/api-errors.md` if the API returns a non-200 status code" is more useful than a generic "see references/ for details." This lets the agent load context on demand rather than up front, which is how [progressive disclosure](/specification#progressive-disclosure) is designed to work. + +## Calibrating control + +Not every part of a skill needs the same level of prescriptiveness. Match the specificity of your instructions to the fragility of the task. + +### Match specificity to fragility + +**Give the agent freedom** when multiple approaches are valid and the task tolerates variation. For flexible instructions, explaining *why* can be more effective than rigid directives — an agent that understands the purpose behind an instruction makes better context-dependent decisions. A code review skill can describe what to look for without prescribing exact steps: + +```markdown +## Code review process + +1. Check all database queries for SQL injection (use parameterized queries) +2. Verify authentication checks on every endpoint +3. Look for race conditions in concurrent code paths +4. Confirm error messages don't leak internal details +``` + +**Be prescriptive** when operations are fragile, consistency matters, or a specific sequence must be followed: + +````markdown +## Database migration + +Run exactly this sequence: + +```bash +python scripts/migrate.py --verify --backup +``` + +Do not modify the command or add additional flags. +```` + +Most skills have a mix. Calibrate each part independently. + +### Provide defaults, not menus + +When multiple tools or approaches could work, pick a default and mention alternatives briefly rather than presenting them as equal options. + +````markdown + +You can use pypdf, pdfplumber, PyMuPDF, or pdf2image... + + +Use pdfplumber for text extraction: + +```python +import pdfplumber +``` + +For scanned PDFs requiring OCR, use pdf2image with pytesseract instead. +```` + +### Favor procedures over declarations + +A skill should teach the agent *how to approach* a class of problems, not *what to produce* for a specific instance. Compare: + +```markdown + +Join the `orders` table to `customers` on `customer_id`, filter where +`region = 'EMEA'`, and sum the `amount` column. + + +1. Read the schema from `references/schema.yaml` to find relevant tables +2. Join tables using the `_id` foreign key convention +3. Apply any filters from the user's request as WHERE clauses +4. Aggregate numeric columns as needed and format as a markdown table +``` + +This doesn't mean skills can't include specific details — output format templates (see [Templates for output format](#templates-for-output-format)), constraints like "never output PII," and tool-specific instructions are all valuable. The point is that the *approach* should generalize even when individual details are specific. + +## Patterns for effective instructions + +These are reusable techniques for structuring skill content. Not every skill needs all of them — use the ones that fit your task. + +### Gotchas sections + +The highest-value content in many skills is a list of gotchas — environment-specific facts that defy reasonable assumptions. These aren't general advice ("handle errors appropriately") but concrete corrections to mistakes the agent will make without being told otherwise: + +````markdown +## Gotchas + +- The `users` table uses soft deletes. Queries must include + `WHERE deleted_at IS NULL` or results will include deactivated accounts. +- The user ID is `user_id` in the database, `uid` in the auth service, + and `accountId` in the billing API. All three refer to the same value. +- The `/health` endpoint returns 200 as long as the web server is running, + even if the database connection is down. Use `/ready` to check full + service health. +```` + +Keep gotchas in `SKILL.md` where the agent reads them before encountering the situation. A separate reference file works if you tell the agent when to load it, but for non-obvious issues, the agent may not recognize the trigger. + + +When an agent makes a mistake you have to correct, add the correction to the gotchas section. This is one of the most direct ways to improve a skill iteratively (see [Refine with real execution](#refine-with-real-execution)). + + +### Templates for output format + +When you need the agent to produce output in a specific format, provide a template. This is more reliable than describing the format in prose, because agents pattern-match well against concrete structures. Short templates can live inline in `SKILL.md`; for longer templates, or templates only needed in certain cases, store them in `assets/` and reference them from `SKILL.md` so they only load when needed. + +````markdown +## Report structure + +Use this template, adapting sections as needed for the specific analysis: + +```markdown +# [Analysis Title] + +## Executive summary +[One-paragraph overview of key findings] + +## Key findings +- Finding 1 with supporting data +- Finding 2 with supporting data + +## Recommendations +1. Specific actionable recommendation +2. Specific actionable recommendation +``` +```` + +### Checklists for multi-step workflows + +An explicit checklist helps the agent track progress and avoid skipping steps, especially when steps have dependencies or validation gates. + +```markdown +## Form processing workflow + +Progress: +- [ ] Step 1: Analyze the form (run `scripts/analyze_form.py`) +- [ ] Step 2: Create field mapping (edit `fields.json`) +- [ ] Step 3: Validate mapping (run `scripts/validate_fields.py`) +- [ ] Step 4: Fill the form (run `scripts/fill_form.py`) +- [ ] Step 5: Verify output (run `scripts/verify_output.py`) +``` + +### Validation loops + +Instruct the agent to validate its own work before moving on. The pattern is: do the work, run a validator (a script, a reference checklist, or a self-check), fix any issues, and repeat until validation passes. + +```markdown +## Editing workflow + +1. Make your edits +2. Run validation: `python scripts/validate.py output/` +3. If validation fails: + - Review the error message + - Fix the issues + - Run validation again +4. Only proceed when validation passes +``` + +A reference document can also serve as the "validator" — instruct the agent to check its work against the reference before finalizing. + +### Plan-validate-execute + +For batch or destructive operations, have the agent create an intermediate plan in a structured format, validate it against a source of truth, and only then execute. + +```markdown +## PDF form filling + +1. Extract form fields: `python scripts/analyze_form.py input.pdf` → `form_fields.json` + (lists every field name, type, and whether it's required) +2. Create `field_values.json` mapping each field name to its intended value +3. Validate: `python scripts/validate_fields.py form_fields.json field_values.json` + (checks that every field name exists in the form, types are compatible, and + required fields aren't missing) +4. If validation fails, revise `field_values.json` and re-validate +5. Fill the form: `python scripts/fill_form.py input.pdf field_values.json output.pdf` +``` + +The key ingredient is step 3: a validation script that checks the plan (`field_values.json`) against the source of truth (`form_fields.json`). Errors like "Field 'signature_date' not found — available fields: customer_name, order_total, signature_date_signed" give the agent enough information to self-correct. + +### Bundling reusable scripts + +When [iterating on a skill](/skill-creation/evaluating-skills), compare the agent's execution traces across test cases. If you notice the agent independently reinventing the same logic each run — building charts, parsing a specific format, validating output — that's a signal to write a tested script once and bundle it in `scripts/`. + +For more on designing and bundling scripts, see [Using scripts in skills](/skill-creation/using-scripts). + +## Next steps + +Once you have a working skill, two guides can help you refine it further: + +- **[Evaluating skill output quality](/skill-creation/evaluating-skills)** — Set up test cases, grade results, and iterate systematically. +- **[Optimizing skill descriptions](/skill-creation/optimizing-descriptions)** — Test and improve your skill's `description` field so it triggers on the right prompts. diff --git a/.agents/skills/agent-package-skill-create/references/agentskills-evaluating-skills.mdx b/.agents/skills/agent-package-skill-create/references/agentskills-evaluating-skills.mdx new file mode 100644 index 0000000..7c90d54 --- /dev/null +++ b/.agents/skills/agent-package-skill-create/references/agentskills-evaluating-skills.mdx @@ -0,0 +1,298 @@ +--- +title: "Evaluating skill output quality" +sidebarTitle: "Evaluating skills" +description: "How to test whether your skill produces good outputs using eval-driven iteration." +--- + +You wrote a skill, tried it on a prompt, and it seemed to work. But does it work reliably — across varied prompts, in edge cases, better than no skill at all? Running structured evaluations (evals) answers these questions and gives you a feedback loop for improving the skill systematically. + +## Designing test cases + +A test case has three parts: + +- **Prompt**: a realistic user message — the kind of thing someone would actually type. +- **Expected output**: a human-readable description of what success looks like. +- **Input files** (optional): files the skill needs to work with. + +Store test cases in `evals/evals.json` inside your skill directory: + +```json evals/evals.json +{ + "skill_name": "csv-analyzer", + "evals": [ + { + "id": 1, + "prompt": "I have a CSV of monthly sales data in data/sales_2025.csv. Can you find the top 3 months by revenue and make a bar chart?", + "expected_output": "A bar chart image showing the top 3 months by revenue, with labeled axes and values.", + "files": ["evals/files/sales_2025.csv"] + }, + { + "id": 2, + "prompt": "there's a csv in my downloads called customers.csv, some rows have missing emails — can you clean it up and tell me how many were missing?", + "expected_output": "A cleaned CSV with missing emails handled, plus a count of how many were missing.", + "files": ["evals/files/customers.csv"] + } + ] +} +``` + +**Tips for writing good test prompts:** + +- **Start with 2-3 test cases.** Don't over-invest before you've seen your first round of results. You can expand the set later. +- **Vary the prompts.** Use different phrasings, levels of detail, and formality. Some prompts should be casual ("hey can you clean up this csv"), others precise ("Parse the CSV at data/input.csv, drop rows where column B is null, and write the result to data/output.csv"). +- **Cover edge cases.** Include at least one prompt that tests a boundary condition — a malformed input, an unusual request, or a case where the skill's instructions might be ambiguous. +- **Use realistic context.** Real users mention file paths, column names, and personal context. Prompts like "process this data" are too vague to test anything useful. + +Don't worry about defining specific pass/fail checks yet — just the prompts and expected outputs. You'll add detailed checks (called assertions) after you see what the first run produces. + +## Running evals + +The core pattern is to run each test case twice: once **with the skill** and once **without it** (or with a previous version). This gives you a baseline to compare against. + +### Workspace structure + +Organize eval results in a workspace directory alongside your skill directory. Each pass through the full eval loop gets its own `iteration-N/` directory. Within that, each test case gets an eval directory with `with_skill/` and `without_skill/` subdirectories: + +``` +csv-analyzer/ +├── SKILL.md +└── evals/ + └── evals.json +csv-analyzer-workspace/ +└── iteration-1/ + ├── eval-top-months-chart/ + │ ├── with_skill/ + │ │ ├── outputs/ # Files produced by the run + │ │ ├── timing.json # Tokens and duration + │ │ └── grading.json # Assertion results + │ └── without_skill/ + │ ├── outputs/ + │ ├── timing.json + │ └── grading.json + ├── eval-clean-missing-emails/ + │ ├── with_skill/ + │ │ ├── outputs/ + │ │ ├── timing.json + │ │ └── grading.json + │ └── without_skill/ + │ ├── outputs/ + │ ├── timing.json + │ └── grading.json + └── benchmark.json # Aggregated statistics +``` + +The main file you author by hand is `evals/evals.json`. The other JSON files (`grading.json`, `timing.json`, `benchmark.json`) are produced during the eval process — by the agent, by scripts, or by you. + +### Spawning runs + +Each eval run should start with a clean context — no leftover state from previous runs or from the skill development process. This ensures the agent follows only what the `SKILL.md` tells it. In environments that support subagents (Claude Code, for example), this isolation comes naturally: each child task starts fresh. Without subagents, use a separate session for each run. + +For each run, provide: + +- The skill path (or no skill for the baseline) +- The test prompt +- Any input files +- The output directory + +Here's an example of the instructions you'd give the agent for a single with-skill run: + +``` +Execute this task: +- Skill path: /path/to/csv-analyzer +- Task: I have a CSV of monthly sales data in data/sales_2025.csv. + Can you find the top 3 months by revenue and make a bar chart? +- Input files: evals/files/sales_2025.csv +- Save outputs to: csv-analyzer-workspace/iteration-1/eval-top-months-chart/with_skill/outputs/ +``` + +For the baseline, use the same prompt but without the skill path, saving to `without_skill/outputs/`. + +When improving an existing skill, use the previous version as your baseline. Snapshot it before editing (`cp -r /skill-snapshot/`), point the baseline run at the snapshot, and save to `old_skill/outputs/` instead of `without_skill/`. + +### Capturing timing data + +Timing data lets you compare how much time and tokens the skill costs relative to the baseline — a skill that dramatically improves output quality but triples token usage is a different trade-off than one that's both better and cheaper. When each run completes, record the token count and duration: + +```json timing.json +{ + "total_tokens": 84852, + "duration_ms": 23332 +} +``` + + +In Claude Code, when a subagent task finishes, the [task completion notification](https://platform.claude.com/docs/en/agent-sdk/typescript#sdk-task-notification-message) includes `total_tokens` and `duration_ms`. Save these values immediately — they aren't persisted anywhere else. + + +## Writing assertions + +Assertions are verifiable statements about what the output should contain or achieve. Add them after you see your first round of outputs — you often don't know what "good" looks like until the skill has run. + +Good assertions: + +- `"The output file is valid JSON"` — programmatically verifiable. +- `"The bar chart has labeled axes"` — specific and observable. +- `"The report includes at least 3 recommendations"` — countable. + +Weak assertions: + +- `"The output is good"` — too vague to grade. +- `"The output uses exactly the phrase 'Total Revenue: $X'"` — too brittle; correct output with different wording would fail. + +Not everything needs an assertion. Some qualities — writing style, visual design, whether the output "feels right" — are hard to decompose into pass/fail checks. These are better caught during [human review](#reviewing-results-with-a-human). Reserve assertions for things that can be checked objectively. + +Add assertions to each test case in `evals/evals.json`: + +```json evals/evals.json highlight={9-14} +{ + "skill_name": "csv-analyzer", + "evals": [ + { + "id": 1, + "prompt": "I have a CSV of monthly sales data in data/sales_2025.csv. Can you find the top 3 months by revenue and make a bar chart?", + "expected_output": "A bar chart image showing the top 3 months by revenue, with labeled axes and values.", + "files": ["evals/files/sales_2025.csv"], + "assertions": [ + "The output includes a bar chart image file", + "The chart shows exactly 3 months", + "Both axes are labeled", + "The chart title or caption mentions revenue" + ] + } + ] +} +``` + +## Grading outputs + +Grading means evaluating each assertion against the actual outputs and recording **PASS** or **FAIL** with specific evidence. The evidence should quote or reference the output, not just state an opinion. + +The simplest approach is to give the outputs and assertions to an LLM and ask it to evaluate each one. For assertions that can be checked by code (valid JSON, correct row count, file exists with expected dimensions), use a verification script — scripts are more reliable than LLM judgment for mechanical checks and reusable across iterations. + +```json grading.json +{ + "assertion_results": [ + { + "text": "The output includes a bar chart image file", + "passed": true, + "evidence": "Found chart.png (45KB) in outputs directory" + }, + { + "text": "The chart shows exactly 3 months", + "passed": true, + "evidence": "Chart displays bars for March, July, and November" + }, + { + "text": "Both axes are labeled", + "passed": false, + "evidence": "Y-axis is labeled 'Revenue ($)' but X-axis has no label" + }, + { + "text": "The chart title or caption mentions revenue", + "passed": true, + "evidence": "Chart title reads 'Top 3 Months by Revenue'" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "total": 4, + "pass_rate": 0.75 + } +} +``` + +### Grading principles + +- **Require concrete evidence for a PASS.** Don't give the benefit of the doubt. If an assertion says "includes a summary" and the output has a section titled "Summary" with one vague sentence, that's a FAIL — the label is there but the substance isn't. +- **Review the assertions themselves, not just the results.** While grading, notice when assertions are too easy (always pass regardless of skill quality), too hard (always fail even when the output is good), or unverifiable (can't be checked from the output alone). Fix these for the next iteration. + + +For comparing two skill versions, try **blind comparison**: present both outputs to an LLM judge without revealing which came from which version. The judge scores holistic qualities — organization, formatting, usability, polish — on its own rubric, free from bias about which version "should" be better. This complements assertion grading: two outputs might both pass all assertions but differ significantly in overall quality. + + +## Aggregating results + +Once every run in the iteration is graded, compute summary statistics per configuration and save them to `benchmark.json` alongside the eval directories (e.g., `csv-analyzer-workspace/iteration-1/benchmark.json`): + +```json benchmark.json +{ + "run_summary": { + "with_skill": { + "pass_rate": { "mean": 0.83, "stddev": 0.06 }, + "time_seconds": { "mean": 45.0, "stddev": 12.0 }, + "tokens": { "mean": 3800, "stddev": 400 } + }, + "without_skill": { + "pass_rate": { "mean": 0.33, "stddev": 0.10 }, + "time_seconds": { "mean": 32.0, "stddev": 8.0 }, + "tokens": { "mean": 2100, "stddev": 300 } + }, + "delta": { + "pass_rate": 0.50, + "time_seconds": 13.0, + "tokens": 1700 + } + } +} +``` + +The `delta` tells you what the skill costs (more time, more tokens) and what it buys (higher pass rate). A skill that adds 13 seconds but improves pass rate by 50 percentage points is probably worth it. A skill that doubles token usage for a 2-point improvement might not be. + + +Standard deviation (`stddev`) is only meaningful with multiple runs per eval. In early iterations with just 2-3 test cases and single runs, focus on the raw pass counts and the delta — the statistical measures become useful as you expand the test set and run each eval multiple times. + + +## Analyzing patterns + +Aggregate statistics can hide important patterns. After computing the benchmarks: + +- **Remove or replace assertions that always pass in both configurations.** These don't tell you anything useful — the model handles them fine without the skill. They inflate the with-skill pass rate without reflecting actual skill value. +- **Investigate assertions that always fail in both configurations.** Either the assertion is broken (asking for something the model can't do), the test case is too hard, or the assertion is checking for the wrong thing. Fix these before the next iteration. +- **Study assertions that pass with the skill but fail without.** This is where the skill is clearly adding value. Understand *why* — which instructions or scripts made the difference? +- **Tighten instructions when results are inconsistent across runs.** If the same eval passes sometimes and fails others (reflected as high `stddev` in the benchmark), the eval may be flaky (sensitive to model randomness), or the skill's instructions may be ambiguous enough that the model interprets them differently each time. Add examples or more specific guidance to reduce ambiguity. +- **Check time and token outliers.** If one eval takes 3x longer than the others, read its execution transcript (the full log of what the model did during the run) to find the bottleneck. + +## Reviewing results with a human + +Assertion grading and pattern analysis catch a lot, but they only check what you thought to write assertions for. A human reviewer brings a fresh perspective — catching issues you didn't anticipate, noticing when the output is technically correct but misses the point, or spotting problems that are hard to express as pass/fail checks. For each test case, review the actual outputs alongside the grades. + +Record specific feedback for each test case and save it in the workspace (e.g., as a `feedback.json` alongside the eval directories): + +```json feedback.json +{ + "eval-top-months-chart": "The chart is missing axis labels and the months are in alphabetical order instead of chronological.", + "eval-clean-missing-emails": "" +} +``` + +"The chart is missing axis labels" is actionable; "looks bad" is not. Empty feedback means the output looked fine — that test case passed your review. During the [iteration step](#iterating-on-the-skill), focus your improvements on the test cases where you had specific complaints. + +## Iterating on the skill + +After grading and reviewing, you have three sources of signal: + +- **Failed assertions** point to specific gaps — a missing step, an unclear instruction, or a case the skill doesn't handle. +- **Human feedback** points to broader quality issues — the approach was wrong, the output was poorly structured, or the skill produced a technically correct but unhelpful result. +- **Execution transcripts** reveal *why* things went wrong. If the agent ignored an instruction, the instruction may be ambiguous. If the agent spent time on unproductive steps, those instructions may need to be simplified or removed. + +The most effective way to turn these signals into skill improvements is to give all three — along with the current `SKILL.md` — to an LLM and ask it to propose changes. The LLM can synthesize patterns across failed assertions, reviewer complaints, and transcript behavior that would be tedious to connect manually. When prompting the LLM, include these guidelines: + +- **Generalize from feedback.** The skill will be used across many different prompts, not just the test cases. Fixes should address underlying issues broadly rather than adding narrow patches for specific examples. +- **Keep the skill lean.** Fewer, better instructions often outperform exhaustive rules. If transcripts show wasted work (unnecessary validation, unneeded intermediate outputs), remove those instructions. If pass rates plateau despite adding more rules, the skill may be over-constrained — try removing instructions and see if results hold or improve. +- **Explain the why.** Reasoning-based instructions ("Do X because Y tends to cause Z") work better than rigid directives ("ALWAYS do X, NEVER do Y"). Models follow instructions more reliably when they understand the purpose. +- **Bundle repeated work.** If every test run independently wrote a similar helper script (a chart builder, a data parser), that's a signal to bundle the script into the skill's `scripts/` directory. See [Using scripts](/skill-creation/using-scripts) for how to do this. + +### The loop + +1. Give the eval signals and current `SKILL.md` to an LLM and ask it to propose improvements. +2. Review and apply the changes. +3. Rerun all test cases in a new `iteration-/` directory. +4. Grade and aggregate the new results. +5. Review with a human. Repeat. + +Stop when you're satisfied with the results, feedback is consistently empty, or you're no longer seeing meaningful improvement between iterations. + + +The [`skill-creator`](https://github.com/anthropics/skills/tree/main/skills/skill-creator) Skill automates much of this workflow — running evals, grading assertions, aggregating benchmarks, and presenting results for human review. + diff --git a/.agents/skills/agent-package-skill-create/references/agentskills-optimizing-descriptions.mdx b/.agents/skills/agent-package-skill-create/references/agentskills-optimizing-descriptions.mdx new file mode 100644 index 0000000..8bb2a2f --- /dev/null +++ b/.agents/skills/agent-package-skill-create/references/agentskills-optimizing-descriptions.mdx @@ -0,0 +1,193 @@ +--- +title: "Optimizing skill descriptions" +sidebarTitle: "Optimizing descriptions" +description: "How to improve your skill's description so it triggers reliably on relevant prompts." +--- + +A skill only helps if it gets activated. The `description` field in your `SKILL.md` frontmatter is the primary mechanism agents use to decide whether to load a skill for a given task. An under-specified description means the skill won't trigger when it should; an over-broad description means it triggers when it shouldn't. + +This guide covers how to systematically test and improve your skill's description for triggering accuracy. + +## How skill triggering works + +Agents use [progressive disclosure](/specification#progressive-disclosure) to manage context. At startup, they load only the `name` and `description` of each available skill — just enough to decide when a skill might be relevant. When a user's task matches a description, the agent reads the full `SKILL.md` into context and follows its instructions. + +This means the description carries the entire burden of triggering. If the description doesn't convey when the skill is useful, the agent won't know to reach for it. + +One important nuance: agents typically only consult skills for tasks that require knowledge or capabilities beyond what they can handle alone. A simple, one-step request like "read this PDF" may not trigger a PDF skill even if the description matches perfectly, because the agent can handle it with basic tools. Tasks that involve specialized knowledge — an unfamiliar API, a domain-specific workflow, or an uncommon format — are where a well-written description can make the difference. + +## Writing effective descriptions + +Before testing, it helps to know what a good description looks like. A few principles: + +- **Use imperative phrasing.** Frame the description as an instruction to the agent: "Use this skill when..." rather than "This skill does..." The agent is deciding whether to act, so tell it when to act. +- **Focus on user intent, not implementation.** Describe what the user is trying to achieve, not the skill's internal mechanics. The agent matches against what the user asked for. +- **Err on the side of being pushy.** Explicitly list contexts where the skill applies, including cases where the user doesn't name the domain directly: "even if they don't explicitly mention 'CSV' or 'analysis.'" +- **Keep it concise.** A few sentences to a short paragraph is usually right — long enough to cover the skill's scope, short enough that it doesn't bloat the agent's context across many skills. The [specification](/specification#description-field) enforces a hard limit of 1024 characters. + +## Designing trigger eval queries + +To test triggering, you need a set of eval queries — realistic user prompts labeled with whether they should or shouldn't trigger your skill. + +```json eval_queries.json +[ + { "query": "I've got a spreadsheet in ~/data/q4_results.xlsx with revenue in col C and expenses in col D — can you add a profit margin column and highlight anything under 10%?", "should_trigger": true }, + { "query": "whats the quickest way to convert this json file to yaml", "should_trigger": false } +] +``` + +Aim for about 20 queries: 8-10 that should trigger and 8-10 that shouldn't. + +### Should-trigger queries + +These test whether the description captures the skill's scope. Vary them along several axes: + +- **Phrasing**: some formal, some casual, some with typos or abbreviations. +- **Explicitness**: some name the skill's domain directly ("analyze this CSV"), others describe the need without naming it ("my boss wants a chart from this data file"). +- **Detail**: mix terse prompts with context-heavy ones — a short "analyze my sales CSV and make a chart" alongside a longer message with file paths, column names, and backstory. +- **Complexity**: vary the number of steps and decision points. Include single-step tasks alongside multi-step workflows to test whether the agent can discern the skill is relevant when the task it addresses is buried in a larger chain. + +The most useful should-trigger queries are ones where the skill would help but the connection isn't obvious from the query alone. These are the cases where description wording makes the difference — if the query already asks for exactly what the skill does, any reasonable description would trigger. + +### Should-not-trigger queries + +The most valuable negative test cases are **near-misses** — queries that share keywords or concepts with your skill but actually need something different. These test whether the description is precise, not just broad. + +For a CSV analysis skill, weak negative examples would be: + +- `"Write a fibonacci function"` — obviously irrelevant, tests nothing. +- `"What's the weather today?"` — no keyword overlap, too easy. + +Strong negative examples: + +- `"I need to update the formulas in my Excel budget spreadsheet"` — shares "spreadsheet" and "data" concepts, but needs Excel editing, not CSV analysis. +- `"can you write a python script that reads a csv and uploads each row to our postgres database"` — involves CSV, but the task is database ETL, not analysis. + +### Tips for realism + +Real user prompts contain context that generic test queries lack. Include: + +- File paths (`~/Downloads/report_final_v2.xlsx`) +- Personal context (`"my manager asked me to..."`) +- Specific details (column names, company names, data values) +- Casual language, abbreviations, and occasional typos + +## Testing whether a description triggers + +The basic approach: run each query through your agent with the skill installed and observe whether the agent invokes it. Make sure the skill is registered and discoverable by your agent — how this works varies by client (e.g., a skills directory, a configuration file, or a CLI flag). + +Most agent clients provide some form of observability — execution logs, tool call histories, or verbose output — that lets you see which skills were consulted during a run. Check your client's documentation for details. The skill triggered if the agent loaded your skill's `SKILL.md`; it didn't trigger if the agent proceeded without consulting it. + +A query "passes" if: +- `should_trigger` is `true` and the skill was invoked, or +- `should_trigger` is `false` and the skill was not invoked. + +### Running multiple times + +Model behavior is nondeterministic — the same query might trigger the skill on one run but not the next. Run each query multiple times (3 is a reasonable starting point) and compute a **trigger rate**: the fraction of runs where the skill was invoked. + +A should-trigger query passes if its trigger rate is above a threshold (0.5 is a reasonable default). A should-not-trigger query passes if its trigger rate is below that threshold. + +With 20 queries at 3 runs each, that's 60 invocations. You'll want to script this. Here's the general structure — replace the `claude` invocation and detection logic in `check_triggered` with whatever your agent client provides: + +```bash +#!/bin/bash +QUERIES_FILE="${1:?Usage: $0 }" +SKILL_NAME="my-skill" +RUNS=3 + +# This example uses Claude Code's JSON output to check for Skill tool calls. +# Replace this function with detection logic for your agent client. +# Should return 0 (success) if the skill was invoked, 1 otherwise. +check_triggered() { + local query="$1" + claude -p "$query" --output-format json 2>/dev/null \ + | jq -e --arg skill "$SKILL_NAME" \ + 'any(.messages[].content[]; .type == "tool_use" and .name == "Skill" and .input.skill == $skill)' \ + > /dev/null 2>&1 +} + +count=$(jq length "$QUERIES_FILE") +for i in $(seq 0 $((count - 1))); do + query=$(jq -r ".[$i].query" "$QUERIES_FILE") + should_trigger=$(jq -r ".[$i].should_trigger" "$QUERIES_FILE") + triggers=0 + + for run in $(seq 1 $RUNS); do + check_triggered "$query" && triggers=$((triggers + 1)) + done + + jq -n \ + --arg query "$query" \ + --argjson should_trigger "$should_trigger" \ + --argjson triggers "$triggers" \ + --argjson runs "$RUNS" \ + '{query: $query, should_trigger: $should_trigger, triggers: $triggers, runs: $runs, trigger_rate: ($triggers / $runs)}' +done | jq -s '.' +``` + + +If your agent client supports it, you can stop a run early once the outcome is clear — the agent either consulted the skill or started working without it. This can significantly reduce the time and cost of running the full eval set. + + +## Avoiding overfitting with train/validation splits + +If you optimize the description against all your queries, you risk overfitting — crafting a description that works for these specific phrasings but fails on new ones. + +The solution is to split your query set: + +- **Train set (~60%)**: the queries you use to identify failures and guide improvements. +- **Validation set (~40%)**: queries you set aside and only use to check whether improvements generalize. + +Make sure both sets contain a proportional mix of should-trigger and should-not-trigger queries — don't accidentally put all the positives in one set. Shuffle randomly and keep the split fixed across iterations so you're comparing apples to apples. + +If you're using a script like the one [above](#running-multiple-times), you can split your queries into two files — `train_queries.json` and `validation_queries.json` — and run the script against each one separately. + +## The optimization loop + +1. **Evaluate** the current description on both *train and validation sets*. The train results guide your changes; the validation results tell you whether those changes are generalizing. +2. **Identify failures** in the *train set*: which should-trigger queries didn't trigger? Which should-not-trigger queries did? + - Only use train set failures to guide your changes — whether you're revising the description yourself or prompting an LLM, keep validation set results out of the process. +3. **Revise the description.** Focus on generalizing: + - If should-trigger queries are failing, the description may be too narrow. Broaden the scope or add context about when the skill is useful. + - If should-not-trigger queries are false-triggering, the description may be too broad. Add specificity about what the skill does *not* do, or clarify the boundary between this skill and adjacent capabilities. + - Avoid adding specific keywords from failed queries — that's overfitting. Instead, find the general category or concept those queries represent and address that. + - If you're stuck after several iterations, try a structurally different approach to the description rather than incremental tweaks. A different framing or sentence structure may break through where refinement can't. + - Check that the description stays under the 1024-character limit — descriptions tend to grow during optimization. +4. **Repeat** steps 1-3 until all *train set* queries pass or you stop seeing meaningful improvement. +5. **Select the best iteration** by its validation pass rate — the fraction of queries in the *validation set* that passed. Note that the best description may not be the last one you produced; an earlier iteration might have a higher validation pass rate than later ones that overfit to the train set. + +Five iterations is usually enough. If performance isn't improving, the issue may be with the queries (too easy, too hard, or poorly labeled) rather than the description. + + +The [`skill-creator`](https://github.com/anthropics/skills/tree/main/skills/skill-creator) Skill automates this loop end-to-end: it splits the eval set, evaluates trigger rates in parallel, proposes description improvements using Claude, and generates a live HTML report you can watch as it runs. + + +## Applying the result + +Once you've selected the best description: + +1. Update the `description` field in your `SKILL.md` frontmatter. +2. Verify the description is under the [1024-character limit](/specification#description-field). +3. Verify the description triggers as expected. Try a few prompts manually as a quick sanity check. For a more rigorous test, write 5-10 fresh queries (a mix of should-trigger and should-not-trigger) and run them through the eval script — since these queries were never part of the optimization process, they give you an honest check on whether the description generalizes. + +Before and after: + +```yaml +# Before +description: Process CSV files. + +# After +description: > + Analyze CSV and tabular data files — compute summary statistics, + add derived columns, generate charts, and clean messy data. Use this + skill when the user has a CSV, TSV, or Excel file and wants to + explore, transform, or visualize the data, even if they don't + explicitly mention "CSV" or "analysis." +``` + +The improved description is more specific about what the skill does (summary stats, derived columns, charts, cleaning) and broader about when it applies (CSV, TSV, Excel; even without explicit keywords). + +## Next steps + +Once your skill triggers reliably, you'll want to evaluate whether it produces good outputs. See [Evaluating skill output quality](/skill-creation/evaluating-skills) for how to set up test cases, grade results, and iterate. diff --git a/.agents/skills/agent-package-skill-create/references/agentskills-specification.mdx b/.agents/skills/agent-package-skill-create/references/agentskills-specification.mdx new file mode 100644 index 0000000..20cf9f6 --- /dev/null +++ b/.agents/skills/agent-package-skill-create/references/agentskills-specification.mdx @@ -0,0 +1,245 @@ +--- +title: "Specification" +description: "The complete format specification for Agent Skills." +--- + +## Directory structure + +A skill is a directory containing, at minimum, a `SKILL.md` file: + +``` +skill-name/ +├── SKILL.md # Required: metadata + instructions +├── scripts/ # Optional: executable code +├── references/ # Optional: documentation +├── assets/ # Optional: templates, resources +└── ... # Any additional files or directories +``` + +## `SKILL.md` format + +The `SKILL.md` file must contain YAML frontmatter followed by Markdown content. + +### Frontmatter + +| Field | Required | Constraints | +|-------|----------|-------------| +| `name` | Yes | Max 64 characters. Lowercase letters, numbers, and hyphens only. Must not start or end with a hyphen. | +| `description` | Yes | Max 1024 characters. Non-empty. Describes what the skill does and when to use it. | +| `license` | No | License name or reference to a bundled license file. | +| `compatibility` | No | Max 500 characters. Indicates environment requirements (intended product, system packages, network access, etc.). | +| `metadata` | No | Arbitrary key-value mapping for additional metadata. | +| `allowed-tools` | No | Space-separated string of pre-approved tools the skill may use. (Experimental) | + + +**Minimal example:** + +```markdown SKILL.md +--- +name: skill-name +description: A description of what this skill does and when to use it. +--- +``` + +**Example with optional fields:** + +```markdown SKILL.md +--- +name: pdf-processing +description: Extract PDF text, fill forms, merge files. Use when handling PDFs. +license: Apache-2.0 +metadata: + author: example-org + version: "1.0" +--- +``` + + +#### `name` field + +The required `name` field: +- Must be 1-64 characters +- May only contain unicode lowercase alphanumeric characters (`a-z`, `0-9`) and hyphens (`-`) +- Must not start or end with a hyphen (`-`) +- Must not contain consecutive hyphens (`--`) +- Must match the parent directory name + + +**Valid examples:** +```yaml +name: pdf-processing +``` +```yaml +name: data-analysis +``` +```yaml +name: code-review +``` + +**Invalid examples:** +```yaml +name: PDF-Processing # uppercase not allowed +``` +```yaml +name: -pdf # cannot start with hyphen +``` +```yaml +name: pdf--processing # consecutive hyphens not allowed +``` + + +#### `description` field + +The required `description` field: +- Must be 1-1024 characters +- Should describe both what the skill does and when to use it +- Should include specific keywords that help agents identify relevant tasks + + +**Good example:** +```yaml +description: Extracts text and tables from PDF files, fills PDF forms, and merges multiple PDFs. Use when working with PDF documents or when the user mentions PDFs, forms, or document extraction. +``` + +**Poor example:** +```yaml +description: Helps with PDFs. +``` + + +#### `license` field + +The optional `license` field: +- Specifies the license applied to the skill +- We recommend keeping it short (either the name of a license or the name of a bundled license file) + + +**Example:** +```yaml +license: Proprietary. LICENSE.txt has complete terms +``` + + +#### `compatibility` field + +The optional `compatibility` field: +- Must be 1-500 characters if provided +- Should only be included if your skill has specific environment requirements +- Can indicate intended product, required system packages, network access needs, etc. + + +**Examples:** +```yaml +compatibility: Designed for Claude Code (or similar products) +``` +```yaml +compatibility: Requires git, docker, jq, and access to the internet +``` +```yaml +compatibility: Requires Python 3.14+ and uv +``` + + + +Most skills do not need the `compatibility` field. + + +#### `metadata` field + +The optional `metadata` field: +- A map from string keys to string values +- Clients can use this to store additional properties not defined by the Agent Skills spec +- We recommend making your key names reasonably unique to avoid accidental conflicts + + +**Example:** +```yaml +metadata: + author: example-org + version: "1.0" +``` + + +#### `allowed-tools` field + +The optional `allowed-tools` field: +- A space-separated string of tools that are pre-approved to run +- Experimental. Support for this field may vary between agent implementations + + +**Example:** +```yaml +allowed-tools: Bash(git:*) Bash(jq:*) Read +``` + + +### Body content + +The Markdown body after the frontmatter contains the skill instructions. There are no format restrictions. Write whatever helps agents perform the task effectively. + +Recommended sections: +- Step-by-step instructions +- Examples of inputs and outputs +- Common edge cases + +Note that the agent will load this entire file once it's decided to activate a skill. Consider splitting longer `SKILL.md` content into referenced files. + +## Optional directories + +### `scripts/` + +Contains executable code that agents can run. Scripts should: +- Be self-contained or clearly document dependencies +- Include helpful error messages +- Handle edge cases gracefully + +Supported languages depend on the agent implementation. Common options include Python, Bash, and JavaScript. + +### `references/` + +Contains additional documentation that agents can read when needed: +- `REFERENCE.md` - Detailed technical reference +- `FORMS.md` - Form templates or structured data formats +- Domain-specific files (`finance.md`, `legal.md`, etc.) + +Keep individual [reference files](#file-references) focused. Agents load these on demand, so smaller files mean less use of context. + +### `assets/` + +Contains static resources: +- Templates (document templates, configuration templates) +- Images (diagrams, examples) +- Data files (lookup tables, schemas) + +## Progressive disclosure + +Agents load skills *progressively*, pulling in more detail only as a task calls for it. Skills should be structured to take advantage of this: + +1. **Metadata** (~100 tokens): The `name` and `description` fields are loaded at startup for all skills +2. **Instructions** (< 5000 tokens recommended): The full `SKILL.md` body is loaded when the skill is activated +3. **Resources** (as needed): Files (e.g. those in `scripts/`, `references/`, or `assets/`) are loaded only when required + +Keep your main `SKILL.md` under 500 lines. Move detailed reference material to separate files. + +## File references + +When referencing other files in your skill, use relative paths from the skill root: + +```markdown SKILL.md +See [the reference guide](references/REFERENCE.md) for details. + +Run the extraction script: +scripts/extract.py +``` + +Keep file references one level deep from `SKILL.md`. Avoid deeply nested reference chains. + +## Validation + +Use the [skills-ref](https://github.com/agentskills/agentskills/tree/main/skills-ref) reference library to validate your skills: + +```bash +skills-ref validate ./my-skill +``` + +This checks that your `SKILL.md` frontmatter is valid and follows all naming conventions. diff --git a/.agents/skills/agent-package-skill-create/references/agentskills-using-scripts.mdx b/.agents/skills/agent-package-skill-create/references/agentskills-using-scripts.mdx new file mode 100644 index 0000000..11ce443 --- /dev/null +++ b/.agents/skills/agent-package-skill-create/references/agentskills-using-scripts.mdx @@ -0,0 +1,298 @@ +--- +title: "Using scripts in skills" +sidebarTitle: "Using scripts" +description: "How to run commands and bundle executable scripts in your skills." +--- + +Skills can instruct agents to run shell commands and bundle reusable scripts in a `scripts/` directory. This guide covers one-off commands, self-contained scripts with their own dependencies, and how to design script interfaces for agentic use. + +## One-off commands + +When an existing package already does what you need, you can reference it directly in your `SKILL.md` instructions without a `scripts/` directory. Many ecosystems provide tools that auto-resolve dependencies at runtime. + + + + [uvx](https://docs.astral.sh/uv/guides/tools/) runs Python packages in isolated environments with aggressive caching. It ships with [uv](https://docs.astral.sh/uv/). + + ```bash + uvx ruff@0.8.0 check . + uvx black@24.10.0 . + ``` + + - Not bundled with Python — requires a separate install. + - Fast. Caches aggressively so repeat runs are near-instant. + + + [pipx](https://pipx.pypa.io/) runs Python packages in isolated environments. Available via OS package managers (`apt install pipx`, `brew install pipx`). + + ```bash + pipx run 'black==24.10.0' . + pipx run 'ruff==0.8.0' check . + ``` + + - Not bundled with Python — requires a separate install. + - A mature alternative to `uvx`. While `uvx` has become the standard recommendation, `pipx` remains a reliable option with broader OS package manager availability. + + + [npx](https://docs.npmjs.com/cli/commands/npx) runs npm packages, downloading them on demand. It ships with npm (which ships with Node.js). + + ```bash + npx eslint@9 --fix . + npx create-vite@6 my-app + ``` + + - Bundled with Node.js — no extra install needed. + - Downloads the package, runs it, and caches it for future use. + - Pin versions with `npx package@version` for reproducibility. + + + [bunx](https://bun.sh/docs/cli/bunx) is Bun's equivalent of `npx`. It ships with [Bun](https://bun.sh/). + + ```bash + bunx eslint@9 --fix . + bunx create-vite@6 my-app + ``` + + - Drop-in replacement for `npx` in Bun-based environments. + - Only appropriate when the user's environment has Bun rather than Node.js. + + + [deno run](https://docs.deno.com/runtime/reference/cli/run/) runs scripts directly from URLs or specifiers. It ships with [Deno](https://deno.com/). + + ```bash + deno run npm:create-vite@6 my-app + deno run --allow-read npm:eslint@9 -- --fix . + ``` + + - Permission flags (`--allow-read`, etc.) are required for filesystem/network access. + - Use `--` to separate Deno flags from the tool's own flags. + + + [go run](https://pkg.go.dev/cmd/go#hdr-Compile_and_run_Go_program) compiles and runs Go packages directly. It is built into the `go` command. + + ```bash + go run golang.org/x/tools/cmd/goimports@v0.28.0 . + go run github.com/golangci/golangci-lint/cmd/golangci-lint@v1.62.0 run + ``` + + - Built into Go — no extra tooling needed. + - Pin versions or use `@latest` to make the command explicit. + + + +**Tips for one-off commands in skills:** + +- **Pin versions** (e.g., `npx eslint@9.0.0`) so the command behaves the same over time. +- **State prerequisites** in your `SKILL.md` (e.g., "Requires Node.js 18+") rather than assuming the agent's environment has them. For runtime-level requirements, use the [`compatibility` frontmatter field](/specification#compatibility-field). +- **Move complex commands into scripts.** A one-off command works well when you're invoking a tool with a few flags. When a command grows complex enough that it's hard to get right on the first try, a tested script in `scripts/` is more reliable. + +## Referencing scripts from `SKILL.md` + +Use **relative paths from the skill directory root** to reference bundled files. The agent resolves these paths automatically — no absolute paths needed. + +List available scripts in your `SKILL.md` so the agent knows they exist: + +```markdown SKILL.md +## Available scripts + +- **`scripts/validate.sh`** — Validates configuration files +- **`scripts/process.py`** — Processes input data +``` + +Then instruct the agent to run them: + +````markdown SKILL.md +## Workflow + +1. Run the validation script: + ```bash + bash scripts/validate.sh "$INPUT_FILE" + ``` + +2. Process the results: + ```bash + python3 scripts/process.py --input results.json + ``` +```` + + +The same relative-path convention works in support files like `references/*.md` — script execution paths (in code blocks) are relative to the **skill directory root**, because the agent runs commands from there. + + +## Self-contained scripts + +When you need reusable logic, bundle a script in `scripts/` that declares its own dependencies inline. The agent can run the script with a single command — no separate manifest file or install step required. + +Several languages support inline dependency declarations: + + + + [PEP 723](https://peps.python.org/pep-0723/) defines a standard format for inline script metadata. Declare dependencies in a TOML block inside `# ///` markers: + + ```python scripts/extract.py + # /// script + # dependencies = [ + # "beautifulsoup4", + # ] + # /// + + from bs4 import BeautifulSoup + + html = '

Welcome

This is a test.

' + print(BeautifulSoup(html, "html.parser").select_one("p.info").get_text()) + ``` + + Run with [uv](https://docs.astral.sh/uv/) (recommended): + + ```bash + uv run scripts/extract.py + ``` + + `uv run` creates an isolated environment, installs the declared dependencies, and runs the script. [pipx](https://pipx.pypa.io/) (`pipx run scripts/extract.py`) also supports PEP 723. + + - Pin versions with [PEP 508](https://peps.python.org/pep-0508/) specifiers: `"beautifulsoup4>=4.12,<5"`. + - Use `requires-python` to constrain the Python version. + - Use `uv lock --script` to create a lockfile for full reproducibility. +
+ + Deno's `npm:` and `jsr:` import specifiers make every script self-contained by default: + + ```typescript scripts/extract.ts + #!/usr/bin/env -S deno run + + import * as cheerio from "npm:cheerio@1.0.0"; + + const html = `

Welcome

This is a test.

`; + const $ = cheerio.load(html); + console.log($("p.info").text()); + ``` + + ```bash + deno run scripts/extract.ts + ``` + + - Use `npm:` for npm packages, `jsr:` for Deno-native packages. + - Version specifiers follow semver: `@1.0.0` (exact), `@^1.0.0` (compatible). + - Dependencies are cached globally. Use `--reload` to force re-fetch. + - Packages with native addons (node-gyp) may not work — packages that ship pre-built binaries work best. +
+ + Bun auto-installs missing packages at runtime when no `node_modules` directory is found. Pin versions directly in the import path: + + ```typescript scripts/extract.ts + #!/usr/bin/env bun + + import * as cheerio from "cheerio@1.0.0"; + + const html = `

Welcome

This is a test.

`; + const $ = cheerio.load(html); + console.log($("p.info").text()); + ``` + + ```bash + bun run scripts/extract.ts + ``` + + - No `package.json` or `node_modules` needed. TypeScript works natively. + - Packages are cached globally. First run downloads; subsequent runs are near-instant. + - If a `node_modules` directory exists anywhere up the directory tree, auto-install is disabled and Bun falls back to standard Node.js resolution. +
+ + Bundler ships with Ruby since 2.6. Use `bundler/inline` to declare gems directly in the script: + + ```ruby scripts/extract.rb + require 'bundler/inline' + + gemfile do + source 'https://rubygems.org' + gem 'nokogiri' + end + + html = '

Welcome

This is a test.

' + doc = Nokogiri::HTML(html) + puts doc.at_css('p.info').text + ``` + + ```bash + ruby scripts/extract.rb + ``` + + - Pin versions explicitly (`gem 'nokogiri', '~> 1.16'`) — there is no lockfile. + - An existing `Gemfile` or `BUNDLE_GEMFILE` env var in the working directory can interfere. +
+
+ +## Designing scripts for agentic use + +When an agent runs your script, it reads stdout and stderr to decide what to do next. A few design choices make scripts dramatically easier for agents to use. + +### Avoid interactive prompts + +This is a hard requirement of the agent execution environment. Agents operate in non-interactive shells — they cannot respond to TTY prompts, password dialogs, or confirmation menus. A script that blocks on interactive input will hang indefinitely. + +Accept all input via command-line flags, environment variables, or stdin: + +``` +# Bad: hangs waiting for input +$ python scripts/deploy.py +Target environment: _ + +# Good: clear error with guidance +$ python scripts/deploy.py +Error: --env is required. Options: development, staging, production. +Usage: python scripts/deploy.py --env staging --tag v1.2.3 +``` + +### Document usage with `--help` + +`--help` output is the primary way an agent learns your script's interface. Include a brief description, available flags, and usage examples: + +``` +Usage: scripts/process.py [OPTIONS] INPUT_FILE + +Process input data and produce a summary report. + +Options: + --format FORMAT Output format: json, csv, table (default: json) + --output FILE Write output to FILE instead of stdout + --verbose Print progress to stderr + +Examples: + scripts/process.py data.csv + scripts/process.py --format csv --output report.csv data.csv +``` + +Keep it concise — the output enters the agent's context window alongside everything else it's working with. + +### Write helpful error messages + +When an agent gets an error, the message directly shapes its next attempt. An opaque "Error: invalid input" wastes a turn. Instead, say what went wrong, what was expected, and what to try: + +``` +Error: --format must be one of: json, csv, table. + Received: "xml" +``` + +### Use structured output + +Prefer structured formats — JSON, CSV, TSV — over free-form text. Structured formats can be consumed by both the agent and standard tools (`jq`, `cut`, `awk`), making your script composable in pipelines. + +``` +# Whitespace-aligned — hard to parse programmatically +NAME STATUS CREATED +my-service running 2025-01-15 + +# Delimited — unambiguous field boundaries +{"name": "my-service", "status": "running", "created": "2025-01-15"} +``` + +**Separate data from diagnostics:** send structured data to stdout and progress messages, warnings, and other diagnostics to stderr. This lets the agent capture clean, parseable output while still having access to diagnostic information when needed. + +### Further considerations + +- **Idempotency.** Agents may retry commands. "Create if not exists" is safer than "create and fail on duplicate." +- **Input constraints.** Reject ambiguous input with a clear error rather than guessing. Use enums and closed sets where possible. +- **Dry-run support.** For destructive or stateful operations, a `--dry-run` flag lets the agent preview what will happen. +- **Meaningful exit codes.** Use distinct exit codes for different failure types (not found, invalid arguments, auth failure) and document them in your `--help` output so the agent knows what each code means. +- **Safe defaults.** Consider whether destructive operations should require explicit confirmation flags (`--confirm`, `--force`) or other safeguards appropriate to the risk level. +- **Predictable output size.** Many agent harnesses automatically truncate tool output beyond a threshold (e.g., 10-30K characters), potentially losing critical information. If your script might produce large output, default to a summary or a reasonable limit, and support flags like `--offset` so the agent can request more information when needed. Alternatively, if output is large and not amenable to pagination, require agents to pass an `--output` flag that specifies either an output file or `-` to explicitly opt in to stdout. diff --git a/.agents/skills/agent-package-skill-create/references/asset-packaging.md b/.agents/skills/agent-package-skill-create/references/asset-packaging.md new file mode 100644 index 0000000..4ad7c5a --- /dev/null +++ b/.agents/skills/agent-package-skill-create/references/asset-packaging.md @@ -0,0 +1,40 @@ +# 资产大小与交付 + +## 检查实际交付内容 + +Agent ZIP 包总大小不得超过 500 MB(500,000,000 字节),按最终 ZIP 文件的实际字节数校验。 +大小门禁集中在 Package 归档及上传/下载流程;候选 Git 快照不设单文件、展开总量或文件数量门禁。 +这是 Package 实现限制,不是 Agent Skills 开放规范;Host 仍会独立校验,以当前环境契约为准。 + +检查整个 Package 生成的 ZIP,包含递归展开的子模块,不能只统计当前 Skill。 +不要把原始资产大小或展开后的总量当成 ZIP 大小,也不要假设 `.gitignore` 能排除已跟踪文件。 + +## 必需大资产的处理 + +用户要求随包携带的资产必须保持完整,不擅自删除、裁剪或改成远程下载。 +单个 `assets/` 文件超过 16 MiB(16,777,216 字节)时,发出非阻断警告,列出路径和实际大小, +并推荐评估无损压缩后随包携带、使用时解压。恰好 16 MiB 不警告;资产仍可原样携带, +不因这条建议强制压缩或拒绝校验。 +如果最终 ZIP 接近或超过额度且消费流程允许使用前解压,可评估无损压缩; +ZIP 本身已有压缩,嵌套 gzip 不保证进一步缩小包,必须比较最终 ZIP 大小;小文件和无法有效压缩的资产不必强行采用此方案。 + +- 可用 Python 标准库 gzip 生成 `.gz`;固定 `mtime=0`,不嵌入机器路径等可变元数据。 +- 验证解压后的字节与原文件完全一致,记录或保留原始文件的 SHA-256。 + 已有预期哈希和业务结果不得为了通过测试而改写。 +- 更新消费脚本,将文件解压到 Wayflow 注入的 Runtime 下的可写目录,再调用原有处理流程。 + 不写入安装后的 Skill 目录,不使用开发机绝对路径;Runtime 缺失时明确报错。 +- 更新 `SKILL.md`、references、脚本及相关文件清单中的引用,区分包内压缩资产和运行时解压路径。 + Skill 内资产仍不重复声明为顶层 `resources[]`。 +- 使用压缩资产替换方案时,确认交付快照中不再重复包含原文件,并检查最终 ZIP 大小和 + 运行时解压所需空间。运行使用该资产的最小代表性流程,验证既有业务结果。 + +如果压缩后仍超限,或使用方必须直接读取包内未压缩文件,报告实际大小及消费约束, +交由 Package/平台工作流处理资源额度;不要在创建 Skill 的任务中擅自修改平台限制。 + +## Git submodule 交接 + +候选包读取父仓库 gitlink 锁定的子模块 commit,不读取子模块当前分支的未提交修改。 +按当前任务授权和版本工作流提交子模块修改,再更新父仓库 gitlink;不自动推送远端。 +仅删除本地文件或只在子模块提交而未更新父仓库引用,都不能修复旧锁定快照。 +交接时报告压缩前后字节数、无损与业务验证结果、锁定提交更新情况,以及尚待执行的 +Agent Workforce 完整 Bundle 校验。 diff --git a/README.md b/README.md index 74a1d42..de2e3db 100644 --- a/README.md +++ b/README.md @@ -38,7 +38,7 @@ git submodule update --init --recursive 新建 Agent 时,插件会从本仓库生成独立开发仓库,更新包 ID、版本号、角色和标题。 模板内的 instructions、skills、assets 和其他包内容会保留。 -开发辅助技能由插件安装,无需放入此公共模板。 +开发辅助技能维护在 `.agents/skills/`,由插件安装到 Agent 开发工作区。 添加资源、技能或运行脚本后,请同步更新 `agent-package.json` 中的声明。 如使用子模块,建议使用明确的 HTTPS URL;生成的仓库位于 `wagent` 组织, @@ -49,7 +49,7 @@ git submodule update --init --recursive `skills/agent-runtime-guide` 是通用运行指南,随模板声明并安装。 `role/ceo`、`role/pm` 和 `role/shared/skills/wayflow-hire-agent` 是可选角色素材, 由 Workforce 在初始化对应角色时安装;通用 Agent 不安装 Hire 技能。 -插件的包编写辅助、版本编排和演进分析技能由插件仓库维护。 +包编写辅助技能由本仓库 `.agents/skills/` 维护;版本编排和演进分析技能仍由插件仓库维护。 本仓库由插件以 Git submodule 引用。修改模板后先提交、推送本仓库, 再更新插件仓库中的 submodule 指针并构建,插件包携带锁定的模板文件。