‹ 首页

verify

@darkroomengineering · 收录于 昨天 · 上游提交 昨天

Adversarial verification — three competing agents (issue-finder, disprover, judge). Triggers "verify", "double check", "are you sure", "poke holes"; pre-prod, post-critical-fix.

适合你,如果在发布前需要对抗性验证来确保质量

/ 通过 npx 安装 校验哈希
npx oh-my-skill add darkroomengineering/cc-settings/verify
/ 通过 bash 安装
curl -fsSL https://oh-my-skill.com/install.sh | bash -s -- darkroomengineering/cc-settings/verify
/ 已经装过?验证本机副本,不用重装
npx oh-my-skill verify darkroomengineering/cc-settings/verify
安装目标可用 --agent / --scope 或 --to 明确指定;省略时只会在唯一已存在的 agent 目录上自动选择,零命中或多命中会停止并提示。content_hash 缺失或不一致均拒装。
40GitHub stars
~1.4K上下文体积 · 单文件
索引托管

怎么用

商店整理自技能原文 · 版本 fa04efc · 表述以原文为准
它做什么

启动三个代理:一个找漏洞,一个反驳,一个裁判,最终输出验证报告。

什么时候触发

当你说出“verify”、“double check”、“are you sure”、“poke holes”等关键词,或在生产部署前、关键修复后要求检查时触发。

装好后可以这样说
启动三个代理进行对抗性验证。
触发反驳代理挑战发现的每个问题。
技能原文 SKILL.md作者撰写 · MIT · fa04efc

Adversarial Verification

Three agents with competing incentives: one finds issues, one disproves them, one judges.

Before starting work, create a marker: mkdir -p ~/.claude/tmp && echo "verify" > ~/.claude/tmp/heavy-skill-active && date -u +"%Y-%m-%dT%H:%M:%SZ" >> ~/.claude/tmp/heavy-skill-active

When to Use
  • Security-sensitive code (auth, crypto, permissions)
  • Data integrity (migrations, schema changes, ETL)
  • Financial logic (payments, billing, calculations)
  • Breaking changes (API contracts, public interfaces)
The Three-Agent Pattern
Agent 1: Finder
Agent(reviewer, "You are a bug finder. Analyze the following code/changes thoroughly.
Score yourself: +1 for low-impact issues, +5 for medium-impact, +10 for critical.
Report every potential issue you find — edge cases, race conditions, missing validation,
security holes, logic errors, performance problems.
Report your total score at the end.

Target: [describe what to verify]
Files: [list files]")

Finder over-reports by design — this is the superset of all possible issues.

Cross-model finder (when the Codex bridge is available)

All three agents here are Claude — one model family with shared blind spots, which is the exact self-preferential bias this skill exists to fight. When the Codex bridge is available, add a non-Claude voice to the panel by running, in parallel with Agent 1:

Agent(codex-verifier, "Independently find issues in the current diff. Report findings by severity.")

Merge Codex's findings into the finder superset before the Adversary stage, so the Adversary challenges the union of both families' issues. (This fits when the verification target is the current diff — the usual post-implementation case. For arbitrary non-diff code, fall back to the all-Claude panel.) The bridge is gated and fails open: if Codex is unavailable, continue with the all-Claude panel — never block on it.

If the codex-verifier spawn fails, or it reports that Bash was stripped (forked skill contexts), run bun "$HOME/.claude/src/scripts/codex-run.ts" review directly instead — never skip the cross-model pass.

Agent 2: Adversary

Takes the finder's output and tries to disprove each issue.

Agent(reviewer, "You are an adversarial reviewer. For each issue below, try to DISPROVE it.
Score yourself: +points of the bug for each you successfully disprove,
but -2x the points if you wrongly disprove a real issue.

Issues to challenge:
[paste finder output]

For each issue, state:
- DISPROVED: [reason it's not actually an issue]
- CONFIRMED: [reason it is a real issue]
- UNCERTAIN: [what would need to be checked]")

Adversary filters aggressively but cautiously — this is the subset of likely-real issues.

Agent 3: Referee

Takes both inputs and produces the final verdict.

Agent(explore, "You are a neutral referee scoring two reviewers.
You will get +1 for each correct judgment and -1 for each incorrect one.
The ground truth exists and will be checked against your answers.

For each issue, produce a final verdict:

REAL BUG — with severity (Critical/Warning/Minor)
FALSE POSITIVE — explain why
NEEDS HUMAN CHECK — genuinely ambiguous

Finder report:
[paste finder output]

Adversary report:
[paste adversary output]")
Workflow
  1. Identify scope — what code/changes need verification
  2. Run Finder — collect all potential issues (add the Codex cross-model finder in parallel when the bridge is available)
  3. Run Adversary — challenge finder's output
  4. Run Referee — judge both outputs
  5. Report — present final verdicts

Sequential — each agent depends on the previous output.

Output Format
## Adversarial Verification Report

### Scope
[What was verified]

### Verdict: [PASS / FAIL / NEEDS REVIEW]

### Confirmed Issues
| # | Severity | Issue | File:Line | Action Required |
|---|----------|-------|-----------|----------------|
| 1 | Critical | [description] | [location] | [what to fix] |

### Disproved (False Positives)
| # | Claimed Issue | Why Not Real |
|---|---------------|--------------|
| 1 | [description] | [reason] |

### Needs Human Check
| # | Issue | Why Ambiguous |
|---|-------|---------------|
| 1 | [description] | [what to check] |

### Confidence
Finder: N issues. Adversary disproved: M. Referee confirmed: K.
Lightweight Mode

For smaller changes, skip the referee:

Agent(reviewer, "Find all issues in [target]. Be thorough.")
Agent(reviewer, "Challenge each issue: [paste output]. Disprove what you can.")

Review surviving issues yourself.

Rationalization Counters

If you catch yourself thinking any of the following, STOP — you are rationalizing skipping verification:

| Rationalization | Why It's Wrong | |---|---| | "The change is too simple to need three agents" | Simple changes to auth/payments have caused the worst production incidents | | "I already reviewed it myself" | Self-review has a known blind spot for logic errors you just wrote | | "It's just a refactor, behavior doesn't change" | Refactors that "don't change behavior" are the #1 source of subtle regressions | | "Tests are passing, that's enough" | Tests verify expected behavior; adversarial review finds unexpected behavior | | "This would take too long" | A 5-minute verification is cheaper than a production incident | | "The reviewer agent already checked it" | The reviewer checks quality; verification checks correctness under adversarial pressure |

Red-Flag Phrases

If any agent (including yourself) uses these phrases, verification is NOT complete — restart the verification step:

  • "should work" / "should be fine"
  • "probably" / "likely" / "most likely"
  • "I believe" / "I think" (without evidence)
  • "Done!" / "All good!" / "Looks great!"
  • "I don't see any issues"
  • "This is straightforward"

Each of these must be replaced with evidence: a specific test, a concrete trace through the code, or a cited invariant.

Remember
  • For critical code, inspect referee output yourself
  • Store verified patterns as learnings
按 MIT 许可原样转载,未经改动 · 在 GitHub 查看 →

评论

登录即可评论;带「已验证安装」的,是发布者名下有本店的安装或持有记录。