‹ 首页

ladder-quality-order

@yogsoth-ai · 收录于 5 天前 · 上游提交 2 周前

Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.

适合你,如果需要按标准对研究设计进行质量排序。

/ 通过 npx 安装 校验哈希
npx oh-my-skill add yogsoth-ai/de-anthropocentric-research-engine/ladder-quality-order
/ 通过 bash 安装
curl -fsSL https://oh-my-skill.com/install.sh | bash -s -- yogsoth-ai/de-anthropocentric-research-engine/ladder-quality-order
/ 已经装过?验证本机副本,不用重装
npx oh-my-skill verify yogsoth-ai/de-anthropocentric-research-engine/ladder-quality-order
安装目标可用 --agent / --scope 或 --to 明确指定;省略时只会在唯一已存在的 agent 目录上自动选择,零命中或多命中会停止并提示。content_hash 缺失或不一致均拒装。
388GitHub stars
~666最小装载
~666含声明引用
~815文本包总量
索引托管

怎么用

商店整理自技能原文 · 版本 59ace64 · 表述以原文为准
它做什么

安装后,Claude 会要求你对 6 个研究设计样本(每个含研究图和结果)进行成对质量比较,基于 D1–D5 标准选出更好的那个,最终得到质量排序报告。

什么时候触发

当你提供 6 个研究设计样本并请求按质量排序时触发。

装好后可以这样说
Claude 将输出赢得比较的样本编号和理由。
Claude 会启动成对比较流程,逐一询问你。
技能原文 SKILL.md作者撰写 · Apache-2.0 · 59ace64

ladder-quality-order (loss-2)

You rank ONE topic's 6 research-design samples (each a research_graph + research_result pair) by quality. The samples arrive SHUFFLED and anonymous — you see 6 positions (0–5), never their true rung id or config. You judge only on the D1–D5 standard:

  • D1 meaningfulness — is the research question real and worth asking?
  • D2 skill-research value — does the design advance skill/methodology research?
  • D3 use-to-DARE — is it usable by the DARE engine?
  • D4 respects the 4-layer architecture (campaign → strategy → tactic → sop)?
  • D5 prerequisites — are the stated prerequisites sound and met?

Judge only on the D1–D5 standard above; never on academic-publication criteria of any kind. You never see any quality-check list.

Pairwise mechanism

You will be asked to compare two positions at a time. For each pair (i, j) decide the winner (the higher-quality position) and give a one-line reason grounded in D1–D5. Do not assign absolute scores — only pick a winner per pair. The graph is structure-aware context; read it holistically, do not run any checklist over it.

The harness enumerates all 15 pairs (i<j over 6 positions), Copeland-aggregates your winners into an induced order, un-shuffles to true ids, and computes Kendall τ against the intended order id0 > id1 > … > id5 (id0 = highest quality). You only emit {winner, reason} per pair.

Endpoint separation

You will also be asked, K independent times, to compare the two extreme samples (the harness picks them and presents them as just two options, A and B). Return {"winner": "A" | "B"} — exactly the label of the higher-quality one. Judge each call independently and honestly; do not try to be consistent with a previous call you don't remember. (This is a two-way A/B label, distinct from the position integers used in the pairwise rank above.)

Confound flat-check (when present)

If the topic carries a same-substance / different-framing triplet, rank it first. The order must NOT change with framing alone (buzzword vs neutral wording is not a quality difference under D1–D5). If your order tracks framing, say so in the reason — the harness will treat this topic's ladder as untrustworthy.

What the harness writes (you do not compute these)

The harness assembles loss2.json: tau, monotonicity_pass (τ ≥ the τ line AND no adjacent endpoint inversion), endpoint_separation_pass (endpoint majority), rigor_floor_flag (endpoints a near-tie — a possible genuine quality floor, NOT a tuning bug), and pairwise_log (your winners + reasons, un-shuffled to true ids). Your only job: honest per-pair winners and D1–D5 reasons.

按 Apache-2.0 许可原样转载,未经改动 · 在 GitHub 查看 →

评论

登录即可评论;带「已验证安装」的,是发布者名下有本店的安装或持有记录。