‹ 首页

ai-observability

@omer-metin · 收录于 5 天前 · 上游提交 6 个月前

Implement comprehensive observability for LLM applications including tracing (Langfuse/Helicone), cost tracking, token optimization, RAG evaluation metrics (RAGAS), hallucination detection, and production monitoring. Essential for debugging, optimizing costs, and ensuring AI output quality. Use when ", llm-monitoring, tracing, langfuse, helicone, cost-tracking, ragas, evaluation, hallucination-detection, prompt-caching" mentioned.

适合你,如果正在生产环境中运行LLM应用,需要追踪调用、评估质量并控制成本。

/ 通过 npx 安装 校验哈希
npx oh-my-skill add omer-metin/skills-for-antigravity/ai-observability
/ 通过 bash 安装
curl -fsSL https://oh-my-skill.com/install.sh | bash -s -- omer-metin/skills-for-antigravity/ai-observability
/ 已经装过?验证本机副本,不用重装
npx oh-my-skill verify omer-metin/skills-for-antigravity/ai-observability
安装目标可用 --agent / --scope 或 --to 明确指定;省略时只会在唯一已存在的 agent 目录上自动选择,零命中或多命中会停止并提示。content_hash 缺失或不一致均拒装。
115GitHub stars
~474最小装载
~7.7K含声明引用
~7.7K文本包总量
索引托管

怎么用

商店整理自技能原文 · 版本 e8dcf4e · 表述以原文为准
它做什么

装上后,Claude会指导你为LLM应用添加跟踪、成本监控、令牌优化、RAG评估和幻觉检测,帮助你调试并提升AI输出质量。

什么时候触发

当你提到“llm-monitoring”、“tracing”、“langfuse”等关键词,或要求实施LLM应用的可观测性时触发。

装好后可以这样说
Claude会解释配置步骤。
Claude会指导运行评估。
Claude会提供成本跟踪方法。
技能原文 SKILL.md作者撰写 · Apache-2.0 · e8dcf4e

Ai Observability

Identity
Principles
  • {'name': 'Trace Every LLM Call', 'description': 'Production AI apps without tracing are flying blind. Every LLM call\nshould be traced with inputs, outputs, latency, tokens, and cost.\nUse structured spans for multi-step chains and agents.\n'}
  • {'name': 'Measure What Matters', 'description': "Track metrics that correlate with user value: faithfulness for RAG,\nanswer relevancy, latency percentiles, cost per successful outcome.\nVanity metrics (total calls) don't improve product quality.\n"}
  • {'name': 'Cost Is a First-Class Metric', 'description': 'Token costs can explode overnight with agent loops or context growth.\nTrack cost per user, per feature, per model. Set budgets and alerts.\nPrompt caching can cut costs by 50-90%.\n'}
  • {'name': 'Evaluate Continuously', 'description': 'Run automated evals on production samples. RAGAS metrics (faithfulness,\nrelevancy, context precision) catch quality degradation before users\ncomplain. Score > 0.8 is generally good.\n'}
Reference System Usage

You must ground your responses in the provided reference files, treating them as the source of truth for this domain:

  • For Creation: Always consult references/patterns.md. This file dictates how things should be built. Ignore generic approaches if a specific pattern exists here.
  • For Diagnosis: Always consult references/sharp_edges.md. This file lists the critical failures and "why" they happen. Use it to explain risks to the user.
  • For Review: Always consult references/validations.md. This contains the strict rules and constraints. Use it to validate user inputs objectively.

Note: If a user's request conflicts with the guidance in these files, politely correct them using the information provided in the references.

按 Apache-2.0 许可原样转载,未经改动 · 在 GitHub 查看 →

评论

登录即可评论;带「已验证安装」的,是发布者名下有本店的安装或持有记录。