Agent 評估與可觀測性
Agent 評估與可觀測性
以資料集、評分器與執行軌跡量測 Agent 品質。
magnus919 · 社群來源 · AI 與 Agents
來源狀態
可用來源
原始名稱: agent-evals-and-observability
原始作者: magnus919
這是第三方 Community Skill;2lus 收錄不代表已完整安全審計或保證安全。
資源類型
Skill
與類似 Skill 有什麼不同?
- Agent 使用體驗稽核
Agent 評估量測任務品質與軌跡;Agent Experience 檢視使用者操作感受。
來源描述的能力
這些是來源描述的可能操作,不代表 2lus 已授予權限或已測試。
尚未列出能力,請審查原始來源。
這個 Skill 是什麼?
建立可重複的離線評估與線上觀測回饋。
可以做什麼?何時適合使用?
- 比較 prompt 或模型變更
- 分析工具呼叫軌跡
- 設計發行品質門檻
如何使用
- 收集代表性案例與標準答案或判斷規則
- 定義評分器與人工複核方式
- 檢查分數分布、失敗軌跡與 production traces
使用前你需要準備
- 任務案例與成功定義
- 可去識別化的執行記錄
你可以替換:
請提供版本、來源資料、目標與限制。
環境與相依需求
- 任務案例與成功定義
- 可去識別化的執行記錄
使用範例與 Prompt
以下是 2lus 撰寫的示範需求;請替換為你有權處理的檔案與專案,不代表已執行或保證結果。
入門
為客服 Agent 建立 10 題小型評估資料集與評分標準。
實務
設計工具選擇與答案正確性的分開評分器。
進階
比較兩版 Agent 的任務結果、軌跡與線上錯誤率,提出發行門檻。
實用提醒
- 同時檢查最差案例與平均分;保留資料集版本。
限制與注意事項
- 單一分數無法代表安全性或真實使用者體驗。
安全注意事項
- 去識別化 trace,避免把機密 prompt 或工具輸出寫入公開資料集。
使用與設定
此條目不提供已確認的通用安裝指令;請依官方文件與 Agent 版本操作。
支援平台
未確認特定 Agent 相容性
來源與授權
來源查核日期(非安全認證): 2026-10-01
MIT
原始來源 ↗ 官方文件 ↗相關 Skills
使用第三方 Skill 前,請先檢查來源、權限與執行內容。安裝指令只供查看與複製,不會由 2lus 執行。
2lus AI Skills Library 提供 Skill 的整理與使用導覽。第三方 Skill 的內容、授權與可用性以原始來源為準。使用或安裝前,請自行確認其權限與執行內容。
Agent Evals & Observability
Measure agent quality with datasets, graders, and execution trajectories.
magnus919 · Community source · AI & Agents
Source status
Active source
Original name: agent-evals-and-observability
Original author: magnus919
This is a third-party community skill. Inclusion by 2lus is not a complete security audit or safety guarantee.
Resource type
Skill
How is this different from similar skills?
- Agent Experience
Agent Evals measure task quality and trajectories; Agent Experience reviews the user-facing workflow.
Documented capabilities
These are operations described upstream, not permissions granted or tested by 2lus.
Capabilities not declared here; review the original source.
What is this skill?
Build repeatable offline evaluation and online observation loops.
Use cases and when to use it
- Compare prompt or model changes
- Analyze tool-call trajectories
- Define release quality gates
How to use it
- Gather representative cases and expected outcomes or judgment rules
- Define graders and human review
- Inspect distributions, failure trajectories, and production traces
What you need
- Task cases and success criteria
- Redacted execution traces
You can replace:
Provide versions, source material, goals, and constraints.
Environment and dependencies
- Task cases and success criteria
- Redacted execution traces
Usage and prompt examples
These example requests were written by 2lus. Substitute files and projects you may use; examples are not executed results or guarantees.
Beginner
Create ten representative evaluation cases and scoring criteria for a support agent.
Practical
Design separate graders for tool selection and answer correctness.
Advanced
Compare two agent versions by outcomes, trajectories, and production error rates; propose a release gate.
Tips
- Inspect worst cases as well as averages; version the dataset.
Limitations
- A single score does not establish safety or real user experience.
Security notes
- Redact traces to keep secrets and sensitive prompts out of public datasets.
Usage and setup
No verified universal installation command is provided for this entry. Follow the official documentation for your agent version.
Supported agents
Specific agent compatibility unknown
Sources and license
Source check date (not a safety certification): 2026-10-01
MIT
Original source ↗ Documentation ↗Related skills
Before using a third-party skill, review its source, permissions and executable content. Commands are for viewing and copying only; 2lus does not execute them.
2lus AI Skills Library provides curated educational guides. Third-party content, licenses and availability are governed by their original sources. Review permissions and executable content before use or installation.