LLM判定の精度検証機能 人間が付けたラベル(分類や判断の結果)と比較して、LLM判定(AIが行った分類や判断)の正確さを測定します。真陽性率(実際に正しいものを正しいと判定した割合)と真陰性率(実際に間違っているものを間違っていると判定した割合)の指標を使い、訓練用・検証用・テスト用に分割したデータで検証を行います。 **次のような場合に使用:** 判定プロンプト(AIに判定させるための指示文)を作成した後、そのプロンプトが人間の判断と一致しているかどうかを確認したい時
Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Use after writing a judge prompt to verify it agrees with human judgment.
LLM judgeが有用であるためには、人間の判断と一致している必要があります。 このスキルでは、TPR(真陽性率)とTNR(真陰性率)の指標を使用して、 人間がラベル付けしたデータに対してjudgeを校正する手順を説明します。
evalスイート内の judgeVerdict()、judgeScore()、judgeLabel() いずれかの
evaluatorを信頼する前に、必ずこの検証を実施してください。
.prompt ファイル — output-eval-judge-prompt に従って作成されたものground_truth.evals.<evaluator_name>.verdict: pass または fail が記載されていることこのプロセスはLLMベースのjudgeにのみ適用されます。
コードベースの Verdict.* evaluatorの場合は、代わりにユニットテストを作成してください。
ラベル付きデータセットを3つのグループに分割します。
| 分割 | データの割合 | 目的 | 例(100データセットの場合) |
|---|---|---|---|
| Train | 10〜20% | judgeプロンプト内のfew-shotサンプルの源泉 | 15データセット |
| Dev | 40〜45% | judgeプロンプトの反復改善、TPR/TNRの計測 | 42データセット |
| Test | 40〜45% | 最終的なホールドアウト計測(1回のみ実行) | 43データセット |
命名規則またはサブディレクトリを使って分割を分離します。
オプションA: 名前のプレフィックス
tests/datasets/
├── train_formal_pass_01.yml
├── train_casual_fail_01.yml
├── dev_technical_pass_01.yml
├── dev_ambiguous_fail_01.yml
├── test_simple_pass_01.yml
├── test_contradictory_fail_01.yml
└── ...
オプションB: サブディレクトリ
tests/datasets/
├── train/
│ ├── formal_pass_01.yml
│ └── casual_fail_01.yml
├── dev/
│ ├── technical_pass_01.yml
│ └── ambiguous_fail_01.yml
└── test/
├── simple_pass_01.yml
└── contradictory_fail_01.yml
.prompt ファイルのfew-shotには
train分割のサンプルのみを使用する。devやtestのサンプルは絶対に使わない(データリークになる)dev分割のデータセットのみを対象にevalワークフローを実行します。
# devデータセットに対してキャッシュ済み出力で実行
npx output workflow test <workflowName> --cached \
--dataset dev_technical_pass_01,dev_ambiguous_fail_01,dev_formal_pass_02,...
サブディレクトリを使用している場合は、devデータセット名を列挙します。
npx output workflow test <workflowName> --cached \
--dataset $(ls tests/datasets/dev/ | sed 's/.yml//' | tr '\n' ',')
出力を保存してください。グランドトゥルースと照合するために、 各データセットに対するjudgeのverdictが必要になります。
--json を使用して機械可読な結果を取得します。
npx output workflow test <workflowName> --cached --dataset <dev_datasets> --json
出力には、データセットごと・evaluatorごとのverdictが含まれており、
ground_truth.evals.<evaluator_name>.verdict と比較できます。
検証対象のevaluatorについて、devの結果から混同行列を作成します。
「fail」を陽性クラス(検出したいもの)として使用します。
| Judgeがfailと判定 | Judgeがpassと判定 | |
|---|---|---|
| 人間がfailと判定 | 真陽性 (TP) | 偽陰性 (FN) |
| 人間がpassと判定 | 偽陽性 (FP) | 真陰性 (TN) |
TPR(真陽性率) = TP / (TP + FN)
TNR(真陰性率) = TN / (TN + FP)
check_tone evaluatorのdevセット結果(42データセット):
| Judge: Fail | Judge: Pass | |
|---|---|---|
| 人間: Fail | 18 (TP) | 3 (FN) |
| 人間: Pass | 2 (FP) | 19 (TN) |
単純な正解率 = (TP + TN) / 総数 = (18 + 19) / 42 = 88.1%
一見問題なさそうに見えますが、実態を隠してしまいます。 仮にデータセットの90%がpass(クラス不均衡)の場合、 常に「pass」と答えるjudgeは正解率90%を達成しながら、 失敗を1件も検出できません(TPR = 0%)。 TPRとTNRは、失敗の検出と誤検出の防止という本質的に重要なことを計測します。
judgeが人間のラベルと一致しない全ケースについて、根本原因を特定します。
judgeが「pass」と判定したが、人間は「fail」と判定したケース。各ケースについて:
judgeが「fail」と判定したが、人間は「pass」と判定したケース。各ケースについて:
プロンプトの反復改善に役立てるため、各不一致を記録します。
| データセット | 人間 | Judge | 根本原因 | 修正内容 |
|---|---|---|---|---|
| dev_technical_pass_03 | pass | fail | 「it's」をカジュアルと判定したが、直接引用部分だった | 例外を追加:「直接引用内の短縮形は許容される」 |
| dev_ambiguous_fail_02 | fail | pass | 第3段落の微妙なトーンの変化を見逃した | 文章途中でのトーン変化を示す境界few-shotサンプルを追加 |
ステップ4の修正をjudgeの .prompt ファイルに適用し、devセットで再実行します。
npx output workflow test <workflowName> --cached --dataset <dev_datasets>
TPRとTNRを再計算します。両指標が目標値を達成するまで繰り返します。
| 指標 | 目標値 | 最低許容値 |
|---|---|---|
| TPR | > 90% | > 80% |
| TNR | > 90% | > 80% |
3〜4回の反復後も80%/80%に届かない場合は、以下を検討してください。
.prompt のフロントマターでHaikuからSonnetに切り替える各反復で確認:
.prompt ファイルに的を絞った修正を適用した(ランダムな変更はしない)devの指標が目標値を達成したら、ホールドアウトしたtestセットでjudgeを1回だけ実行します。
npx output workflow test <workflowName> --cached --dataset <test_datasets> --json
テスト結果からTPRとTNRを計算し、最終指標として記録します。
最終的な検証結果をjudgeプロンプトの隣に文書化します。
# Validation: check_tone (judge_tone@v1.prompt)
# Date: 2026-03-25
# Model: claude-haiku-4-5-20251001
## Dev Set (42 datasets)
- TPR: 90.5% (19/21)
- TNR: 95.2% (20/21)
## Test Set (43 datasets)
- TPR: 88.0% (22/25)
- TNR: 94.4% (17/18)
## Conclusion: APPROVED — both metrics above 80% minimum
この内容を VALIDATION.md ファイルとして、judgeプロンプトの隣またはevaluatorのドキュメント内に保存します。
An LLM judge is only useful if it agrees with human judgment. This skill walks you through calibrating a judge against human-labeled data using True Positive Rate (TPR) and True Negative Rate (TNR) metrics. Do this before trusting any judgeVerdict(), judgeScore(), or judgeLabel() evaluator in your eval suite.
.prompt file — Written following output-eval-judge-promptground_truth.evals.<evaluator_name>.verdict: pass or failThis process applies only to LLM-based judges. For code-based Verdict.* evaluators, write unit tests instead.
Split your labeled datasets into three groups:
| Split | % of Data | Purpose | Example (100 datasets) |
|---|---|---|---|
| Train | 10-20% | Source of few-shot examples in the judge prompt | 15 datasets |
| Dev | 40-45% | Iterate on judge prompt, measure TPR/TNR | 42 datasets |
| Test | 40-45% | Final held-out measurement, run once | 43 datasets |
Use a naming convention or subdirectories to separate splits:
Option A: Name prefixes
tests/datasets/
├── train_formal_pass_01.yml
├── train_casual_fail_01.yml
├── dev_technical_pass_01.yml
├── dev_ambiguous_fail_01.yml
├── test_simple_pass_01.yml
├── test_contradictory_fail_01.yml
└── ...
Option B: Subdirectories
tests/datasets/
├── train/
│ ├── formal_pass_01.yml
│ └── casual_fail_01.yml
├── dev/
│ ├── technical_pass_01.yml
│ └── ambiguous_fail_01.yml
└── test/
├── simple_pass_01.yml
└── contradictory_fail_01.yml
.prompt file. Never use dev or test examples — that's data leakageExecute the eval workflow against only the dev-split datasets:
# Run with cached output on dev datasets
npx output workflow test <workflowName> --cached \
--dataset dev_technical_pass_01,dev_ambiguous_fail_01,dev_formal_pass_02,...
Or if using subdirectories, list the dev dataset names:
npx output workflow test <workflowName> --cached \
--dataset $(ls tests/datasets/dev/ | sed 's/.yml//' | tr '\n' ',')
Save the output. You need the judge's verdict for each dataset to compare against ground truth.
Use --json to get machine-readable results:
npx output workflow test <workflowName> --cached --dataset <dev_datasets> --json
The output includes per-dataset, per-evaluator verdicts that you can compare against ground_truth.evals.<evaluator_name>.verdict.
For the evaluator you're validating, build a confusion matrix from the dev results.
Using "fail" as the positive class (what you're trying to detect):
| Judge says Fail | Judge says Pass | |
|---|---|---|
| Human says Fail | True Positive (TP) | False Negative (FN) |
| Human says Pass | False Positive (FP) | True Negative (TN) |
TPR (True Positive Rate) = TP / (TP + FN)
TNR (True Negative Rate) = TN / (TN + FP)
Dev set results for check_tone evaluator (42 datasets):
| Judge: Fail | Judge: Pass | |
|---|---|---|
| Human: Fail | 18 (TP) | 3 (FN) |
| Human: Pass | 2 (FP) | 19 (TN) |
Raw accuracy = (TP + TN) / total = (18 + 19) / 42 = 88.1%
This looks fine, but masks problems. If your dataset were 90% pass (class imbalance), a judge that always says "pass" would get 90% accuracy while catching zero failures (TPR = 0%). TPR and TNR measure what actually matters: catching failures and not crying wolf.
For every case where the judge disagrees with the human label, determine the root cause.
The judge said "pass" but the human said "fail." For each:
The judge said "fail" but the human said "pass." For each:
Track each disagreement to guide prompt iteration:
| Dataset | Human | Judge | Root Cause | Fix |
|---|---|---|---|---|
| dev_technical_pass_03 | pass | fail | Judge flagged "it's" as casual but context was a direct quote | Add exception: "Contractions within direct quotes are acceptable" |
| dev_ambiguous_fail_02 | fail | pass | Judge missed subtle tone shift in paragraph 3 | Add borderline few-shot example showing mid-text tone drift |
Apply the fixes from Step 4 to the judge .prompt file. Then re-run on the dev set:
npx output workflow test <workflowName> --cached --dataset <dev_datasets>
Recompute TPR and TNR. Repeat until both metrics meet the target.
| Metric | Target | Minimum Acceptable |
|---|---|---|
| TPR | > 90% | > 80% |
| TNR | > 90% | > 80% |
If you can't reach 80%/80% after 3-4 iterations:
.prompt frontmatterEach iteration:
.prompt file (not random changes)Once dev metrics meet the target, run the judge on the held-out test set exactly once:
npx output workflow test <workflowName> --cached --dataset <test_datasets> --json
Compute TPR and TNR on the test results. Record these as the final metrics.
Document the final validation results alongside the judge prompt:
# Validation: check_tone (judge_tone@v1.prompt)
# Date: 2026-03-25
# Model: claude-haiku-4-5-20251001
## Dev Set (42 datasets)
- TPR: 90.5% (19/21)
- TNR: 95.2% (20/21)
## Test Set (43 datasets)
- TPR: 88.0% (22/25)
- TNR: 94.4% (17/18)
## Conclusion: APPROVED — both metrics above 80% minimum
Store this in a VALIDATION.md file next to the judge prompt or in the evaluator's documentation.
output-eval-judge-prompt — Design the judge prompt being validatedoutput-eval-error-analysis — Source of human-labeled data for validationoutput-eval-dataset-design — Generate additional labeled datasets if you need more dataoutput-dev-eval-testing — output workflow test CLI, --cached and --dataset flagsoutput-eval-audit — Audit whether existing judges have been validated原文・著作権は Anthropic および各プラグイン作者に帰属します。日本語訳は Claude API による自動翻訳です。