AIボイスエージェント(音声で自動応答するAI)のプロンプト(指示文)を検査し、以下の問題がないか確認します。特にSynthflow Single-Promptを対象とします: - 矛盾や不整合 - 文脈の欠落 - 不安全なツール・動作 - 対応の段階的エスカレーション(問題の段階的な引き継ぎ)における欠陥 - コンプライアンス(規制対応)上のリスク - 機能低下や劣化のリスク 次のような場合に使用:ユーザーがプロンプトを導入前に検査・監査・検証・改善・比較することを依頼した場合
Review AI voice-agent prompts, especially Synthflow Single-Prompt, for contradictions, missing context, unsafe tool/action behavior, escalation gaps, compliance risk, and regression risk. Use this when the user asks to review, audit, validate, improve, or compare a prompt before deployment.
AIエージェントのプロンプト(指示文)を、テスト・本番環境へのデプロイ・編集の前に確認します。プロンプトが信頼性があり、具体的で、安全であり、利用可能なツール・アクション(機能)と一致しており、実際の顧客対応に適しているかどうかに焦点を当てます。
特に以下の場合に有用です:Synthflowの音声エージェント、シングルプロンプトエージェント、フローデザイナーのノードプロンプト、転送フロー、エスカレーション(問題の段階的引き上げ)ロジック、サポート・営業・適格性判定・予約エージェント、そして通話動作を制御するあらゆるプロンプト。
Synthflowワークスペースに格納されているプロンプトの場合、Synthflow MCPツールで取得します:get_agentで現在のプロンプトと設定を、get_agent_actions/get_actionでプロンプトが従う必要があるアクション説明を、list_agent_versions/get_agent_version_diffで編集時の回帰リスク(変更による副作用)を比較する際に使用します。Synthflowの機能の実装を確認するには、search_docsで公式ドキュメントを検索するか、ワークスペースツールがまだ接続されていない場合はsynthflow-docsサーバーのsearchDocsツールで検索します。
優れたプロンプトは、背景情報がない人間のオペレーターにも理解できます。
「別の人が読むテスト」(OpenAIが推奨する確認方法):プロンプトを同僚に読み聞かせてください。同僚がエージェントの役割を平易な言葉で言い直せなければ、言語モデルも確実には実行できません。
背景情報ゼロで18歳が同じやり方を2回実行できる明確さを目指してください。曖昧な性格描写より、具体的な行動指示を優先します。
曖昧な動詞を、具体的でテスト可能な行動に置き換えます:
| ❌ 曖昧 | ✅ 具体的 |
|---|---|
| 「発信者に専門的に対応する」 | 「1文/ターン、相手の調子に合わせ、俗語なし」 |
| 「適切なバッテリーを見つけるのを手伝う」 | 「年式/メーカー/車種/グレードを質問、在庫ある2商品をSKU・価格・保証付きで提示」 |
| 「顧客を確認する」 | 「名前・電話・注文参照番号を収集、各項目を読み返し、『合っていますか、はい或いはいいえ?』と質問」 |
| 「親切であること」 | (削除 — 実行不可能) |
Synthflowのプロンプトは、以下の3つのセクションにこの順序で、見出し付きで分けて構成すべきです:
| セクション | 内容 |
|---|---|
| 背景 | エージェントの身分、企業、状況、店舗・アカウントのコンテキスト、方針、役割、目標、発信者の文脈、変数、顧客固有の制限 |
| 動作 | 条件分岐ロジック、フロー、「Xの場合Yを実行」、ツール・アクショントリガー(起動条件)、予約・転送・エスカレーション・フォールバック・通話終了ルール、反論・エラー対応 |
| 出力 | 文の長さ、書式、言語、数字・日付・金額の音声表現 — 書式設定のみ |
報告すべき典型的な配置ミス:出力セクション内の条件ルール(「商用発信者の場合は転送」)、最上部の「重大ルール」ブロック(3つすべてを混在)。フローの「出力」がツール呼び出しである場合、それは動作セクションに属します。
注記 — 決定的なステップバイステップフロー: 多くのエンタープライズ顧客は「ステップ1、2、7」形式のプロンプトを作成します。分け方は同じですが、出力は書式設定・スタイルのみに縮小します。構造変更を強制するのではなく、この仮定を述べてください。
これらは客観的で自動化できるほどです。失敗があれば報告してください。
{vars}インライン参照はプロンプトキャッシング(中程度の副作用)を破壊。詳細と修正表:**references/mechanical-checks.md**を参照。
これらの分類すべてについてプロンプトをレビューしてください。分類ごとの詳細チェックリストは**references/review-checklists.md**にあります — 結果を検討する際に確認してください。
利用可能なツールコンテキストが必要とマークし、不完全と判定してください — 推測で補わないでください。アクション説明の作成ルール:**references/action-descriptions.md**を参照。{vars}、物理的に不可能な指示)**references/output-template.md**の構造を使用してレビューを作成します(エグゼクティブサマリ → 結果 → 不足情報 → 回帰リスク → 推奨テスト → 改善プロンプト案・任意)。該ファイルには採点ルーブリックと具体例も含まれます。
結果は実行可能に — すべての結果に具体的な推奨を含め、「これをもっと明確に」ではなく。
レビューが完全なのは、以下に答える場合のみです:(1)実通話で何が悪くなるか、(2)プロンプトに基づく理由、(3)プロンプトの変更方法、(4)デプロイ前のテスト内容。
改善提案は最大5個。最終レポート後、ユーザーに変更の適用を希望するかどうか確認してください。
Review an AI agent prompt before it is tested, deployed, or edited in production. Focus on whether the prompt is reliable, specific, safe, aligned with available tools/actions, and suitable for real customer conversations.
Especially useful for Synthflow voice agents, Single-Prompt agents, Flow Designer node prompts, transfer flows, escalation logic, support/sales/qualification/booking agents, and any prompt that controls call behavior.
When the prompt lives in a Synthflow workspace rather than in the conversation, pull it with the Synthflow MCP tools: get_agent for the live prompt and settings, get_agent_actions / get_action for the action descriptions the prompt must agree with, and list_agent_versions / get_agent_version_diff when comparing an edit for regression risk. To check how a Synthflow feature actually works, search the official docs with search_docs, or with the searchDocs tool on the synthflow-docs server if the workspace tools aren't connected yet.
A good prompt is understandable by a human operator with no hidden context.
The "another human" test (OpenAI's recommended check): read the prompt aloud to a teammate. If they can't restate what the agent should do in plain English, the LLM won't reliably do it either.
Write so an 18-year-old with zero context could execute it the same way twice. Prefer concrete behavioral instructions over vague personality instructions.
Be helpful and professional.Use one or two sentences per turn. Ask one question at a time. If the caller asks about pricing, give the approved price range and ask whether they want to book a consultation.Replace vague verbs with specific, testable behavior:
| ❌ Vague | ✅ Specific |
|---|---|
| "Handle the caller professionally" | "One sentence per turn, mirror their tone, no slang" |
| "Help them find the right battery" | "Ask year/make/model/trim, then present 2 in-stock SKUs with price + warranty" |
| "Verify the customer" | "Collect name, phone, order ref. Read each back. Ask 'is that correct, yes or no?'" |
| "Be helpful" | (delete — not actionable) |
A Synthflow prompt should split cleanly into three sections, in this order, with headers:
| Section | Contains |
|---|---|
| Background | Who the agent is, company, situation, store/account context, policies, role, goal, caller context, variables, customer-specific restrictions |
| Behavior | Conditional logic, flows, "if X then Y", tool/action triggers, booking/transfer/escalation/fallback/end-call rules, objection and error handling |
| Output | Sentence length, format, language, spoken forms for numbers/dates/money — formatting only |
Common misplacements to flag: conditional rules inside Output ("if commercial caller, transfer"); a "Critical Rules" block at the top mixing all three. If a flow's "output" is a tool call, that lives in Behavior, not Output.
Caveat — deterministic step-by-step flows: many enterprise customers write "step 1, 2, 7" prompts. The split still applies, but Output collapses to formatting/styling only. State this assumption rather than forcing a restructure.
These are objective and near-automatable. Flag any failure.
{vars} break prompt caching (Medium regression).Full detail and the fix-it table: references/mechanical-checks.md.
Review every prompt for these categories. Detailed checklists per category are in references/review-checklists.md — read it when working through findings.
Needs available-tool context and incomplete — do not invent them. Action-description authoring rules: see references/action-descriptions.md.{vars}; physically-impossible instruction.)Produce the review using the structure in references/output-template.md (executive summary → findings → missing information → regression risks → recommended tests → optional improved prompt). That file also contains the scoring rubric and a worked example finding.
Keep findings actionable — every finding includes a concrete recommendation, not "make this clearer."
A review is complete only when it answers: (1) what could go wrong in a real call, (2) why, based on the prompt, (3) how the prompt should change, (4) what to test before deployment.
Give a maximum of 5 improvement suggestions. After the final report, ask the user whether they want you to apply the changes.
原文・著作権は Anthropic および各プラグイン作者に帰属します。日本語訳は Claude API による自動翻訳です。