• Projects
  • Service
  • About
  • branding.bz
  • Podcast
  • Tips
  • FAQ
  • Recruit
  • Download
  • Contact
  • branding.bz(ブランド構築SaaS)
  • DESIGN NOW(デザインメディア)
  • X
  • LinkedIn
  • Spotify
  • Facebook

213-0011 神奈川県川崎市高津区久本3-6-7-303

© 2026 ID INC. All rights reserved

claude-skills/スキル
SKILLOfficialdevelopment

debugging-mwaa-workflow

プラグイン
aws-data-analytics
ソース
GitHub で見る ↗
説明

Amazon MWAAのワークフロー実行失敗を診断し、根本原因を特定するスキルです。 **対応環境** - プロビジョンド環境(Python DAGを使用) - サーバーレス環境(YAML形式のワークフローを使用) **診断内容** 実行失敗やタスク失敗、DAGが表示されない、インポートエラー、ワーカーのメモリ不足、IAM権限の拒否、依存関係の不整合など、様々な問題に対応しています。 **次のような場合に使用:** - DAGやタスク、ワークフロー実行が失敗した - MWAAエラーが発生している - DAGをデバッグしたい、またはワークフローが失敗した理由を知りたい - DAGが表示されていない - インポートエラーや依存関係のインストール失敗が起きている - ワーカーがクラッシュしている - サーバーレスでの実行に失敗している **対応していない場合** ワークフロー作成時の支援、ワークフロー実行後の検証・確認テスト、CI/CDパイプラインでのデプロイ失敗については、それぞれ専用のスキルで対応しています。

原文を表示

Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments. Provisioned uses aws mwaa invoke-rest-api, CloudWatch log groups, and get-environment; Serverless uses aws mwaa-serverless API (GetWorkflowRun, ListWorkflowRuns, GetTaskInstance) and CloudWatch logs. Covers failed runs and tasks, DAGs not appearing, import errors, worker OOM, IAM denials, and dependency drift. Triggers on: DAG failed, task failed, workflow run failed, MWAA error, debug my DAG, why did my workflow fail, DAG not showing up, MWAA import error, requirements failing, worker crashed, serverless run failed. Not applicable to authoring workflows (handled by authoring-mwaa-workflow), running or smoke-testing a workflow (handled by testing-mwaa-workflow), or CI-CD deploy failures.

ユースケース
  • DAGやタスク、ワークフロー実行が失敗した
  • MWAAエラーが発生している
  • DAGをデバッグしたい、またはワークフロー失敗の原因を知りたい
  • DAGが表示されていない
  • インポートエラーや依存関係のインストール失敗が起きている
  • ワーカーがクラッシュしている
本文(日本語訳)

MWAA ワークフローのデバッグ

AWS MCP サーバー(オプション・推奨): このスキルで AWS CLI コマンドを AWS MCP サーバー経由で実行すると、サンドボックス環境での実行と監査ログが可能になります。ここのコマンドはすべてプレーン AWS CLI でも動作するため、このスキルは MCP サーバーや MCP 専用ツールを必須としません。

Amazon MWAA ワークフローの障害を診断し、根本原因を特定します。その後、根本原因、影響範囲、推奨される対処方法を報告します。環境のタイプで振り分けた後、共通の 4 ステップ診断プロセスを実行します。

守則 — このスキル自体のファイルの置き場所(MCP 対 ローカルインストール)

このスキルは 2 つの方法で読み込むことができ、スキル自体のバンドルファイルをそれぞれ異なる場所から取得します。参照を読む前に、スキルの読み込み方法を確認してください。

  • AWS MCP の retrieve_skill ツール経由で読み込まれた場合: スキルはローカルファイルシステムにインストールされていません。retrieve_skill を使用して、file パラメーター(例:file="references/failure-catalog.md")で各参照を取得し、返されたコンテンツを読む必要があります。これらのパスをローカルで file_read しないでください。ディスクに存在しません。
  • ローカルインストール(例:.kiro/skills/debugging-mwaa-workflow/ または ~/.claude/skills/debugging-mwaa-workflow/): スキルディレクトリから相対パスを使用してファイルを読んでください。

この区別は、スキル独自のパッケージファイルにのみ適用されます。ユーザーデータとセッション成果物は常にユーザーの作業ディレクトリから読み書きします。顧客データを retrieve_skill 経由で取得または書き込みしないでください。

ステップ 0:環境タイプを判定し、障害の範囲を特定する

環境タイプを判定する

  1. aws mwaa get-environment で解決できる環境名は、その環境が プロビジョニング型 です。
  2. workflow/... ARN または aws mwaa-serverless コンテキストは、その環境が サーバーレス型 です。単なる実行 ID だけでは環境タイプの判定にはなりません。プロビジョニング型の DAG 実行にも実行 ID があります。
  3. どちらのシグナルも見当たらない場合は、次を確認してください:対象は MWAA プロビジョニング型(Python DAG) ですか、それとも MWAA サーバーレス型(YAML ワークフロー) ですか?

複雑さで振り分ける

  • シンプル — 名前付きの単一タスクまたは実行が明確な例外で失敗した。ステップ 2 に進みます。
  • 標準 — 実行が失敗し、原因が不明。ステップ 1~4 の全体スイープを実行します。
  • 複雑 — 断続的、または環境全体の障害(複数の DAG、「昨日は動作していた」、何も表示されない)。重点を置いてステップ 3 を強化した全体スイープを実行します。

ステップ 1:障害を特定する

プロビジョニング型: aws mwaa invoke-rest-api を使用して、失敗した DAG 実行とタスクインスタンスをリストアップします(パス:/dags/{id}/dagRuns および /dags/{id}/dagRuns/{run_id}/taskInstances)。invoke-rest-api がエラーになった場合(RestApiClientException)、Scheduler および DAGProcessing ロググループにフォールバックします。バージョンと設定は aws mwaa get-environment から取得します。references/provisioned-diagnostics.md を参照。

サーバーレス型: aws mwaa-serverless list-workflow-runs を実行し、その後 get-workflow-run を実行します。RunDetail.ErrorMessage を読みます。TaskInstances が空でパーサーメッセージがあれば定義エラー、Workflow execution failed で TaskInstances が入っていればタスク実行の失敗です。references/serverless-diagnostics.md を参照。

ステップ 2:エラー詳細を取得し、カテゴリー分けする

テンプレート文言を除いた実際の例外を抽出します。

プロビジョニング型: Task ロググループを最初に読み、症状に応じて Worker/Scheduler/DAGProcessing を読みます。

サーバーレス型: list-task-instances を実行し、その後 get-task-instance で各タスクの LogStream を取得してから、CloudWatch でそのストリームを読みます。

次に、優先順位に従って カテゴリー分けします。infra、drift、code-data の順で references/failure-catalog.md を使用します。カテゴリーがステップ 3 の確認項目を決定します。

ステップ 3:コンテキストを確認する(なぜ起こったのか)

references/failure-catalog.md から、マッチしたカテゴリーのコンテキスト確認項目を実行します。表面の例外で止まらないでください。SIGKILL はメモリ不足のサイン、コード変更のない新規インポートエラーは環境変化のサイン、センサータイムアウトはアップストリーム側の健全性問題のサインかもしれません。

ステップ 4:実行可能な出力を提供する

この正確な構造で報告します。

Root Cause(根本原因): <それを証明する証拠を含めた 1 行の診断>
Impact(影響): <何が失敗したか、どの実行か、影響範囲>
Immediate Fix(今すぐできる対処): <ブロックを解除する最小限の変更>
Prevention(再発防止): <再発を防ぐ変更>
Commands(コマンド): <実行した読み取り専用コマンド、および ユーザーが実行する対処用コマンド>

読み取り専用の操作のみを実行します。状態を変更する対処(プロビジョニング型のクリア/再実行/バックフィル、サーバーレス型の workflow 実行開始または修正・再デプロイ)はユーザーが実行するコマンドとして提示し、影響を明記します。本番環境の安全性のため、自律的に実行しないでください。

修正・再デプロイの経路については、authoring-mwaa-workflow を使用して準拠したアーティファクトを再生成します。

よくある落とし穴

  • サーバーレス型: Airflow ウェブ UI、REST API、CLI トークンがありません。create-web-login-token、invoke-rest-api、またはサーバーレス型に対する Airflow REST パスは使用しないでください。
  • プロビジョニング型: 常に aws mwaa invoke-rest-api を使用してください(create-web-login-token + curl ではなく)。invoke-rest-api はネットワークアクセスのない VPC 専用ウェブサーバーに到達できます。
  • GetWorkflowRun.RunDetail.ErrorMessage は定義エラー(TaskInstances が空)とタスク実行の失敗(Workflow execution failed、TaskInstances が入っている)を区別します。タスクログを取得する前に読みます。
  • サーバーレス型のロググループは /aws/mwaa-serverless/{workflow-id}/ がデフォルトですが、カスタムグループも可能です。パスを想定する前に get-workflow の LoggingConfiguration で確認します。
  • DAG が表示されない原因はいくつかあります。インポート/パースエラー、スケジューラースキャン間隔(scheduler.dag_dir_list_interval、Airflow 3.x では dag_processor.refresh_interval)がまだ経過していない、dag_id の競合、S3 同期の遅延など。これは DAG が壊れていることは稀です。invoke-rest-api で GET /importErrors と GET /dags/{dag_id} を確認し、DAGProcessing ログも見てください。失敗カタログの「DAG が UI に表示されない」チェックリストを参照してからコードが間違っていると結論づけます。
  • ワーカー SIGKILL はメモリ不足のシグナルです。Glue/EMR/Lambda に作業を移す方法をお勧めします。ワーカーのスケーリングだけではタスク単位のメモリプレッシャーは修正されません。
  • MWAA は環境更新時に依存関係を再解決するため、コード変更のない DAG が新たにインポートで失敗し始めることがあります。コード変更のないインポート失敗は環境変化として扱います。
  • サーバーレス型の PythonOperator/BashOperator タスクは --code パッケージからカスタムコードを実行します。パッケージ抽出に失敗する実行または ModuleNotFoundError は、YAML 定義エラーではなくパッケージング問題(プラットフォーム不一致ホイール、欠けている依存関係、不正なレイアウト)です。失敗カタログのサーバーレスカスタムコード部分を参照してください。

トラブルシューティング

エラー 原因 修正
RestApiClientException(プロビジョニング型) 実行ロールのスコープ誤りまたはサービスエラー Scheduler/DAGProcessing ロググループにフォールバック
ResourceNotFoundException(get-workflow-run) ワークフロー ARN または実行 ID が誤り list-workflow-runs で再度リストアップ
タスクログストリームが空(サーバーレス型) ロググループを誤って想定 get-workflow から LoggingConfiguration を読み取る
タスクログなし、実行は FAILED 定義/パースエラー RunDetail.ErrorMessage を読み、YAML を修正

参照

  • references/provisioned-diagnostics.md — プロビジョニング型のデータソースと読み取り専用コマンド
  • references/serverless-diagnostics.md — サーバーレス型の API、ロググループ、障害クラス
  • references/failure-catalog.md — カテゴリー別のコンテキスト確認と対処

セキュリティ上の考慮事項

  • デフォルトは読み取り専用: 診断は読み取り/リスト/説明呼び出しのみを使用します。対処(クリア/再実行/バックフィル、IAM またはキーポリシー変更)はユーザーが実行するコマンドとして提示され、自律実行されません。
  • 最小権限の IAM: AccessDenied が本来のパーミッション不足の場合、エラーの最小限の Action/Resource をお勧めします。ワイルドカード使用は避けてください。存在しないリソースの誤入力と区別し、IAM を拡大しないでください。
  • クロスアカウント: KMS キーポリシー / assume-role 変更は人が確認し、リソース所有者と調整します。
  • シークレット非公開: ログまたは API 応答から認証情報または接続文字列を診断出力に含めないでください。
原文(English)を表示

Debugging MWAA Workflows

AWS MCP server (optional but recommended): running the AWS CLI commands in this skill through the AWS MCP server gives sandboxed execution and audit logging. Every command here also works with the plain AWS CLI, so the skill does not require the MCP server or any MCP-only tools.

Diagnose and root-cause Amazon MWAA workflow failures, then report root cause, impact, and recommended remediation. Routes by flavor, then runs a shared 4-step diagnostic spine.

Guardrail — where this skill's own files live (MCP vs local install)

This skill can be loaded two ways, and they resolve the skill's own bundled files from different places. Determine how the skill was loaded before reading a reference:

  • Loaded through the AWS MCP retrieve_skill tool: The skill is not installed on the local filesystem. You MUST fetch each reference via retrieve_skill with the file parameter (e.g. file="references/failure-catalog.md") and read the returned content. Do NOT file_read these paths locally — they do not exist on disk.
  • Installed locally (e.g. .kiro/skills/debugging-mwaa-workflow/ or ~/.claude/skills/debugging-mwaa-workflow/): Read the files from the local skill directory using relative paths.

This distinction applies only to the skill's own packaged files. User data and session artifacts are always read from and written to the user's working directory. Never fetch or write customer data through retrieve_skill.

Step 0: Detect Flavor and Scope the Failure

Detect flavor

  1. An environment name resolvable via aws mwaa get-environment means the environment is Provisioned.
  2. A workflow/... ARN or any aws mwaa-serverless context means the environment is Serverless. A bare run identifier does NOT indicate flavor — Provisioned DAG runs also have run ids.
  3. If neither signal is present, ask: is the target MWAA Provisioned (Python DAG) or MWAA Serverless (YAML workflow)?

Route by complexity

  • Simple — a single named task or run failed with a clear exception. Jump to Step 2 for that task.
  • Standard — a run failed and the cause is unknown. Run the full Step 1 to Step 4 sweep.
  • Complex — intermittent or environment-wide (multiple DAGs, "worked yesterday", nothing appearing). Run the full sweep with emphasis on Step 3.

Step 1: Identify the Failure

Provisioned: list failed DAG runs and task instances via aws mwaa invoke-rest-api (paths /dags/{id}/dagRuns and /dags/{id}/dagRuns/{run_id}/taskInstances). If invoke-rest-api errors (RestApiClientException), fall back to the Scheduler and DAGProcessing log groups. Get version and config from aws mwaa get-environment. See references/provisioned-diagnostics.md.

Serverless: aws mwaa-serverless list-workflow-runs, then get-workflow-run. Read RunDetail.ErrorMessage — an empty TaskInstances with a parser message is a definition error; Workflow execution failed with populated TaskInstances is a task-execution failure. See references/serverless-diagnostics.md.

Step 2: Get Error Details and Categorize

Pull the real exception past boilerplate:

Provisioned: read the Task log group first, then Worker/Scheduler/ DAGProcessing as the symptom directs.

Serverless: list-task-instances then get-task-instance to get each task's LogStream, then read that stream in CloudWatch.

Then categorize in priority order — infra, then drift, then code-data — using references/failure-catalog.md. The category determines the Step 3 checks.

Step 3: Check Context (Why It Happened)

Run the context checks for the matched category from references/failure-catalog.md. Do not stop at the surface exception: a SIGKILL is an OOM story, a fresh import error on unchanged code is a drift story, a sensor timeout is an upstream-health story.

Step 4: Provide Actionable Output

Report in this exact structure:

Root Cause: <one-line diagnosis with the evidence that proves it>
Impact: <what failed, which runs, blast radius>
Immediate Fix: <the smallest change that unblocks>
Prevention: <the change that stops recurrence>
Commands: <exact read-only commands run, plus remediation commands for the user to run>

Run only read-only operations. Present state-mutating remediation (clear/rerun/backfill for Provisioned; start-workflow-run or fix-and-redeploy for Serverless) as commands for the user to run, with the impact stated. Never execute them autonomously (production safety).

For the fix-and-redeploy path, use authoring-mwaa-workflow to regenerate a compliant artifact.

Gotchas

  • Serverless has no Airflow web UI, no REST API, and no CLI token. Do not attempt create-web-login-token, invoke-rest-api, or any Airflow REST path for Serverless.
  • For Provisioned, always use aws mwaa invoke-rest-api (not create-web-login-token + curl). invoke-rest-api reaches VPC-only web servers without network access.
  • GetWorkflowRun.RunDetail.ErrorMessage distinguishes a definition error (empty TaskInstances) from a task-execution failure (Workflow execution failed, populated TaskInstances). Read it before pulling task logs.
  • The Serverless log group defaults to /aws/mwaa-serverless/{workflow-id}/ but can be a custom group; confirm via get-workflow LoggingConfiguration before assuming the path.
  • A DAG not appearing has several causes — an import/parse error, the scheduler scan interval (scheduler.dag_dir_list_interval, or dag_processor.refresh_interval on Airflow 3.x) not yet elapsed, a dag_id collision, or S3-sync delay — and is rarely a broken DAG. Check GET /importErrors and GET /dags/{dag_id} via invoke-rest-api (and the DAGProcessing logs); see the failure catalog's "DAG not appearing in the UI" checklist before concluding the code is wrong.
  • A worker SIGKILL is an OOM signal. Recommend moving work to Glue/EMR/Lambda; scaling workers alone does not fix per-task memory pressure.
  • MWAA re-resolves dependencies on environment update, so an unchanged DAG can start failing on import with no code change. Treat no-code-change import failures as drift.
  • Serverless PythonOperator/BashOperator tasks run custom code from a --code package. A run that fails to extract the package or hits ModuleNotFoundError is a packaging problem (wrong-platform wheel, missing dep, bad layout), not a YAML definition error. See the failure catalog's serverless custom-code section.

Troubleshooting

Error Cause Fix
RestApiClientException (Provisioned) Mis-scoped execution role or service error Fall back to Scheduler/DAGProcessing log groups
ResourceNotFoundException on get-workflow-run Wrong workflow ARN or run id Re-list with list-workflow-runs
Task log stream empty (Serverless) Wrong log group assumed Read LoggingConfiguration from get-workflow
No task logs but run FAILED Definition/parse error Read RunDetail.ErrorMessage; fix the YAML

References

  • references/provisioned-diagnostics.md — Provisioned data sources and read-only commands
  • references/serverless-diagnostics.md — Serverless API, log group, failure classes
  • references/failure-catalog.md — category-keyed context checks and remediation

Security Considerations

  • Read-only by default: diagnosis uses only read/list/describe calls. Remediation (clear/rerun/backfill, IAM or key-policy changes) is presented as commands for the user to run, never executed autonomously.
  • Least-privilege IAM: when an AccessDenied is a genuine permission gap, recommend the minimal Action/Resource from the error — never a wildcard; distinguish it from a nonexistent-resource typo (do not broaden IAM then).
  • Cross-account: KMS key-policy / assume-role changes are human-gated and coordinated with the resource owner.
  • No secret exposure: do not surface credentials or connection strings from logs or API responses in the diagnosis output.

原文・著作権は Anthropic および各プラグイン作者に帰属します。日本語訳は Claude API による自動翻訳です。