HyperPod クラスター(EKS および Slurm)から診断ログと設定情報を収集し、包括的な問題報告書を生成します。トラブルシューティング(問題解決)および AWS サポートへのサポートケースに活用できます。 **次のような場合に使用:** - ユーザーが HyperPod クラスターのノード(コンピュータのグループ構成要素)から診断情報を収集する必要がある - AWS サポート向けの問題報告書を生成する - ノードの障害やパフォーマンス問題を調査する - クラスターの状態を記録する - 診断スナップショット(特定時点の状態の記録)を作成する 複数のノードからログとシステム情報を集める必要がある問題報告書の生成、診断情報の収集、サポートケースの準備、またはクラスターのトラブルシューティングに関するリクエストが対象となります。
Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases. Use when users need to collect diagnostics from HyperPod cluster nodes, generate issue reports for AWS Support, investigate node failures or performance problems, document cluster state, or create diagnostic snapshots. Triggers on requests involving issue reports, diagnostic collection, support case preparation, or cluster troubleshooting that requires gathering logs and system information from multiple nodes.
HyperPodクラスタのノード(コンピュート資源を構成する個別マシン)から診断ログを収集し、S3(Amazon クラウドストレージサービス)に保存します。EKSとSlurm(計算ジョブ管理システム)の両方のクラスタに対応し、自動判定します。同梱の scripts/hyperpod_issue_report.py を使用して、信頼性の高い並列収集を実現します。
sagemaker:DescribeCluster、sagemaker:ListClusterNodes、ssm:StartSession、s3:PutObject、s3:GetObject、eks:DescribeClusters3:GetObject/s3:PutObject が必要ユーザーから以下を確認:
arn:aws:sagemaker:us-west-2:123456789012:cluster/abc123)s3://bucket/prefix)。ユーザーがバケットを持っていない場合は作成(例: s3://hyperpod-diagnostics-<account-id>-<region>)aws sts get-caller-identity
aws sagemaker describe-cluster --cluster-name <name-or-arn> --region <region>
S3バケットが存在しない場合は作成:
aws s3 mb s3://<bucket-name> --region <region>
EKSクラスタの場合(describe-clusterの出力で Orchestrator.Eks を確認):
kubectlがインストール済みか確認(which kubectl)。なければ現在のプラットフォーム用にインストール
describe-clusterの応答から得たEKSクラスタ名を使用して kubeconfig を設定:
aws eks update-kubeconfig --name <eks-cluster-name> --region <region>
uv run scripts/hyperpod_issue_report.py \
--cluster <cluster-name-or-arn> \
--region <region> \
--s3-path s3://<bucket>[/prefix]
--help で全オプション(--instance-groups、--nodes、--max-workers、--debugなど)を表示。--instance-groups と --nodes は同時指定不可。ノード識別子はインスタンスID(i-*)、EKS名(hyperpod-i-*)、またはSlurm名(ip-*)に対応。
収集終了後、スクリプトが統計情報を表示し、対話的なダウンロードを提供します。S3の保存先を報告し、以下の対応を提案:
エラー対応、大規模クラスタの調整方法、既知の制限事項については references/troubleshooting.md を参照してください。
Collect diagnostic logs from HyperPod cluster nodes via SSM, store results in S3. Supports both EKS and Slurm clusters with auto-detection. Uses the bundled scripts/hyperpod_issue_report.py for reliable parallel collection.
sagemaker:DescribeCluster, sagemaker:ListClusterNodes, ssm:StartSession, s3:PutObject, s3:GetObject, eks:DescribeClusters3:GetObject/s3:PutObject on the report bucketCollect from the user:
arn:aws:sagemaker:us-west-2:123456789012:cluster/abc123)s3://bucket/prefix). If the user doesn't have a bucket, create one (e.g., s3://hyperpod-diagnostics-<account-id>-<region>)aws sts get-caller-identity
aws sagemaker describe-cluster --cluster-name <name-or-arn> --region <region>
If the S3 bucket doesn't exist, create it:
aws s3 mb s3://<bucket-name> --region <region>
For EKS clusters (check Orchestrator.Eks in describe-cluster output):
Ensure kubectl is installed (which kubectl). If missing, install it for the current platform.
Configure kubeconfig using the EKS cluster name from the describe-cluster response:
aws eks update-kubeconfig --name <eks-cluster-name> --region <region>
uv run scripts/hyperpod_issue_report.py \
--cluster <cluster-name-or-arn> \
--region <region> \
--s3-path s3://<bucket>[/prefix]
Use --help for all options including --instance-groups, --nodes, --max-workers, and --debug. Note: --instance-groups and --nodes are mutually exclusive. Node identifiers accept instance IDs (i-*), EKS names (hyperpod-i-*), or Slurm names (ip-*).
After collection, the script shows statistics and offers interactive download. Report the S3 location and offer to:
See references/troubleshooting.md for error handling, large cluster tuning, and known limitations.
原文・著作権は Anthropic および各プラグイン作者に帰属します。日本語訳は Claude API による自動翻訳です。