Azure Data Lake Storage Gen2(`abfss://`)からのデータ読み書きをAIDP ノートブックで実行します。ユーザーが ADLS、Azure Data Lake、abfss を言及したり、マルチクラウド対応の Azure ソースからデータを取り込みたい場合に使用します。認証方式は OAuth クライアント認証情報(Service Principal のクライアント ID、シークレット、テナント)を使用します。
Read and write Azure Data Lake Storage Gen2 (`abfss://`) from an AIDP notebook. Use when the user mentions ADLS, Azure Data Lake, abfss, or wants to ingest from a multi-cloud Azure source. Auth is OAuth client-credentials (Service Principal client_id + secret + tenant).
aidp-azure-adls — OAuthクライアント資格情報を使用したAzure ADLS Gen2連携Service Principalを使用して、AIDP Sparkから abfss://<container>@<storage_account>.dfs.core.windows.net/... パスの読み書きを行います。
aidp-object-storageaidp-aws-s3Service Principalの資格情報を使用してSpark Hadoopコネクタを設定します。 セッション/ジョブごとに1回実施してください。
import os
storage_account = os.environ["ADLS_STORAGE_ACCOUNT"] # アカウント名のみ(.dfs... は不要)
client_id = os.environ["ADLS_CLIENT_ID"] # SP アプリケーション(クライアント)ID
client_secret = os.environ["ADLS_CLIENT_SECRET"] # SP シークレット値
tenant = os.environ["ADLS_TENANT"] # Azure AD テナント ID(GUID)
base = f"fs.azure.account"
host = f"{storage_account}.dfs.core.windows.net"
spark.conf.set(f"{base}.auth.type.{host}", "OAuth")
spark.conf.set(f"{base}.oauth.provider.type.{host}", "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider")
spark.conf.set(f"{base}.oauth2.client.id.{host}", client_id)
spark.conf.set(f"{base}.oauth2.client.secret.{host}", client_secret)
spark.conf.set(f"{base}.oauth2.client.endpoint.{host}", f"https://login.microsoftonline.com/{tenant}/oauth2/token")
container = os.environ["ADLS_CONTAINER"]
data_file = os.environ["ADLS_DATA_FILE"] # 例: "data/2025/january/orders.csv"
df = (spark.read
.format("csv")
.option("header", True)
.load(f"abfss://{container}@{storage_account}.dfs.core.windows.net/{data_file}"))
df.show()
(df.write
.mode("overwrite")
.format("delta")
.saveAsTable("default.default.data_from_adls"))
Service PrincipalにストレージアカウントへのRBACを付与する必要があります。
コンテナまたはアカウントに対して Storage Blob Data Contributor(読み取り専用の場合は Reader)を割り当ててください。
abfss:// を使用するには、ストレージアカウントで階層型名前空間(Hierarchical Namespace)を有効にする必要があります。
(ADLS Gen2 = HNS が有効なストレージアカウント)
シークレットは環境変数で管理し、ノートブックにハードコードしないでください。
.gitignore に登録した .env ファイル、またはOCI Vaultの oracle_ai_data_platform_connectors.auth.secrets.get(name) 経由で取得することを推奨します。
エンドポイントURL — login.microsoftonline.com/<tenant>/oauth2/token はv1エンドポイントであり、ClientCredsTokenProvider が期待する形式です。ここではv2エンドポイントを使用しないでください。
abfs:// ではなく abfss:// を使用してください。 常にTLS対応のバリアントを使用してください。
aidp-azure-adls — Azure ADLS Gen2 via OAuth client-credentialsRead or write abfss://<container>@<storage_account>.dfs.core.windows.net/... paths from AIDP Spark using a Service Principal.
aidp-object-storage.aidp-aws-s3.Configure the Spark Hadoop connector with Service-Principal credentials. Do this once per session/job:
import os
storage_account = os.environ["ADLS_STORAGE_ACCOUNT"] # account name only, no .dfs...
client_id = os.environ["ADLS_CLIENT_ID"] # SP application (client) id
client_secret = os.environ["ADLS_CLIENT_SECRET"] # SP secret value
tenant = os.environ["ADLS_TENANT"] # Azure AD tenant id (GUID)
base = f"fs.azure.account"
host = f"{storage_account}.dfs.core.windows.net"
spark.conf.set(f"{base}.auth.type.{host}", "OAuth")
spark.conf.set(f"{base}.oauth.provider.type.{host}", "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider")
spark.conf.set(f"{base}.oauth2.client.id.{host}", client_id)
spark.conf.set(f"{base}.oauth2.client.secret.{host}", client_secret)
spark.conf.set(f"{base}.oauth2.client.endpoint.{host}", f"https://login.microsoftonline.com/{tenant}/oauth2/token")
container = os.environ["ADLS_CONTAINER"]
data_file = os.environ["ADLS_DATA_FILE"] # e.g. "data/2025/january/orders.csv"
df = (spark.read
.format("csv")
.option("header", True)
.load(f"abfss://{container}@{storage_account}.dfs.core.windows.net/{data_file}"))
df.show()
(df.write
.mode("overwrite")
.format("delta")
.saveAsTable("default.default.data_from_adls"))
Storage Blob Data Contributor (or Reader for read-only) on the container or the account.abfss:// to work (ADLS Gen2 = HNS-on storage account)..env file gitignored, or from OCI Vault via oracle_ai_data_platform_connectors.auth.secrets.get(name).login.microsoftonline.com/<tenant>/oauth2/token is the v1 endpoint and is what the ClientCredsTokenProvider expects. Don't use the v2 endpoint here.abfss:// not abfs:// — always use the TLS variant.原文・著作権は Anthropic および各プラグイン作者に帰属します。日本語訳は Claude API による自動翻訳です。