Databricks Vector Search を使用して、ベクトル検索(意味に基づいた検索)やRAG(生成AIに必要な情報を自動取得する仕組み)を実装するためのエンドポイント(接続窓口)とインデックス(検索用の索引)を提供します。 対応する内容: - インデックスの種類 - 検索モード - RAGの一連の実装パターン
Databricks Vector Search endpoints and indexes for RAG and semantic search; covers index types, search modes, end-to-end RAG patterns
はじめに: CLI の基本操作・認証・プロファイル選択には、親スキルである databricks-core を先に参照してください。
RAG およびセマンティック検索アプリケーション向けに、Vector Search インデックスの作成・管理・クエリを行うためのパターン集です。
次のような場合に使用:
Databricks Vector Search は、自動 Embedding 生成と Delta Lake 連携を備えた、マネージド型のベクトル類似検索サービスです。
| コンポーネント | 説明 |
|---|---|
| Endpoint | インデックスをホストするコンピュートリソース(Standard または Storage-Optimized) |
| Index | 類似検索用のベクトルデータ構造 |
| Delta Sync | ソースの Delta テーブルと自動同期 |
| Direct Access | ベクトルに対する手動 CRUD 操作 |
| 種類 | レイテンシ | 容量 | コスト | 最適なユースケース |
|---|---|---|---|---|
| Standard | 20〜50ms | 3.2億ベクトル(768次元) | 高め | リアルタイム・低レイテンシ |
| Storage-Optimized | 300〜500ms | 10億以上のベクトル(768次元) | 7分の1 | 大規模・コスト重視 |
| 種類 | Embedding | 同期方式 | ユースケース |
|---|---|---|---|
| Delta Sync(マネージド) | Databricks が生成 | Delta から自動同期 | 最も簡単なセットアップ |
| Delta Sync(セルフマネージド) | ユーザーが提供 | Delta から自動同期 | カスタム Embedding |
| Direct Access | ユーザーが提供 | 手動 CRUD | リアルタイム更新 |
from databricks.sdk import WorkspaceClient
w = WorkspaceClient()
# 標準エンドポイントを作成
endpoint = w.vector_search_endpoints.create_endpoint(
name="my-vs-endpoint",
endpoint_type="STANDARD" # または "STORAGE_OPTIMIZED"
)
# 注意: エンドポイント作成は非同期です。get_endpoint() でステータスを確認してください。
# ソーステーブルには主キー列とテキスト列が必要です
index = w.vector_search_indexes.create_index(
name="catalog.schema.my_index",
endpoint_name="my-vs-endpoint",
primary_key="id",
index_type="DELTA_SYNC",
delta_sync_index_spec={
"source_table": "catalog.schema.documents",
"embedding_source_columns": [
{
"name": "content", # Embedding 対象のテキスト列
"embedding_model_endpoint_name": "databricks-gte-large-en"
}
],
"pipeline_type": "TRIGGERED" # または "CONTINUOUS"
}
)
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "content", "metadata"],
query_text="What is machine learning?",
num_results=5
)
for doc in results.result.data_array:
score = doc[-1] # 類似スコアは最後の列
print(f"Score: {score}, Content: {doc[1][:100]}...")
# 大規模かつコスト効率を重視したデプロイ向け
endpoint = w.vector_search_endpoints.create_endpoint(
name="my-storage-endpoint",
endpoint_type="STORAGE_OPTIMIZED"
)
# ソーステーブルには主キー列と Embedding ベクトル列が必要です
index = w.vector_search_indexes.create_index(
name="catalog.schema.my_index",
endpoint_name="my-vs-endpoint",
primary_key="id",
index_type="DELTA_SYNC",
delta_sync_index_spec={
"source_table": "catalog.schema.documents",
"embedding_vector_columns": [
{
"name": "embedding", # 事前計算済みの Embedding 列
"embedding_dimension": 768
}
],
"pipeline_type": "TRIGGERED"
}
)
import json
# 手動 CRUD 用インデックスを作成
index = w.vector_search_indexes.create_index(
name="catalog.schema.direct_index",
endpoint_name="my-vs-endpoint",
primary_key="id",
index_type="DIRECT_ACCESS",
direct_access_index_spec={
"embedding_vector_columns": [
{"name": "embedding", "embedding_dimension": 768}
],
"schema_json": json.dumps({
"id": "string",
"text": "string",
"embedding": "array<float>",
"metadata": "string"
})
}
)
# データの Upsert
w.vector_search_indexes.upsert_data_vector_index(
index_name="catalog.schema.direct_index",
inputs_json=json.dumps([
{"id": "1", "text": "Hello", "embedding": [0.1, 0.2, ...], "metadata": "doc1"},
{"id": "2", "text": "World", "embedding": [0.3, 0.4, ...], "metadata": "doc2"},
])
)
# データの削除
w.vector_search_indexes.delete_data_vector_index(
index_name="catalog.schema.direct_index",
primary_keys=["1", "2"]
)
# クエリ用 Embedding が事前計算済みの場合
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "text"],
query_vector=[0.1, 0.2, 0.3, ...], # 768次元のベクトル
num_results=10
)
ハイブリッド検索は、ベクトル類似(ANN)と BM25 キーワードスコアリングを組み合わせます。 SKU・エラーコード・固有名詞・技術用語など、純粋なセマンティック検索ではキーワード固有の結果を取りこぼす可能性がある場合に、クエリに含まれる特定の語句を確実にマッチさせたいときに使用します。 ANN とハイブリッド検索の選択基準については references/search-modes.md を参照してください。
# ベクトル類似とキーワードマッチングを組み合わせる
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "content"],
query_text="SPARK-12345 executor memory error",
query_type="HYBRID",
num_results=10
)
# filters_json は辞書形式を使用
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "content"],
query_text="machine learning",
num_results=10,
filters_json='{"category": "ai", "status": ["active", "pending"]}'
)
Storage-Optimized エンドポイントでは、databricks-vectorsearch パッケージの filters パラメーター(文字列を受け付け)を通じて SQL ライクなフィルター構文を使用します。
from databricks.vector_search.client import VectorSearchClient
vsc = VectorSearchClient()
index = vsc.get_index(endpoint_name="my-storage-endpoint", index_name="catalog.schema.my_index")
# Storage-Optimized エンドポイント向けの SQL ライクフィルター構文
results = index.similarity_search(
query_text="machine learning",
columns=["id", "content"],
num_results=10,
filters="category = 'ai' AND status IN ('active', 'pending')"
)
# フィルター記述例
# filters="price > 100 AND price < 500"
# filters="department LIKE 'eng%'"
# filters="created_at >= '2024-01-01'"
# TRIGGERED パイプラインタイプの場合、手動で同期を実行
w.vector_search_indexes.sync_index(
index_name="catalog.schema.my_index"
)
# 全ベクトルを取得(デバッグ・エクスポート用)
scan_result = w.vector_search_indexes.scan_index(
index_name="catalog.schema.my_index",
num_results=100
)
| トピック | ファイル | 説明 |
|---|---|---|
| インデックスの種類 | references/index-types.md | Delta Sync(マネージド/セルフマネージド)と Direct Access の詳細比較 |
| エンドツーエンド RAG | references/end-to-end-rag.md | ソーステーブル → エンドポイント → インデックス → クエリ → Agent 連携の完全ガイド |
| 検索モード | references/search-modes.md | セマンティック(ANN)とハイブリッド検索の使い分け・判断ガイド |
| 運用管理 | references/troubleshooting-and-operations.md | モニタリング・コスト最適化・容量計画・移行 |
# エンドポイントの一覧表示
databricks vector-search-endpoints list-endpoints
# エンドポイントの作成(位置引数: NAME ENDPOINT_TYPE)
databricks vector-search-endpoints create-endpoint my-endpoint STANDARD
# エンドポイント上のインデックス一覧表示(位置引数: ENDPOINT_NAME)
databricks vector-search-indexes list-indexes my-endpoint
# インデックスのステータス確認(位置引数: INDEX_NAME)
databricks vector-search-indexes get-index catalog.schema.my_index
# インデックスの同期(位置引数: INDEX_NAME)
databricks vector-search-indexes sync-index catalog.schema.my_index
# インデックスの削除(位置引数: INDEX_NAME)
databricks vector-search-indexes delete-index catalog.schema.my_index
| 問題 | 解決策 |
|---|---|
| インデックス同期が遅い | Storage-Optimized エンドポイントを使用(インデックス作成が20倍高速) |
| クエリレイテンシが高い | 100ms 未満が必要な場合は Standard エンドポイントを使用 |
| filters_json が機能しない | Storage-Optimized では databricks-vectorsearch パッケージの filters パラメーターを通じた SQL ライク文字列フィルターを使用 |
| Embedding の次元数が一致しない | クエリとインデックスの次元数が同じであることを確認 |
| インデックスが更新されない | pipeline_type を確認し、TRIGGERED の場合は sync_index() を使用 |
| 容量不足 | Storage-Optimized(10億以上のベクトル対応)にアップグレード |
query_vector が切り詰められる |
大きなベクトル(例: 1024次元)は JSON シリアライズ時に切り詰められることがあります。マネージド Embedding インデックスでは query_text を使用するか、Databricks SDK を使ってベクトルを直接渡してください |
Databricks は以下の組み込み Embedding モデルを提供しています。
| モデル | 次元数 | コンテキストウィンドウ | ユースケース |
|---|---|---|---|
databricks-gte-large-en |
1024 | 8192 トークン | 英語テキスト・高品質 |
databricks-bge-large-en |
1024 | 512 トークン | 英語テキスト・汎用 |
# マネージド Embedding で使用
embedding_source_columns=[
{
"name": "content",
"embedding_model_endpoint_name": "databricks-gte-large-en"
}
]
FIRST: Use the parent databricks-core skill for CLI basics, authentication, and profile selection.
Patterns for creating, managing, and querying vector search indexes for RAG and semantic search applications.
Use this skill when:
Databricks Vector Search provides managed vector similarity search with automatic embedding generation and Delta Lake integration.
| Component | Description |
|---|---|
| Endpoint | Compute resource hosting indexes (Standard or Storage-Optimized) |
| Index | Vector data structure for similarity search |
| Delta Sync | Auto-syncs with source Delta table |
| Direct Access | Manual CRUD operations on vectors |
| Type | Latency | Capacity | Cost | Best For |
|---|---|---|---|---|
| Standard | 20-50ms | 320M vectors (768 dim) | Higher | Real-time, low-latency |
| Storage-Optimized | 300-500ms | 1B+ vectors (768 dim) | 7x lower | Large-scale, cost-sensitive |
| Type | Embeddings | Sync | Use Case |
|---|---|---|---|
| Delta Sync (managed) | Databricks computes | Auto from Delta | Easiest setup |
| Delta Sync (self-managed) | You provide | Auto from Delta | Custom embeddings |
| Direct Access | You provide | Manual CRUD | Real-time updates |
from databricks.sdk import WorkspaceClient
w = WorkspaceClient()
# Create a standard endpoint
endpoint = w.vector_search_endpoints.create_endpoint(
name="my-vs-endpoint",
endpoint_type="STANDARD" # or "STORAGE_OPTIMIZED"
)
# Note: Endpoint creation is asynchronous; check status with get_endpoint()
# Source table must have: primary key column + text column
index = w.vector_search_indexes.create_index(
name="catalog.schema.my_index",
endpoint_name="my-vs-endpoint",
primary_key="id",
index_type="DELTA_SYNC",
delta_sync_index_spec={
"source_table": "catalog.schema.documents",
"embedding_source_columns": [
{
"name": "content", # Text column to embed
"embedding_model_endpoint_name": "databricks-gte-large-en"
}
],
"pipeline_type": "TRIGGERED" # or "CONTINUOUS"
}
)
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "content", "metadata"],
query_text="What is machine learning?",
num_results=5
)
for doc in results.result.data_array:
score = doc[-1] # Similarity score is last column
print(f"Score: {score}, Content: {doc[1][:100]}...")
# For large-scale, cost-effective deployments
endpoint = w.vector_search_endpoints.create_endpoint(
name="my-storage-endpoint",
endpoint_type="STORAGE_OPTIMIZED"
)
# Source table must have: primary key + embedding vector column
index = w.vector_search_indexes.create_index(
name="catalog.schema.my_index",
endpoint_name="my-vs-endpoint",
primary_key="id",
index_type="DELTA_SYNC",
delta_sync_index_spec={
"source_table": "catalog.schema.documents",
"embedding_vector_columns": [
{
"name": "embedding", # Pre-computed embedding column
"embedding_dimension": 768
}
],
"pipeline_type": "TRIGGERED"
}
)
import json
# Create index for manual CRUD
index = w.vector_search_indexes.create_index(
name="catalog.schema.direct_index",
endpoint_name="my-vs-endpoint",
primary_key="id",
index_type="DIRECT_ACCESS",
direct_access_index_spec={
"embedding_vector_columns": [
{"name": "embedding", "embedding_dimension": 768}
],
"schema_json": json.dumps({
"id": "string",
"text": "string",
"embedding": "array<float>",
"metadata": "string"
})
}
)
# Upsert data
w.vector_search_indexes.upsert_data_vector_index(
index_name="catalog.schema.direct_index",
inputs_json=json.dumps([
{"id": "1", "text": "Hello", "embedding": [0.1, 0.2, ...], "metadata": "doc1"},
{"id": "2", "text": "World", "embedding": [0.3, 0.4, ...], "metadata": "doc2"},
])
)
# Delete data
w.vector_search_indexes.delete_data_vector_index(
index_name="catalog.schema.direct_index",
primary_keys=["1", "2"]
)
# When you have pre-computed query embedding
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "text"],
query_vector=[0.1, 0.2, 0.3, ...], # Your 768-dim vector
num_results=10
)
Hybrid search combines vector similarity (ANN) with BM25 keyword scoring. Use it when queries contain exact terms that must match — SKUs, error codes, proper nouns, or technical terminology — where pure semantic search might miss keyword-specific results. See references/search-modes.md for detailed guidance on choosing between ANN and hybrid search.
# Combines vector similarity with keyword matching
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "content"],
query_text="SPARK-12345 executor memory error",
query_type="HYBRID",
num_results=10
)
# filters_json uses dictionary format
results = w.vector_search_indexes.query_index(
index_name="catalog.schema.my_index",
columns=["id", "content"],
query_text="machine learning",
num_results=10,
filters_json='{"category": "ai", "status": ["active", "pending"]}'
)
Storage-Optimized endpoints use SQL-like filter syntax via the databricks-vectorsearch package's filters parameter (accepts a string):
from databricks.vector_search.client import VectorSearchClient
vsc = VectorSearchClient()
index = vsc.get_index(endpoint_name="my-storage-endpoint", index_name="catalog.schema.my_index")
# SQL-like filter syntax for storage-optimized endpoints
results = index.similarity_search(
query_text="machine learning",
columns=["id", "content"],
num_results=10,
filters="category = 'ai' AND status IN ('active', 'pending')"
)
# More filter examples
# filters="price > 100 AND price < 500"
# filters="department LIKE 'eng%'"
# filters="created_at >= '2024-01-01'"
# For TRIGGERED pipeline type, manually sync
w.vector_search_indexes.sync_index(
index_name="catalog.schema.my_index"
)
# Retrieve all vectors (for debugging/export)
scan_result = w.vector_search_indexes.scan_index(
index_name="catalog.schema.my_index",
num_results=100
)
| Topic | File | Description |
|---|---|---|
| Index Types | references/index-types.md | Detailed comparison of Delta Sync (managed/self-managed) vs Direct Access |
| End-to-End RAG | references/end-to-end-rag.md | Complete walkthrough: source table → endpoint → index → query → agent integration |
| Search Modes | references/search-modes.md | When to use semantic (ANN) vs hybrid search, decision guide |
| Operations | references/troubleshooting-and-operations.md | Monitoring, cost optimization, capacity planning, migration |
# List endpoints
databricks vector-search-endpoints list-endpoints
# Create endpoint (positional args: NAME ENDPOINT_TYPE)
databricks vector-search-endpoints create-endpoint my-endpoint STANDARD
# List indexes on endpoint (positional arg: ENDPOINT_NAME)
databricks vector-search-indexes list-indexes my-endpoint
# Get index status (positional arg: INDEX_NAME)
databricks vector-search-indexes get-index catalog.schema.my_index
# Sync index (positional arg: INDEX_NAME)
databricks vector-search-indexes sync-index catalog.schema.my_index
# Delete index (positional arg: INDEX_NAME)
databricks vector-search-indexes delete-index catalog.schema.my_index
| Issue | Solution |
|---|---|
| Index sync slow | Use Storage-Optimized endpoints (20x faster indexing) |
| Query latency high | Use Standard endpoint for <100ms latency |
| filters_json not working | Storage-Optimized uses SQL-like string filters via databricks-vectorsearch package's filters parameter |
| Embedding dimension mismatch | Ensure query and index dimensions match |
| Index not updating | Check pipeline_type; use sync_index() for TRIGGERED |
| Out of capacity | Upgrade to Storage-Optimized (1B+ vectors) |
query_vector truncated |
Large vectors (e.g. 1024-dim) can be truncated when serialized as JSON. Use query_text instead (for managed embedding indexes), or use the Databricks SDK to pass raw vectors |
Databricks provides built-in embedding models:
| Model | Dimensions | Context Window | Use Case |
|---|---|---|---|
databricks-gte-large-en |
1024 | 8192 tokens | English text, high quality |
databricks-bge-large-en |
1024 | 512 tokens | English text, general purpose |
# Use with managed embeddings
embedding_source_columns=[
{
"name": "content",
"embedding_model_endpoint_name": "databricks-gte-large-en"
}
]
columns_to_sync matters — only synced columns are available in query results; include all columns you needdatabricks-vectorsearch package's filters parameter which accepts both formatsVectorSearchRetrieverTool原文・著作権は Anthropic および各プラグイン作者に帰属します。日本語訳は Claude API による自動翻訳です。