Query foundation and embedding models
Databricks provides several ways to query foundation models. Choose an interface based on whether you need a provider-independent API, provider-specific features, or batch inference.
Ways to query a model
Approach | When to use it |
|---|---|
Use interfaces compatible with OpenResponses and OpenAI Chat Completions. Databricks translates each request to the downstream model's native format, so you can switch models across providers without changing client code. | |
Use provider-specific features or existing OpenAI, Anthropic, or Google SDK code. | |
Run batch inference from SQL or Python. |
Quickstart
Query a model service in two steps:
Step 1: Pick a model
Databricks provides models out of the box, such as claude-sonnet-4-5 or gpt-5-6-sol. Databricks-provided models are registered in Unity Catalog under system.ai.
Step 2: Send a request using the unified OpenAI-compatible API
Use the MLflow Chat Completions API with the OpenAI Python SDK:
from openai import OpenAI
import os
DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')
client = OpenAI(
api_key=DATABRICKS_TOKEN, # your personal access token
base_url="https://<workspace-url>/ai-gateway/mlflow/v1" # your Databricks workspace instance
)
chat_completion = client.chat.completions.create(
messages=[
{"role": "user", "content": "What is Databricks?"},
],
model="system.ai.claude-sonnet-4-5",
max_tokens=256
)
print(chat_completion.choices[0].message.content)
Requirements
- A Databricks workspace in a Unity Gateway supported region.
- Unity Catalog enabled for your workspace. See Enable a workspace for Unity Catalog.
- Workspace entitlement to query: Workspace access, or Consumer access with the Consumer access to Unity Gateway preview enabled for your account (Public Preview). See Manage entitlements and Manage Databricks previews.
Query model services with unified APIs
Unified APIs offer an OpenAI-compatible interface to query models on Databricks. Use unified APIs to seamlessly switch between models from different providers without changing your code.
Unified Responses API
Unified Responses API
The Unified Responses API (/mlflow/v1/responses) is an OpenResponses-compatible, provider-agnostic API for querying models on Databricks. Databricks recommends it over the MLflow Chat Completions API. Pick the best model for your use case across providers, without changing your code.
- Python
- REST API
from openai import OpenAI
import os
DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')
client = OpenAI(
api_key=DATABRICKS_TOKEN,
base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)
response = client.responses.create(
model="<model-service>",
input=[{"role": "user", "content": "What is Databricks?"}]
)
print(response.output_text)
curl \
-u token:$DATABRICKS_TOKEN \
-X POST \
-H "Content-Type: application/json" \
-d '{
"model": "<model-service>",
"input": [
{"role": "user", "content": "What is Databricks?"}
]
}' \
https://<workspace-url>/ai-gateway/mlflow/v1/responses
Replace <workspace-url> with your Databricks workspace URL and <model-service> with the fully qualified name of your model service.
MLflow Chat Completions API
MLflow Chat Completions API
- Python
- REST API
from openai import OpenAI
import os
DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')
client = OpenAI(
api_key=DATABRICKS_TOKEN,
base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)
chat_completion = client.chat.completions.create(
messages=[
{"role": "user", "content": "Hello!"},
{"role": "assistant", "content": "Hello! How can I assist you today?"},
{"role": "user", "content": "What is Databricks?"},
],
model="system.ai.gpt-5-6-sol",
max_tokens=256
)
print(chat_completion.choices[0].message.content)
curl \
-u token:$DATABRICKS_TOKEN \
-X POST \
-H "Content-Type: application/json" \
-d '{
"model": "system.ai.gpt-5-6-sol",
"max_tokens": 256,
"messages": [
{"role": "user", "content": "Hello!"},
{"role": "assistant", "content": "Hello! How can I assist you today?"},
{"role": "user", "content": "What is Databricks?"}
]
}' \
https://<workspace-url>/ai-gateway/mlflow/v1/chat/completions
Replace <workspace-url> with your Databricks workspace URL.
MLflow Embeddings API
MLflow Embeddings API
- Python
- REST API
from openai import OpenAI
import os
DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')
client = OpenAI(
api_key=DATABRICKS_TOKEN,
base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)
embeddings = client.embeddings.create(
input="What is Databricks?",
model="<model-service>"
)
print(embeddings.data[0].embedding)
curl \
-u token:$DATABRICKS_TOKEN \
-X POST \
-H "Content-Type: application/json" \
-d '{
"model": "<model-service>",
"input": "What is Databricks?"
}' \
https://<workspace-url>/ai-gateway/mlflow/v1/embeddings
Replace <workspace-url> with your Databricks workspace URL and <model-service> with the fully qualified name of your model service.
Query model services with ai_query
You can use the ai_query function to query model services directly from SQL or Python. This allows you to capture usage tracking information for your batch inference workloads.
To query a model service with ai_query, run ai_query against a model service:
SELECT ai_query(
'system.ai.claude-sonnet-4-5',
'Summarize the following text: ' || text_column
) AS summary
FROM my_table
LIMIT 10
The usage tracking system table (system.ai_gateway.usage) captures requests made through ai_query to model services. These requests also appear in the built-in usage dashboard.
For full ai_query syntax and parameter reference, see ai_query function. For best practices and supported models, see Use ai_query.
Limitations
ai_querysupport for Unity Gateway is only available for Databricks-provided models. Pass thesystem.aimodel service name (for example,system.ai.claude-sonnet-4-5orsystem.ai.gpt-5-6-sol). Model services that you create in Unity Gateway are not yet supported.- Only usage tracking applies to
ai_querybatch inference workloads. Other Unity Gateway features such as rate limits, guardrails implemented with service policies, inference tables, and fallbacks do not apply.
Query model services with native APIs
Native APIs offer provider-specific interfaces to query models on Databricks. Use native APIs to access the latest provider-specific features.
Each native API works only with model services whose underlying model uses the matching API format:
- Use the OpenAI Responses API to query model services backed by OpenAI (GPT) models.
- Use the Anthropic Messages API to query model services backed by Claude models.
- Use the Google Gemini API to query model services backed by Gemini models.
To query a model service regardless of its underlying model, use the unified APIs instead.