Skip to main content

Query foundation and embedding models

Databricks provides several ways to query foundation models. Choose an interface based on whether you need a provider-independent API, provider-specific features, or batch inference.

Ways to query a model​

Approach

When to use it

Unified APIs

Use interfaces compatible with OpenResponses and OpenAI Chat Completions. Databricks translates each request to the downstream model's native format, so you can switch models across providers without changing client code.

Provider-native APIs

Use provider-specific features or existing OpenAI, Anthropic, or Google SDK code.

ai_query

Run batch inference from SQL or Python.

Approach

When to use it

Unified APIs

Use interfaces compatible with OpenResponses and OpenAI Chat Completions. Databricks translates each request to the downstream model's native format, so you can switch models across providers without changing client code.

Provider-native APIs

Use provider-specific features or existing OpenAI, Anthropic, or Google SDK code.

ai_query

Run batch inference from SQL or Python.

Quickstart​

Query a model service in two steps:

Step 1: Pick a model​

Databricks provides models out of the box, such as claude-sonnet-4-5 or gpt-5-6-sol. Databricks-provided models are registered in Unity Catalog under system.ai.

Step 2: Send a request using the unified OpenAI-compatible API​

Use the MLflow Chat Completions API with the OpenAI Python SDK:

Python
from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
api_key=DATABRICKS_TOKEN, # your personal access token
base_url="https://<workspace-url>/ai-gateway/mlflow/v1" # your Databricks workspace instance
)

chat_completion = client.chat.completions.create(
messages=[
{"role": "user", "content": "What is Databricks?"},
],
model="system.ai.claude-sonnet-4-5",
max_tokens=256
)

print(chat_completion.choices[0].message.content)

Requirements​

Query model services with unified APIs​

Unified APIs offer an OpenAI-compatible interface to query models on Databricks. Use unified APIs to seamlessly switch between models from different providers without changing your code.

Unified Responses API

Unified Responses API​

The Unified Responses API (/mlflow/v1/responses) is an OpenResponses-compatible, provider-agnostic API for querying models on Databricks. Databricks recommends it over the MLflow Chat Completions API. Pick the best model for your use case across providers, without changing your code.

Python
from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
api_key=DATABRICKS_TOKEN,
base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)

response = client.responses.create(
model="<model-service>",
input=[{"role": "user", "content": "What is Databricks?"}]
)

print(response.output_text)

Replace <workspace-url> with your Databricks workspace URL and <model-service> with the fully qualified name of your model service.

MLflow Chat Completions API

MLflow Chat Completions API​

Python
from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
api_key=DATABRICKS_TOKEN,
base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)

chat_completion = client.chat.completions.create(
messages=[
{"role": "user", "content": "Hello!"},
{"role": "assistant", "content": "Hello! How can I assist you today?"},
{"role": "user", "content": "What is Databricks?"},
],
model="system.ai.gpt-5-6-sol",
max_tokens=256
)

print(chat_completion.choices[0].message.content)

Replace <workspace-url> with your Databricks workspace URL.

MLflow Embeddings API

MLflow Embeddings API​

Python
from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
api_key=DATABRICKS_TOKEN,
base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)

embeddings = client.embeddings.create(
input="What is Databricks?",
model="<model-service>"
)

print(embeddings.data[0].embedding)

Replace <workspace-url> with your Databricks workspace URL and <model-service> with the fully qualified name of your model service.

Query model services with ai_query​

You can use the ai_query function to query model services directly from SQL or Python. This allows you to capture usage tracking information for your batch inference workloads.

To query a model service with ai_query, run ai_query against a model service:

SQL
SELECT ai_query(
'system.ai.claude-sonnet-4-5',
'Summarize the following text: ' || text_column
) AS summary
FROM my_table
LIMIT 10

The usage tracking system table (system.ai_gateway.usage) captures requests made through ai_query to model services. These requests also appear in the built-in usage dashboard.

For full ai_query syntax and parameter reference, see ai_query function. For best practices and supported models, see Use ai_query.

Limitations​

  • ai_query support for Unity Gateway is only available for Databricks-provided models. Pass the system.ai model service name (for example, system.ai.claude-sonnet-4-5 or system.ai.gpt-5-6-sol). Model services that you create in Unity Gateway are not yet supported.
  • Only usage tracking applies to ai_query batch inference workloads. Other Unity Gateway features such as rate limits, guardrails implemented with service policies, inference tables, and fallbacks do not apply.

Query model services with native APIs​

Native APIs offer provider-specific interfaces to query models on Databricks. Use native APIs to access the latest provider-specific features.

Each native API works only with model services whose underlying model uses the matching API format:

To query a model service regardless of its underlying model, use the unified APIs instead.

Next steps​