Skip to main content
Harbor runs AI agents against tasks in sandboxed environments. Its Phoenix plugin records those jobs as versioned datasets and experiments. You can compare agents, models, and repetitions in Phoenix, then open the trace behind an individual score. Harbor remains responsible for running agents and verifiers. Phoenix stores and displays the resulting tasks, runs, rewards, errors, and traces. The plugin does not rerun tasks or calculate a replacement reward.

When to use the plugin

Use the plugin when Harbor runs your benchmark and you want to:
  • compare agents or models over the same task set;
  • track results across repeated benchmark jobs;
  • separate behavioral scores from infrastructure failures;
  • inspect an Agent Trajectory Interchange Format (ATIF) trace for a scored run; or
  • keep completed results when a long job stops early.
Omit the plugin when you want a Harbor-only job. Selecting the plugin makes Phoenix recording part of the job contract. A setup or result-write failure stops the job instead of continuing with unrecorded trials.
The integration requires Python 3.12 or newer, Harbor 0.21.0 or newer, and Phoenix server 15.0 or newer.

Install and run

Install the Phoenix client and Harbor in the same Python environment:
Set the connection to your Phoenix instance. Self-hosted Phoenix uses http://localhost:6006 by default.
You can omit PHOENIX_API_KEY when your instance does not require authentication. See What is my Phoenix endpoint? for hosted and self-hosted endpoint formats. Add the plugin to a Harbor job:

What you’ll see in Phoenix

At job start, the plugin creates or reuses a versioned dataset for the resolved task set. Each Harbor task becomes a dataset example, and each agent and model configuration gets its own experiment. As each final logical trial finishes, Phoenix records:
  • an experiment run linked to the task’s dataset example;
  • the agent’s final textual response when its terminal ATIF trajectory contains one;
  • Harbor’s verifier rewards as experiment evaluations;
  • an infra_ok evaluation for execution health;
  • a run error when Harbor recorded an exception; and
  • a link to the ATIF trace when tracing succeeds.
Start with the default atif mode when your Harbor agent writes ATIF trajectories. It captures agent execution without adding tracing code or giving the sandbox network access to Phoenix.

How Harbor data maps to Phoenix

Each Harbor task becomes one dataset example, including a multi-step task. For a multi-step task, the example input also contains the ordered step names and instructions. Reference outputs are optional. Without an explicit reference file, the dataset example’s output stays empty. When the terminal ATIF trajectory ends with a textual agent turn, the run output uses Phoenix’s chat-message format and the experiment comparison renders the response as Markdown. For structured ATIF messages, the plugin joins text parts in order and omits media parts. It leaves the output empty when the final turn is a tool call, contains only media, or is otherwise not a user-facing text response. This is expected for tasks whose result is an environment change rather than a written answer. Run output extraction is independent of trace upload. trace_mode=null still records a saved terminal ATIF response without creating a trace. Missing or invalid trajectories leave the output empty. Multi-step tasks use the last attempted step, continued trajectories use their terminal continuation, and Harbor retries contribute only the terminal physical attempt. Successful experiment runs are immutable. Runs created by earlier plugin versions keep their legacy Harbor metadata output when a job resumes; the plugin recognizes and reuses them instead of treating the changed output format as a conflict.

Optional reference outputs

To display an expected answer beside an experiment result, add this setting to the task’s task.toml:
The file must contain either a JSON string or a JSON object. A string becomes an assistant message in the example’s output:
An object is stored unchanged. Use it for structured expected results, chat messages, or existing verifier references:
Wrap top-level arrays, numbers, booleans, or null in an object, such as {"count": 117}. The plugin does not interpret object keys or use the reference to calculate rewards. Paths resolve relative to the downloaded task root on the host running Harbor, not the agent sandbox or the job’s working directory. Absolute paths, .. components, and symlinks that resolve outside the task root are rejected. Keep references under tests/ beside verifier assets; do not copy them into the agent’s workspace. The plugin reads files during setup and never runs a solution to obtain a reference. The same setting works for a task selected from a local dataset or passed directly:
It also travels with tasks in registry, package, and repository datasets. Each task opts in independently, so jobs can mix tasks with and without references. No job-wide Phoenix option is needed. For a multi-step task, put the setting in the root task.toml as above. There is one reference output for the whole task. It should describe the expected final result. If useful, an object can describe expected intermediate results as well:
The plugin stores this object as one example output; it does not split it into step examples or match its keys to step names. Omitting the setting leaves the reference blank, even if a file named expected.json exists. A configured file that is missing, unreadable, invalid JSON, or has an unsupported top-level type stops setup before trials begin. Errors identify the task and configured path. Changing reference content updates the dataset example and creates a new dataset version; existing experiments retain their original version.

Dataset versions

The plugin uses the Harbor task ID as the stable example ID. It synchronizes the full task set each time a job starts.
  • An unchanged task set reuses the current dataset version.
  • Adding, removing, or changing a task creates a new dataset version.
  • An experiment stays pinned to the dataset version used when the experiment was created.
The plugin infers a Phoenix dataset name for each supported single-source job. The inferred name depends on the Harbor task source: Provide dataset=<name> only when a job contains several direct tasks, which have no shared collection name, or when you want to customize the dataset’s display name in Phoenix. Add this setting to the job’s existing command:
Use one task collection per job. The plugin rejects jobs that mix a configured dataset with direct tasks or include several configured datasets.

Read scores correctly

Harbor tasks can use different verifiers, so the plugin keeps summary metrics separate from task-specific diagnostics. Phoenix does not run a second evaluator. The plugin stores the rewards returned by Harbor as experiment evaluations on each run. A multi-step run can have verifier rewards and an exception at the same time. Phoenix keeps both: the reward remains available, while the run has an error and infra_ok=0. Trial-level evaluations for a multi-step task include the resolved multi_step_reward_strategy in their metadata. Harbor uses mean when the task does not set a strategy; an explicit final value remains final. Step evaluations and infra_ok do not carry this metadata. For comparisons, check reward coverage before calculating an aggregate. Then use infra_ok to separate agent behavior from broken environments, timeouts, or verifier failures. Step-level scores show where a multi-step task failed.

Understand ATIF traces

ATIF is the default trace mode. The plugin reads saved trajectories after the final trial attempt, converts them to OpenInference spans, and uploads them to the experiment’s Phoenix project. The sandbox does not need a Phoenix endpoint or Phoenix credentials. One Harbor trial becomes one trace and one Phoenix session. A multi-step trial adds a span for each attempted step:
Single-step trajectories attach directly to the harbor.trial root. Each multi-step harbor.step span records the step instruction, timing, exception status, and any verifier rewards. Its trajectories appear beneath it. This keeps an attempted step visible even when Harbor did not save a trajectory for that step. Agent, model, and tool spans use their names from ATIF. Fresh agent operations use iteration N; context-management operations use compaction N; and other operational system steps use system event N. An agent step with llm_call_count: 0 has no LLM span, but it still keeps its operation and tool spans. Continuation roots use <agent> (continuation N). Referenced subagents attach to the matching tool call when source_call_id proves that relationship, or to the referencing operation when it does not. The converter supports ATIF v1.0 through v1.7. It reconstructs LLM inputs from ATIF messages and marks them with metadata.atif.input_source = "reconstructed"; it does not parse provider-native message formats. User and system prompts and copied context contribute to those inputs without creating duplicate execution spans. An observation becomes a tool result only when its source_call_id matches the call. Multiple results for one call remain in order. Unmatched step observations stay on the operation span, while unassigned feedback remains structured in the reconstructed input without an invented message role or tool association. Structured text and image parts remain in serialized messages, but the plugin does not read or upload media bytes. ATIF v1.8 audio fields are not supported. Only LLM spans carry llm.* attributes. The converter keeps trajectory-level final_metrics on the agent root so Phoenix does not count the same tokens twice. It maps producer-specific cache-write and reasoning token counts when they are present. ATIF timestamps describe events rather than complete operation durations. The plugin uses request timings only when it can map every measurement to one LLM step. It leaves ambiguous LLM and tool durations at zero instead of inventing timing or concurrency. Trace discovery and conversion are best-effort. If the agent does not save a valid trajectory, the plugin logs a warning and records the run and evaluations without a trace. A successful Phoenix run is immutable, so replay cannot add a missing trace link later. Use trace_mode=null when the agent has no ATIF output or when you do not want traces:
Live OpenTelemetry Protocol (OTLP) support is deferred to a follow-up. This release accepts atif or null, and does not link live OpenTelemetry traces from Harbor agents to experiment runs.

Name experiments

The default experiment name is:
For a job with one agent configuration, set an exact display name:
For a job with several agent configurations, use a template:
Available fields are {job.name}, {job.id}, {dataset.name}, {agent.name}, {agent.model}, and {agent.short_digest}. Agent names do not need to be unique. Two agents with the same name but different effective configurations each get an experiment. If their templates render the same experiment name, the plugin appends the short agent configuration digest to distinguish them. Experiment display names do not define identity. The plugin identifies an experiment by the Harbor job ID and the effective agent configuration. Use a new Harbor job for a new benchmark execution, even when you want to reuse the same display name.

Configure the plugin

Pass settings with Harbor’s --plugin-kwarg option.

Resume and failure behavior

The plugin writes each trial when it reaches its final state. This gives you live progress and preserves completed runs when the job stops. On resume or replay, the plugin recovers the matching experiment, reuses matching successful runs, retries failed runs, and upserts their evaluations. If another Harbor job created a newer version of the shared dataset, the recovered experiment remains pinned to its original version. Deterministic task, run, and trace identities prevent duplicate records during sequential ingestion. Run only one process for a given Harbor job. Experiment recovery is not atomic across multiple ingesters. The plugin handles failures as follows:
  • Phoenix setup failures stop the job before Harbor spends trial compute.
  • Run or evaluation write failures stop the job. Records from completed trials remain in Phoenix and Harbor keeps its terminal results for resume.
  • Missing or invalid ATIF data does not stop the job. The run remains available without a trace.

Current limits

The plugin does not support:
  • Harbor regrade jobs, which run a new verifier against recorded agent work;
  • post-hoc import of a finished job;
  • live OTLP trace linkage, which is deferred to a follow-up;
  • several configured datasets in one job;
  • a mixture of configured datasets and direct tasks; or
  • concurrent ingestion of the same Harbor job.
For Harbor task, dataset, agent, and job configuration, see the Harbor documentation.

Give a coding agent Harbor context

Install the phoenix-harbor skill when a coding agent will configure or interpret the integration:
For example:
See Coding agents for supported agents and installation options.