Skip to content

[experiments] Metric charts on the experiment compare page #16780

Description

@mikeldking

Part of #11637.

Add metric charts to the experiment compare page's metrics view (?view=metrics, ExperimentCompareMetricsPage), styled like the charts on the experiments page:

  • Per-evaluator charts: one chart per evaluator present on the selected experiments. Numeric evaluators show the mean score; label evaluators show the label distribution
  • Latency: average run latency per experiment
  • Errors: share of runs that errored per experiment
  • Token breakdown: prompt tokens and completion tokens per experiment, stacked by token type

Context

The experiments page charts (pages/dataset/metrics/chartCatalog.tsx) get their data from useExperimentMetricsData(datasetId). That hook always loads the dataset's last EXPERIMENT_METRICS_EXPERIMENT_COUNT experiments plus the dataset baseline. On the compare page, the charts need to show only the selected experiments and use the base experiment as the reference.

Scope

Data

  • Load chart data for the selected experiment IDs (base + compare). Dataset.experiments(filterIds:, includeEphemeral: true) already supports this; the compare metrics query uses it.
  • Let the chart views accept their experiments + baseline data from either source (dataset recent experiments or compare selection) without forking each chart.
  • Get evaluator names from the selected experiments, not from the whole dataset.
  • Charts update when compared experiments are added or removed, without a full reload.

Presentation

  • On the x-axis, order experiments base first, then compare experiments in selection order. Color each one with useExperimentColors so the charts match the rest of the compare page.
  • Draw the base experiment as the reference, using the same treatment as the dataset baseline (ExperimentBaselineReference).
  • Latency: mirror ExperimentLatencyChart.
  • Errors: mirror ExperimentErrorRateChart. Hovering shows error rate and run count.
  • Token breakdown: mirror ExperimentPromptTokenDetailsChart / ExperimentCompletionTokenDetailsChart, using data from costDetailSummaryEntries. Use the labels and colors from tokenDetailUtils, and key details off the raw token type.
  • Evaluators: reuse ExperimentAnnotationMetricPanel. Defer offscreen charts (DeferredChartPanel) so many evaluators don't all fetch up front.

Acceptance

  • Metrics view shows latency, error rate, prompt token, and completion token charts for exactly the selected experiments
  • One chart per evaluator present on the selected experiments; numeric and label evaluators each render their appropriate view
  • An experiment missing an evaluator or metric shows as no data, not as zero
  • The base experiment is visually marked as the reference
  • Experiment colors match the compare page's experiment colors
  • The experiments page charts behave as before
  • Storybook story covers the compare chart variant

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

  • Status
    📘 Todo

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions