Skip to content

[BUG]: experiment warnings miscount failed runs and resumed evaluations #16754

Description

@4ktLuffy

Where do you use Phoenix

Self-hosted

What version of Phoenix are you using?

main (9212a42)

What happened?

Two warnings in phoenix.client experiments are wrong:

  1. run_experiment: "Only X out of Y expected runs were completed successfully" never fires when a task fails, because the runs re-fetched from the server include errored runs and all of them are counted.
  2. resume_evaluation: the warning compares completed evaluations against runs × all evaluators, but only each run's missing evaluations are re-run. Re-running one missing evaluation successfully prints "Only 1 out of 2 incomplete evaluations were completed successfully."

Expected: (1) warns when any run failed; (2) warns only when an attempted evaluation failed.

Additional information

packages/phoenix-client/src/phoenix/client/resources/experiments/__init__.py, sync and async (actual_runs = len(task_runs); total_processed * len(evaluators_by_name)).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't workingc/clientlanguage: pythonIssues with python as the primary language of use

Type

No type

Projects

  • Status
    📘 Todo

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions