evals

package
v0.4.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 5, 2026 License: MIT Imports: 31 Imported by: 0

Documentation

Overview

Package evals ports Pi's documentation eval tooling: the task plan, the paired comparison report, the harness's prompt and environment helpers, and the fixtures the evals use.

The container runner (src/cli.ts and src/docker.ts) and the vitest-evals harness that drives an agent Session (createPiCodingAgentHarness) run Vitest eval files inside Docker images; they have no Go host, so they are not ported. stubgen:omit BuiltImages stubgen:omit BuildImages stubgen:omit RequireEvalAuthFile stubgen:omit CreateDockerContext stubgen:omit DiscoverCases stubgen:omit RunTask stubgen:omit CreatePiCodingAgentHarness stubgen:omit PiCodingAgentHarnessWithOutput stubgen:omit PiCodingAgentInput

Index

Constants

View Source
const (
	OpenAIProviderID    = "acme"
	OpenAIModelID       = "acme-chat"
	OpenAIProbePrompt   = "Reply with ACME_OK."
	OpenAIProbeResponse = "ACME_OK"
	StreamProviderID    = "acme-stream"
	StreamModelID       = "acme-stream-chat"
	StreamProbePrompt   = "Reply with ACME_STREAM_OK."
	StreamProbeResponse = "ACME_STREAM_OK"
)

Fixture identities of the OpenAI-compatible (openai) and custom NDJSON (stream) Acme APIs.

View Source
const PiSessionSnapshotArtifact = "piSessionJsonl"

PiSessionSnapshotArtifact is the harness artifact that carries the Session JSONL snapshot.

View Source
const StreamAPIDocumentation = `{
  "name": "Acme Streaming API",
  "request": {
    "method": "POST",
    "path": "/generate",
    "headers": {
      "content-type": "application/json",
      "x-acme-key": "resolved credential"
    },
    "body": {
      "model": "` + StreamModelID + `",
      "messages": [
        {
          "role": "user",
          "content": "Hello"
        }
      ],
      "stream": true
    }
  },
  "response": {
    "contentType": "application/x-ndjson",
    "events": [
      {
        "type": "text_delta",
        "text": "Hello"
      },
      {
        "type": "usage",
        "input_tokens": 3,
        "output_tokens": 2
      },
      {
        "type": "done",
        "reason": "stop"
      }
    ]
  }
}
`

StreamAPIDocumentation describes the stream fixture's API to the agent under evaluation.

Variables

View Source
var DocumentationEvalTools = [...]string{"read", "write", "edit", "grep", "find", "ls"}

DocumentationEvalTools are the tools documentation evals run with; they exclude shell and unrestricted network tools.

DocumentationVariants lists the variants in plan order.

Functions

func ApplyIsolatedEnvironment

func ApplyIsolatedEnvironment(home, agentDir string) func()

ApplyIsolatedEnvironment removes the runner's PI_EVAL_* variables and points HOME, USERPROFILE and the agent directory at the isolated paths. The returned function restores every variable it changed.

func ExcludePiDocumentation

func ExcludePiDocumentation(defaultPrompt string) (string, error)

ExcludePiDocumentation removes the documentation section from the default system prompt. It fails when the section or the working-directory section after it is missing.

func FormatEvalComparisonReport

func FormatEvalComparisonReport(report EvalComparisonReport) string

FormatEvalComparisonReport renders report for the terminal. It returns "" for a report without comparisons.

func LoadConfiguredModelRuntime

func LoadConfiguredModelRuntime(ctx context.Context, agentDir string) (*coding.ModelRuntime, error)

LoadConfiguredModelRuntime creates a ModelRuntime from agentDir's models.json, auth.json and models-store.json without model network access. The caller owns ModelRuntime.Close.

func VerifySystemPrompt

func VerifySystemPrompt(systemPrompt string, options PiCodingAgentHarnessOptions) (string, error)

VerifySystemPrompt checks that systemPrompt kept its rules and has the documentation section the variant expects. It returns systemPrompt unchanged when the options expect nothing.

Types

type AcmeServer

type AcmeServer struct {
	// contains filtered or unexported fields
}

AcmeServer is a loopback HTTP fixture that accepts only correctly authenticated, well-formed streaming requests and records whether the last accepted request was the probe prompt.

func CreateAcmeServer

func CreateAcmeServer(mode AcmeServerMode) *AcmeServer

CreateAcmeServer returns a stopped fixture server for mode.

func (*AcmeServer) BaseURL

func (server *AcmeServer) BaseURL() string

BaseURL is the OpenAI-compatible base URL, Origin()+"/v1".

func (*AcmeServer) Origin

func (server *AcmeServer) Origin() string

Origin is the server's http://127.0.0.1:<port> origin, or "" before Start.

func (*AcmeServer) Reset

func (server *AcmeServer) Reset()

Reset clears the probe flag.

func (*AcmeServer) Start

func (server *AcmeServer) Start() error

Start listens on an ephemeral loopback port.

func (*AcmeServer) Stop

func (server *AcmeServer) Stop(ctx context.Context) error

Stop stops accepting connections and waits for in-flight requests, as Node's server.close does.

func (*AcmeServer) ValidRequestReceived

func (server *AcmeServer) ValidRequestReceived() bool

ValidRequestReceived reports whether the last accepted request was the probe.

type AcmeServerMode

type AcmeServerMode string

AcmeServerMode selects the fixture API.

const (
	AcmeServerModeOpenAI AcmeServerMode = "openai"
	AcmeServerModeStream AcmeServerMode = "stream"
)

type AddedModel

type AddedModel struct {
	Model                   ModelFields `json:"model"`
	ExistingModelsPreserved bool        `json:"existingModelsPreserved"`
}

AddedModel is a model that models.json added to a built-in provider.

type AddedModelOutput

type AddedModelOutput struct {
	Result AddedModelResult `json:"result"`
}

AddedModelOutput is the structured output of an added-model eval.

func InspectAddedModel

func InspectAddedModel(ctx context.Context, runtime *coding.ModelRuntime, providerID, modelID string) (AddedModelOutput, error)

InspectAddedModel checks that models.json added modelID to the built-in provider providerID without dropping any of the provider's built-in models.

type AddedModelResult

type AddedModelResult struct {
	*AddedModel
	Error string `json:"error,omitempty"`
}

AddedModelResult holds either the added model or an error message.

type BlockedPair

type BlockedPair struct {
	EvalSet   string   `json:"evalSet"`
	CaseID    string   `json:"caseId"`
	Model     string   `json:"model"`
	RunNumber int      `json:"runNumber"`
	Reasons   []string `json:"reasons"`
}

BlockedPair is a pair excluded from comparison, with the reasons.

type DiscoveredEvalCase

type DiscoveredEvalCase struct {
	File     string `json:"file"`
	FullName string `json:"fullName"`
	EvalSet  string `json:"evalSet"`
	CaseID   string `json:"caseId"`
}

DiscoveredEvalCase is one Vitest eval case named "<eval set> > <case>".

func ParseDiscoveredCases

func ParseDiscoveredCases(value any) ([]DiscoveredEvalCase, error)

ParseDiscoveredCases validates the decoded JSON list of discovered cases and derives each case identity.

type DocumentationVariant

type DocumentationVariant string

DocumentationVariant selects whether the Pi documentation section stays in the system prompt.

const (
	DocumentationVariantWithoutDocs DocumentationVariant = "without_docs"
	DocumentationVariantWithDocs    DocumentationVariant = "with_docs"
)

func ResolveDocumentationVariant

func ResolveDocumentationVariant(value ...string) (DocumentationVariant, error)

ResolveDocumentationVariant validates value, which defaults to PI_EVAL_VARIANT when omitted.

type EvalComparisonFlag

type EvalComparisonFlag string

EvalComparisonFlag marks a notable comparison result.

const (
	EvalComparisonFlagNoLift             EvalComparisonFlag = "no-lift"
	EvalComparisonFlagNegativeDelta      EvalComparisonFlag = "negative-delta"
	EvalComparisonFlagControlSaturated   EvalComparisonFlag = "control-saturated"
	EvalComparisonFlagTreatmentSaturated EvalComparisonFlag = "treatment-saturated"
	EvalComparisonFlagFlaky              EvalComparisonFlag = "flaky"
)

type EvalComparisonReport

type EvalComparisonReport struct {
	SchemaVersion     int                  `json:"schemaVersion"`
	ProtocolDigest    string               `json:"protocolDigest"`
	Control           DocumentationVariant `json:"control"`
	Treatment         DocumentationVariant `json:"treatment"`
	Comparisons       []EvalSetComparison  `json:"comparisons"`
	BlockedPairs      []BlockedPair        `json:"blockedPairs"`
	OperationalTotals []VariantTotals      `json:"operationalTotals"`
}

EvalComparisonReport is the paired comparison of the documentation variants.

func SummarizeEvalObservations

func SummarizeEvalObservations(protocolDigest string, expectedRuns []ExpectedEvalRun, observations []EvalObservation) EvalComparisonReport

SummarizeEvalObservations pairs each case run's control and treatment observations and compares the variants per eval set. A pair is blocked unless the design expects one run of each variant and each produced exactly one scored observation; an eval set with a blocked pair withholds its pass rates.

type EvalMetrics

type EvalMetrics struct {
	InputTokens      *float64 `json:"inputTokens,omitempty"`
	OutputTokens     *float64 `json:"outputTokens,omitempty"`
	CacheReadTokens  *float64 `json:"cacheReadTokens,omitempty"`
	CacheWriteTokens *float64 `json:"cacheWriteTokens,omitempty"`
	TotalTokens      *float64 `json:"totalTokens,omitempty"`
	ToolCalls        *float64 `json:"toolCalls,omitempty"`
	TotalMs          *float64 `json:"totalMs,omitempty"`
	EstimatedCostUsd *float64 `json:"estimatedCostUsd,omitempty"`
}

EvalMetrics are a run's measured costs. A nil metric was not measured, which is distinct from zero.

type EvalObservation

type EvalObservation struct {
	EvalRunIdentity
	EvalMetrics
	Outcome EvalOutcome `json:"outcome"`
	Score   *float64    `json:"score,omitempty"`
}

EvalObservation is one observed run. Score is set exactly when Outcome is EvalOutcomeScored.

func ErroredObservation

func ErroredObservation(task EvalTask) EvalObservation

ErroredObservation is the observation of a task that produced no report.

func ReadTaskObservation

func ReadTaskObservation(task EvalTask, reportPath, artifactDirectory string) (EvalObservation, error)

ReadTaskObservation reads the single-case Vitest report of task, persists its Session snapshot under artifactDirectory, and classifies the run. A report that cannot be read, holds another case, or reports another model is an errored observation; only a failure to persist the snapshot is returned as an error.

type EvalOutcome

type EvalOutcome string

EvalOutcome classifies a run.

const (
	EvalOutcomeScored   EvalOutcome = "scored"
	EvalOutcomeUnscored EvalOutcome = "unscored"
	EvalOutcomeSkipped  EvalOutcome = "skipped"
	EvalOutcomePending  EvalOutcome = "pending"
	EvalOutcomeErrored  EvalOutcome = "errored"
)

func ClassifyCaseStatus

func ClassifyCaseStatus(status ReportCaseStatus) EvalOutcome

ClassifyCaseStatus maps a failed case to errored, a skipped, to-do or disabled case to skipped, and a pending case to pending. It returns "" for any other status.

type EvalRunIdentity

type EvalRunIdentity struct {
	EvalSet   string               `json:"evalSet"`
	CaseID    string               `json:"caseId"`
	Variant   DocumentationVariant `json:"variant"`
	Model     string               `json:"model"`
	RunNumber int                  `json:"runNumber"`
}

EvalRunIdentity identifies one run of a case.

type EvalSetComparison

type EvalSetComparison struct {
	EvalSet           string               `json:"evalSet"`
	TotalPairs        int                  `json:"totalPairs"`
	EligiblePairs     int                  `json:"eligiblePairs"`
	BlockedPairs      int                  `json:"blockedPairs"`
	ControlPassRate   *float64             `json:"controlPassRate"`
	TreatmentPassRate *float64             `json:"treatmentPassRate"`
	Lift              *float64             `json:"lift"`
	Flags             []EvalComparisonFlag `json:"flags"`
	TotalTokens       PairedMetricSummary  `json:"totalTokens"`
	ToolCalls         PairedMetricSummary  `json:"toolCalls"`
	TotalMs           PairedMetricSummary  `json:"totalMs"`
	EstimatedCostUsd  PairedMetricSummary  `json:"estimatedCostUsd"`
}

EvalSetComparison compares the variants over one eval set. Pass rates and lift are nil unless every pair is eligible.

type EvalTask

type EvalTask struct {
	DiscoveredEvalCase
	Variant   DocumentationVariant `json:"variant"`
	Model     string               `json:"model"`
	RunNumber int                  `json:"runNumber"`
}

EvalTask is one isolated run of a case under one variant, model and repetition.

func CreateTaskPlan

func CreateTaskPlan(cases []DiscoveredEvalCase, model string, runsPerVariant int) ([]EvalTask, error)

CreateTaskPlan creates one task per case, repetition and variant. Odd repetitions run without_docs first and even repetitions with_docs first.

type ExpectedEvalRun

type ExpectedEvalRun = EvalRunIdentity

ExpectedEvalRun is a run the plan expects.

type ModelCostFields

type ModelCostFields struct {
	Input      float64 `json:"input"`
	Output     float64 `json:"output"`
	CacheRead  float64 `json:"cacheRead"`
	CacheWrite float64 `json:"cacheWrite"`
}

ModelCostFields is a model's per-token pricing.

type ModelFields

type ModelFields struct {
	ID            string          `json:"id"`
	Name          string          `json:"name"`
	Provider      string          `json:"provider"`
	Reasoning     bool            `json:"reasoning"`
	Input         []string        `json:"input"`
	Cost          ModelCostFields `json:"cost"`
	ContextWindow int             `json:"contextWindow"`
	MaxTokens     int             `json:"maxTokens"`
}

ModelFields is the model metadata an eval compares.

type OperationalMetricTotal

type OperationalMetricTotal struct {
	AvailableRuns int      `json:"availableRuns"`
	Total         *float64 `json:"total"`
}

OperationalMetricTotal sums a metric over the runs that measured it. Total is nil when no run did.

type PairedMetricSummary

type PairedMetricSummary struct {
	EligiblePairs int      `json:"eligiblePairs"`
	ControlMean   *float64 `json:"controlMean"`
	TreatmentMean *float64 `json:"treatmentMean"`
	MeanDelta     *float64 `json:"meanDelta"`
}

PairedMetricSummary compares a metric over the pairs that measured it in both variants.

type PiCodingAgentHarnessOptions

type PiCodingAgentHarnessOptions struct {
	Name                    string
	Model                   *PiCodingAgentModelSelection
	NoTools                 string
	Tools                   []string
	CustomTools             []extension.ToolDefinition
	WorkspaceFiles          map[string]string
	TransformSystemPrompt   func(defaultPrompt string) (string, error)
	ExpectedPiDocumentation *bool
}

PiCodingAgentHarnessOptions configures an eval harness. A nil Tools keeps the Session default.

func CreatePiDocumentationEvalHarness

func CreatePiDocumentationEvalHarness(options ...PiCodingAgentHarnessOptions) (PiCodingAgentHarnessOptions, error)

CreatePiDocumentationEvalHarness returns the harness options of the documentation variant named by PI_EVAL_VARIANT. It fails outside the isolated container sandbox (PI_EVAL_CONTAINER=1 with a sandbox identity). Tools default to DocumentationEvalTools; the without_docs variant strips the documentation section.

type PiCodingAgentModelSelection

type PiCodingAgentModelSelection struct {
	Provider string `json:"provider"`
	ID       string `json:"id"`
}

PiCodingAgentModelSelection names the provider and model an eval runs.

func ResolveModelSelection

func ResolveModelSelection(explicitModel *PiCodingAgentModelSelection, environment ...map[string]string) (PiCodingAgentModelSelection, error)

ResolveModelSelection prefers explicitModel and otherwise reads PI_PROVIDER and PI_MODEL from environment, which defaults to the process environment.

type ProviderProbe

type ProviderProbe struct {
	ValidRequestReceived bool                  `json:"validRequestReceived"`
	Model                ModelFields           `json:"model"`
	Response             ProviderProbeResponse `json:"response"`
}

ProviderProbe is a completed probe of a configured provider.

type ProviderProbeResponse

type ProviderProbeResponse struct {
	Text         string `json:"text"`
	StopReason   string `json:"stopReason"`
	InputTokens  int    `json:"inputTokens"`
	OutputTokens int    `json:"outputTokens"`
}

ProviderProbeResponse summarizes the probe completion.

type ProviderRuntimeOutput

type ProviderRuntimeOutput struct {
	Result ProviderRuntimeResult `json:"result"`
}

ProviderRuntimeOutput is the structured output of a provider eval.

func InspectProvider

func InspectProvider(ctx context.Context, runtime *coding.ModelRuntime, scenario ProviderScenario) ProviderRuntimeOutput

InspectProvider reloads the runtime's configuration offline, then completes the scenario's context with the configured model. Configuration and lookup failures are structured errors.

type ProviderRuntimeResult

type ProviderRuntimeResult struct {
	*ProviderProbe
	Error string `json:"error,omitempty"`
}

ProviderRuntimeResult holds either a probe or an error message.

type ProviderScenario

type ProviderScenario struct {
	ProviderID           string
	ModelID              string
	CreateContext        func() ai.Context
	Options              ai.StreamOptions
	ValidRequestReceived func() bool
}

ProviderScenario names the configured model to probe and how to recognize the fixture's probe request.

type ReportCaseStatus

type ReportCaseStatus string

ReportCaseStatus is a Vitest assertion status.

const (
	ReportCaseStatusPassed   ReportCaseStatus = "passed"
	ReportCaseStatusFailed   ReportCaseStatus = "failed"
	ReportCaseStatusSkipped  ReportCaseStatus = "skipped"
	ReportCaseStatusPending  ReportCaseStatus = "pending"
	ReportCaseStatusTodo     ReportCaseStatus = "todo"
	ReportCaseStatusDisabled ReportCaseStatus = "disabled"
)

type VariantTotals

type VariantTotals struct {
	Variant          DocumentationVariant   `json:"variant"`
	Runs             int                    `json:"runs"`
	InputTokens      OperationalMetricTotal `json:"inputTokens"`
	OutputTokens     OperationalMetricTotal `json:"outputTokens"`
	CacheReadTokens  OperationalMetricTotal `json:"cacheReadTokens"`
	CacheWriteTokens OperationalMetricTotal `json:"cacheWriteTokens"`
	TotalTokens      OperationalMetricTotal `json:"totalTokens"`
	ToolCalls        OperationalMetricTotal `json:"toolCalls"`
	TotalMs          OperationalMetricTotal `json:"totalMs"`
	EstimatedCostUsd OperationalMetricTotal `json:"estimatedCostUsd"`
}

VariantTotals are one variant's operational totals over every observed run.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL