Documentation
¶
Overview ¶
Package evals ports Pi's documentation eval tooling: the task plan, the paired comparison report, the harness's prompt and environment helpers, and the fixtures the evals use.
The container runner (src/cli.ts and src/docker.ts) and the vitest-evals harness that drives an agent Session (createPiCodingAgentHarness) run Vitest eval files inside Docker images; they have no Go host, so they are not ported. stubgen:omit BuiltImages stubgen:omit BuildImages stubgen:omit RequireEvalAuthFile stubgen:omit CreateDockerContext stubgen:omit DiscoverCases stubgen:omit RunTask stubgen:omit CreatePiCodingAgentHarness stubgen:omit PiCodingAgentHarnessWithOutput stubgen:omit PiCodingAgentInput
Index ¶
- Constants
- Variables
- func ApplyIsolatedEnvironment(home, agentDir string) func()
- func ExcludePiDocumentation(defaultPrompt string) (string, error)
- func FormatEvalComparisonReport(report EvalComparisonReport) string
- func LoadConfiguredModelRuntime(ctx context.Context, agentDir string) (*coding.ModelRuntime, error)
- func VerifySystemPrompt(systemPrompt string, options PiCodingAgentHarnessOptions) (string, error)
- type AcmeServer
- type AcmeServerMode
- type AddedModel
- type AddedModelOutput
- type AddedModelResult
- type BlockedPair
- type DiscoveredEvalCase
- type DocumentationVariant
- type EvalComparisonFlag
- type EvalComparisonReport
- type EvalMetrics
- type EvalObservation
- type EvalOutcome
- type EvalRunIdentity
- type EvalSetComparison
- type EvalTask
- type ExpectedEvalRun
- type ModelCostFields
- type ModelFields
- type OperationalMetricTotal
- type PairedMetricSummary
- type PiCodingAgentHarnessOptions
- type PiCodingAgentModelSelection
- type ProviderProbe
- type ProviderProbeResponse
- type ProviderRuntimeOutput
- type ProviderRuntimeResult
- type ProviderScenario
- type ReportCaseStatus
- type VariantTotals
Constants ¶
const ( OpenAIProviderID = "acme" OpenAIModelID = "acme-chat" OpenAIProbePrompt = "Reply with ACME_OK." OpenAIProbeResponse = "ACME_OK" StreamProviderID = "acme-stream" StreamModelID = "acme-stream-chat" StreamProbePrompt = "Reply with ACME_STREAM_OK." StreamProbeResponse = "ACME_STREAM_OK" )
Fixture identities of the OpenAI-compatible (openai) and custom NDJSON (stream) Acme APIs.
const PiSessionSnapshotArtifact = "piSessionJsonl"
PiSessionSnapshotArtifact is the harness artifact that carries the Session JSONL snapshot.
const StreamAPIDocumentation = `{ "name": "Acme Streaming API", "request": { "method": "POST", "path": "/generate", "headers": { "content-type": "application/json", "x-acme-key": "resolved credential" }, "body": { "model": "` + StreamModelID + `", "messages": [ { "role": "user", "content": "Hello" } ], "stream": true } }, "response": { "contentType": "application/x-ndjson", "events": [ { "type": "text_delta", "text": "Hello" }, { "type": "usage", "input_tokens": 3, "output_tokens": 2 }, { "type": "done", "reason": "stop" } ] } } `
StreamAPIDocumentation describes the stream fixture's API to the agent under evaluation.
Variables ¶
var DocumentationEvalTools = [...]string{"read", "write", "edit", "grep", "find", "ls"}
DocumentationEvalTools are the tools documentation evals run with; they exclude shell and unrestricted network tools.
var DocumentationVariants = [...]DocumentationVariant{DocumentationVariantWithoutDocs, DocumentationVariantWithDocs}
DocumentationVariants lists the variants in plan order.
Functions ¶
func ApplyIsolatedEnvironment ¶
func ApplyIsolatedEnvironment(home, agentDir string) func()
ApplyIsolatedEnvironment removes the runner's PI_EVAL_* variables and points HOME, USERPROFILE and the agent directory at the isolated paths. The returned function restores every variable it changed.
func ExcludePiDocumentation ¶
ExcludePiDocumentation removes the documentation section from the default system prompt. It fails when the section or the working-directory section after it is missing.
func FormatEvalComparisonReport ¶
func FormatEvalComparisonReport(report EvalComparisonReport) string
FormatEvalComparisonReport renders report for the terminal. It returns "" for a report without comparisons.
func LoadConfiguredModelRuntime ¶
LoadConfiguredModelRuntime creates a ModelRuntime from agentDir's models.json, auth.json and models-store.json without model network access. The caller owns ModelRuntime.Close.
func VerifySystemPrompt ¶
func VerifySystemPrompt(systemPrompt string, options PiCodingAgentHarnessOptions) (string, error)
VerifySystemPrompt checks that systemPrompt kept its rules and has the documentation section the variant expects. It returns systemPrompt unchanged when the options expect nothing.
Types ¶
type AcmeServer ¶
type AcmeServer struct {
// contains filtered or unexported fields
}
AcmeServer is a loopback HTTP fixture that accepts only correctly authenticated, well-formed streaming requests and records whether the last accepted request was the probe prompt.
func CreateAcmeServer ¶
func CreateAcmeServer(mode AcmeServerMode) *AcmeServer
CreateAcmeServer returns a stopped fixture server for mode.
func (*AcmeServer) BaseURL ¶
func (server *AcmeServer) BaseURL() string
BaseURL is the OpenAI-compatible base URL, Origin()+"/v1".
func (*AcmeServer) Origin ¶
func (server *AcmeServer) Origin() string
Origin is the server's http://127.0.0.1:<port> origin, or "" before Start.
func (*AcmeServer) Start ¶
func (server *AcmeServer) Start() error
Start listens on an ephemeral loopback port.
func (*AcmeServer) Stop ¶
func (server *AcmeServer) Stop(ctx context.Context) error
Stop stops accepting connections and waits for in-flight requests, as Node's server.close does.
func (*AcmeServer) ValidRequestReceived ¶
func (server *AcmeServer) ValidRequestReceived() bool
ValidRequestReceived reports whether the last accepted request was the probe.
type AcmeServerMode ¶
type AcmeServerMode string
AcmeServerMode selects the fixture API.
const ( AcmeServerModeOpenAI AcmeServerMode = "openai" AcmeServerModeStream AcmeServerMode = "stream" )
type AddedModel ¶
type AddedModel struct {
Model ModelFields `json:"model"`
ExistingModelsPreserved bool `json:"existingModelsPreserved"`
}
AddedModel is a model that models.json added to a built-in provider.
type AddedModelOutput ¶
type AddedModelOutput struct {
Result AddedModelResult `json:"result"`
}
AddedModelOutput is the structured output of an added-model eval.
func InspectAddedModel ¶
func InspectAddedModel(ctx context.Context, runtime *coding.ModelRuntime, providerID, modelID string) (AddedModelOutput, error)
InspectAddedModel checks that models.json added modelID to the built-in provider providerID without dropping any of the provider's built-in models.
type AddedModelResult ¶
type AddedModelResult struct {
*AddedModel
Error string `json:"error,omitempty"`
}
AddedModelResult holds either the added model or an error message.
type BlockedPair ¶
type BlockedPair struct {
EvalSet string `json:"evalSet"`
CaseID string `json:"caseId"`
Model string `json:"model"`
RunNumber int `json:"runNumber"`
Reasons []string `json:"reasons"`
}
BlockedPair is a pair excluded from comparison, with the reasons.
type DiscoveredEvalCase ¶
type DiscoveredEvalCase struct {
File string `json:"file"`
FullName string `json:"fullName"`
EvalSet string `json:"evalSet"`
CaseID string `json:"caseId"`
}
DiscoveredEvalCase is one Vitest eval case named "<eval set> > <case>".
func ParseDiscoveredCases ¶
func ParseDiscoveredCases(value any) ([]DiscoveredEvalCase, error)
ParseDiscoveredCases validates the decoded JSON list of discovered cases and derives each case identity.
type DocumentationVariant ¶
type DocumentationVariant string
DocumentationVariant selects whether the Pi documentation section stays in the system prompt.
const ( DocumentationVariantWithoutDocs DocumentationVariant = "without_docs" DocumentationVariantWithDocs DocumentationVariant = "with_docs" )
func ResolveDocumentationVariant ¶
func ResolveDocumentationVariant(value ...string) (DocumentationVariant, error)
ResolveDocumentationVariant validates value, which defaults to PI_EVAL_VARIANT when omitted.
type EvalComparisonFlag ¶
type EvalComparisonFlag string
EvalComparisonFlag marks a notable comparison result.
const ( EvalComparisonFlagNoLift EvalComparisonFlag = "no-lift" EvalComparisonFlagNegativeDelta EvalComparisonFlag = "negative-delta" EvalComparisonFlagControlSaturated EvalComparisonFlag = "control-saturated" EvalComparisonFlagTreatmentSaturated EvalComparisonFlag = "treatment-saturated" EvalComparisonFlagFlaky EvalComparisonFlag = "flaky" )
type EvalComparisonReport ¶
type EvalComparisonReport struct {
SchemaVersion int `json:"schemaVersion"`
ProtocolDigest string `json:"protocolDigest"`
Control DocumentationVariant `json:"control"`
Treatment DocumentationVariant `json:"treatment"`
Comparisons []EvalSetComparison `json:"comparisons"`
BlockedPairs []BlockedPair `json:"blockedPairs"`
OperationalTotals []VariantTotals `json:"operationalTotals"`
}
EvalComparisonReport is the paired comparison of the documentation variants.
func SummarizeEvalObservations ¶
func SummarizeEvalObservations(protocolDigest string, expectedRuns []ExpectedEvalRun, observations []EvalObservation) EvalComparisonReport
SummarizeEvalObservations pairs each case run's control and treatment observations and compares the variants per eval set. A pair is blocked unless the design expects one run of each variant and each produced exactly one scored observation; an eval set with a blocked pair withholds its pass rates.
type EvalMetrics ¶
type EvalMetrics struct {
InputTokens *float64 `json:"inputTokens,omitempty"`
OutputTokens *float64 `json:"outputTokens,omitempty"`
CacheReadTokens *float64 `json:"cacheReadTokens,omitempty"`
CacheWriteTokens *float64 `json:"cacheWriteTokens,omitempty"`
TotalTokens *float64 `json:"totalTokens,omitempty"`
ToolCalls *float64 `json:"toolCalls,omitempty"`
TotalMs *float64 `json:"totalMs,omitempty"`
EstimatedCostUsd *float64 `json:"estimatedCostUsd,omitempty"`
}
EvalMetrics are a run's measured costs. A nil metric was not measured, which is distinct from zero.
type EvalObservation ¶
type EvalObservation struct {
EvalRunIdentity
EvalMetrics
Outcome EvalOutcome `json:"outcome"`
Score *float64 `json:"score,omitempty"`
}
EvalObservation is one observed run. Score is set exactly when Outcome is EvalOutcomeScored.
func ErroredObservation ¶
func ErroredObservation(task EvalTask) EvalObservation
ErroredObservation is the observation of a task that produced no report.
func ReadTaskObservation ¶
func ReadTaskObservation(task EvalTask, reportPath, artifactDirectory string) (EvalObservation, error)
ReadTaskObservation reads the single-case Vitest report of task, persists its Session snapshot under artifactDirectory, and classifies the run. A report that cannot be read, holds another case, or reports another model is an errored observation; only a failure to persist the snapshot is returned as an error.
type EvalOutcome ¶
type EvalOutcome string
EvalOutcome classifies a run.
const ( EvalOutcomeScored EvalOutcome = "scored" EvalOutcomeUnscored EvalOutcome = "unscored" EvalOutcomeSkipped EvalOutcome = "skipped" EvalOutcomePending EvalOutcome = "pending" EvalOutcomeErrored EvalOutcome = "errored" )
func ClassifyCaseStatus ¶
func ClassifyCaseStatus(status ReportCaseStatus) EvalOutcome
ClassifyCaseStatus maps a failed case to errored, a skipped, to-do or disabled case to skipped, and a pending case to pending. It returns "" for any other status.
type EvalRunIdentity ¶
type EvalRunIdentity struct {
EvalSet string `json:"evalSet"`
CaseID string `json:"caseId"`
Variant DocumentationVariant `json:"variant"`
Model string `json:"model"`
RunNumber int `json:"runNumber"`
}
EvalRunIdentity identifies one run of a case.
type EvalSetComparison ¶
type EvalSetComparison struct {
EvalSet string `json:"evalSet"`
TotalPairs int `json:"totalPairs"`
EligiblePairs int `json:"eligiblePairs"`
BlockedPairs int `json:"blockedPairs"`
ControlPassRate *float64 `json:"controlPassRate"`
TreatmentPassRate *float64 `json:"treatmentPassRate"`
Lift *float64 `json:"lift"`
Flags []EvalComparisonFlag `json:"flags"`
TotalTokens PairedMetricSummary `json:"totalTokens"`
ToolCalls PairedMetricSummary `json:"toolCalls"`
TotalMs PairedMetricSummary `json:"totalMs"`
EstimatedCostUsd PairedMetricSummary `json:"estimatedCostUsd"`
}
EvalSetComparison compares the variants over one eval set. Pass rates and lift are nil unless every pair is eligible.
type EvalTask ¶
type EvalTask struct {
DiscoveredEvalCase
Variant DocumentationVariant `json:"variant"`
Model string `json:"model"`
RunNumber int `json:"runNumber"`
}
EvalTask is one isolated run of a case under one variant, model and repetition.
func CreateTaskPlan ¶
func CreateTaskPlan(cases []DiscoveredEvalCase, model string, runsPerVariant int) ([]EvalTask, error)
CreateTaskPlan creates one task per case, repetition and variant. Odd repetitions run without_docs first and even repetitions with_docs first.
type ExpectedEvalRun ¶
type ExpectedEvalRun = EvalRunIdentity
ExpectedEvalRun is a run the plan expects.
type ModelCostFields ¶
type ModelCostFields struct {
Input float64 `json:"input"`
Output float64 `json:"output"`
CacheRead float64 `json:"cacheRead"`
CacheWrite float64 `json:"cacheWrite"`
}
ModelCostFields is a model's per-token pricing.
type ModelFields ¶
type ModelFields struct {
ID string `json:"id"`
Name string `json:"name"`
Provider string `json:"provider"`
Reasoning bool `json:"reasoning"`
Input []string `json:"input"`
Cost ModelCostFields `json:"cost"`
ContextWindow int `json:"contextWindow"`
MaxTokens int `json:"maxTokens"`
}
ModelFields is the model metadata an eval compares.
type OperationalMetricTotal ¶
type OperationalMetricTotal struct {
AvailableRuns int `json:"availableRuns"`
Total *float64 `json:"total"`
}
OperationalMetricTotal sums a metric over the runs that measured it. Total is nil when no run did.
type PairedMetricSummary ¶
type PairedMetricSummary struct {
EligiblePairs int `json:"eligiblePairs"`
ControlMean *float64 `json:"controlMean"`
TreatmentMean *float64 `json:"treatmentMean"`
MeanDelta *float64 `json:"meanDelta"`
}
PairedMetricSummary compares a metric over the pairs that measured it in both variants.
type PiCodingAgentHarnessOptions ¶
type PiCodingAgentHarnessOptions struct {
Name string
Model *PiCodingAgentModelSelection
NoTools string
Tools []string
CustomTools []extension.ToolDefinition
WorkspaceFiles map[string]string
TransformSystemPrompt func(defaultPrompt string) (string, error)
ExpectedPiDocumentation *bool
}
PiCodingAgentHarnessOptions configures an eval harness. A nil Tools keeps the Session default.
func CreatePiDocumentationEvalHarness ¶
func CreatePiDocumentationEvalHarness(options ...PiCodingAgentHarnessOptions) (PiCodingAgentHarnessOptions, error)
CreatePiDocumentationEvalHarness returns the harness options of the documentation variant named by PI_EVAL_VARIANT. It fails outside the isolated container sandbox (PI_EVAL_CONTAINER=1 with a sandbox identity). Tools default to DocumentationEvalTools; the without_docs variant strips the documentation section.
type PiCodingAgentModelSelection ¶
PiCodingAgentModelSelection names the provider and model an eval runs.
func ResolveModelSelection ¶
func ResolveModelSelection(explicitModel *PiCodingAgentModelSelection, environment ...map[string]string) (PiCodingAgentModelSelection, error)
ResolveModelSelection prefers explicitModel and otherwise reads PI_PROVIDER and PI_MODEL from environment, which defaults to the process environment.
type ProviderProbe ¶
type ProviderProbe struct {
ValidRequestReceived bool `json:"validRequestReceived"`
Model ModelFields `json:"model"`
Response ProviderProbeResponse `json:"response"`
}
ProviderProbe is a completed probe of a configured provider.
type ProviderProbeResponse ¶
type ProviderProbeResponse struct {
Text string `json:"text"`
StopReason string `json:"stopReason"`
InputTokens int `json:"inputTokens"`
OutputTokens int `json:"outputTokens"`
}
ProviderProbeResponse summarizes the probe completion.
type ProviderRuntimeOutput ¶
type ProviderRuntimeOutput struct {
Result ProviderRuntimeResult `json:"result"`
}
ProviderRuntimeOutput is the structured output of a provider eval.
func InspectProvider ¶
func InspectProvider(ctx context.Context, runtime *coding.ModelRuntime, scenario ProviderScenario) ProviderRuntimeOutput
InspectProvider reloads the runtime's configuration offline, then completes the scenario's context with the configured model. Configuration and lookup failures are structured errors.
type ProviderRuntimeResult ¶
type ProviderRuntimeResult struct {
*ProviderProbe
Error string `json:"error,omitempty"`
}
ProviderRuntimeResult holds either a probe or an error message.
type ProviderScenario ¶
type ProviderScenario struct {
ProviderID string
ModelID string
CreateContext func() ai.Context
Options ai.StreamOptions
ValidRequestReceived func() bool
}
ProviderScenario names the configured model to probe and how to recognize the fixture's probe request.
type ReportCaseStatus ¶
type ReportCaseStatus string
ReportCaseStatus is a Vitest assertion status.
const ( ReportCaseStatusPassed ReportCaseStatus = "passed" ReportCaseStatusFailed ReportCaseStatus = "failed" ReportCaseStatusSkipped ReportCaseStatus = "skipped" ReportCaseStatusPending ReportCaseStatus = "pending" ReportCaseStatusTodo ReportCaseStatus = "todo" ReportCaseStatusDisabled ReportCaseStatus = "disabled" )
type VariantTotals ¶
type VariantTotals struct {
Variant DocumentationVariant `json:"variant"`
Runs int `json:"runs"`
InputTokens OperationalMetricTotal `json:"inputTokens"`
OutputTokens OperationalMetricTotal `json:"outputTokens"`
CacheReadTokens OperationalMetricTotal `json:"cacheReadTokens"`
CacheWriteTokens OperationalMetricTotal `json:"cacheWriteTokens"`
TotalTokens OperationalMetricTotal `json:"totalTokens"`
ToolCalls OperationalMetricTotal `json:"toolCalls"`
TotalMs OperationalMetricTotal `json:"totalMs"`
EstimatedCostUsd OperationalMetricTotal `json:"estimatedCostUsd"`
}
VariantTotals are one variant's operational totals over every observed run.