Documentation
¶
Overview ¶
Package evaluation implements the JPS Core §§7–10 evaluator contract behind an explicitly experimental surface (ADR-0007).
It implements the evaluator class Core 0.2.0-draft defines: the input preflight of §8.2, the portable disposition of §8.3 with its RFC 8785 byte-agreement, the four error classes and fixed precedence of §8.4, and the §10 limits of limits.go. The conformance claim for this evaluator is stated, in full and only, in CONFORMANCE.md (ADR-0011); nothing in this package restates it, and every payload carries a reference to that file rather than a claim of its own (result.EvaluationClaimReference). The surface may still change or be removed without compatibility promise, which is what "experimental" names here.
Only a pack declaring specVersion 0.2.0-draft is evaluated. §11 makes the declared value exact and says an unedited 0.1.0-draft pack "must be re-declared before an implementation claiming this draft evaluates it", so a pack declaring any other version is refused as pack-not-conformant in the preflight phase (declaredSpecVersion) rather than evaluated under a contract it has not opted into. Every payload still names the contract's own version (result.EvaluatorSpecVersion) beside the pack's, because they are two facts and a consumer should read the applied contract rather than infer it.
Evaluation is a pure function of its inputs: a conformant pack, one facts document, the tri-state availability of declared evidence, and the supported-extension set. Errors are not dispositions: a non-conformant pack, a malformed input, an unsupported required extension, or an exhausted documented limit terminates the evaluation with no disposition at all (§8.4).
The limits this evaluator enforces divide by §10's phase split. Admission is bounded by the carrier limits — 10 MiB per input document, 250,000 parsed nodes, depth 128, 1 MiB strings — and reaching one of those refuses the input rather than processing part of it, which is malformed-input in the preflight phase. The evaluation phase has the two limits §10 requires of the class, stated and derived in limits.go: an evaluation-work limit charged over every condition node, §8 iteration, pointer resolution, and compared byte (DefaultCoreWorkLimit, configurable per evaluation), and a collection-size limit that is the carrier's parsed-node cap because every collection this path traverses comes from a document admitted under it (CoreCollectionSizeLimit). Reaching the work limit is resource-exhaustion, the one class of the evaluation phase, and it is reachable on the ordinary Core path and not only under the draft RFC 0008 opt-in, which has its own smaller budget (ADR-0009).
Index ¶
- Constants
- Variables
- func ArrayIndex(token string, length int) (int, bool)
- func DecimalCompare(fact, operand any) (int, bool)
- func DecimalKey(value any) (string, bool)
- func DecimalString(value *big.Rat) (string, bool)
- func DecimalValue(value any) (*big.Rat, bool)
- func DecodeDisposition(raw json.RawMessage) (result.Disposition, error)
- func DecodeHandoffTarget(raw json.RawMessage) (*result.HandoffTarget, error)
- func OrderedOperator(operator string) bool
- func PointerTokens(pointer string) ([]string, bool)
- func ResolvePointer(document any, pointer string) (any, bool)
- func SameHandoffTarget(expected, actual *result.HandoffTarget) bool
- type AdmittedPack
- type Engine
- func (e *Engine) AdmitPack(pack []byte) *AdmittedPack
- func (e *Engine) Admits(pack []byte, supported []string) bool
- func (e *Engine) Evaluate(pack, facts, evidence []byte, supported []string, command string) (result.Evaluation, *Failure)
- func (e *Engine) EvaluateAdmitted(admitted *AdmittedPack, facts, evidence []byte, options Options) (result.Evaluation, *Failure)
- func (e *Engine) EvaluateWith(pack, facts, evidence []byte, options Options) (result.Evaluation, *Failure)
- func (e *Engine) HandoffTargetRenders() int64
- func (e *Engine) PackHandoffTarget(pack []byte) HandoffTargetRendering
- func (e *Engine) RunCase(pack []byte, item MatrixCase, command string) result.EvaluationCorpusCase
- func (e *Engine) RunCaseAdmitted(admitted *AdmittedPack, item MatrixCase, declared HandoffTargetRendering, ...) result.EvaluationCorpusCase
- func (e *Engine) RunCorpus(specVersion, command string) (result.EvaluationCorpus, *Failure)
- type Failure
- type HandoffTargetRendering
- type MatrixCase
- type Options
Constants ¶
const ( ReasonNotApplicable = "not-applicable" ReasonMissingEvidence = "missing-required-evidence" ReasonUnknown = "unknown" ReasonConflict = "conflict" ReasonNoMatch = "no-match" ReasonExceptionEscalation = "exception-escalation" )
Reason vocabulary of the §8 experiment. exception-escalation is a direct request rather than a trigger-selected one. Exported because the resolver is not the vocabulary's only reader: a project matrix's coverage report derives which of these reasons a pack's own declarations make reachable, and a vocabulary stated in two packages could disagree with itself.
const ComparableFactsCode = "JPS-FACTS-COMPARABLE-REQUIRED"
ComparableFactsCode is the refusal a project that requires comparable facts gives an evaluation in which a fact some comparison of the pack reads has a JSON type that comparison can never match (ADR-0046).
const CoreCollectionSizeLimit = carrier.DefaultMaxNodes
CoreCollectionSizeLimit is this runtime's §10 collection-size limit on the Core path: the carrier's parsed-node cap, which every input is admitted under.
It is a determination rather than a second mechanism, and the determination is the substance. The cap is a budget over one whole parsed document, so no admitted document contains a collection with more than this many members — a collection of n members is n+1 parsed nodes — and the sum of every collection in one document is bounded by it too. Every collection a Core evaluation traverses comes from an admitted document: the facts value a fact condition selects, the candidate array an in condition names, the object members §7.4 equality descends through, and the pack's own rules, exceptions, and evidenceRequirements arrays. Core constructs no collection of its own, and it has no operator that iterates one — the draft RFC 0008 aggregates do, behind their own opt-in and their own budget (ADR-0009).
Because the bound is enforced while admitting each input, §10's phase split makes reaching it malformed-input in the preflight phase — pack-not-conformant for the pack, which §8.4 orders first in any case — rather than resource-exhaustion during evaluation. An evaluation-phase counter over the same bound could therefore only report a condition the preflight has already refused, and refused more strictly: §2.1 rejects the document whole instead of admitting it and then stopping partway through it. What DefaultCoreWorkLimit adds is the half a size bound cannot supply: the *work* that traversing an admitted collection costs, which the collection's size alone does not bound once one condition can read it many times.
const DefaultCoreWorkLimit = int(2 * carrier.HardMaxBytes)
DefaultCoreWorkLimit is this runtime's §10 evaluation-work limit for one evaluation on the Core path, in the work units of the accounting model documented on evaluate. Options.WorkBudget overrides it exactly as it overrides the draft grammar's DefaultWorkBudget, so the limit is configurable per evaluation and the default is what an unconfigured caller gets.
The number is derived in this expression rather than written down: it is exactly twice the carrier's byte cap, which is the only admission limit a work charge is measured against, so there is no independent constant here to drift from carrier.HardMaxBytes. At the current cap of 10 MiB per input it is 20,971,520 units.
A work unit is one visited condition node, one §8 iteration over an authored rule, exception, or evidence requirement, one step of a pointer resolution, or one byte of a path, a member name, or a scalar token a comparison has to read.
The ratio is an arithmetic fact and nothing else: this default is exactly twice the per-document carrier byte cap, computed here from that cap. It buys no guarantee about a whole evaluation, and none is offered:
- Three documents may be admitted, not two — the pack, the facts document, and an optional evidence-availability document — each under the same byte cap, so the admitted bytes of one evaluation can already exceed twice that cap.
- A unit is charged per *use*, not once per admitted byte. The bytes of a pointer and of a selected value are charged again each time a condition reads them, and the per-node and per-iteration charges of §8 are charged on top, so one admitted byte can cost many units and the count is not bounded by the input's length.
An earlier draft of this comment inferred from "every unit is backed by at least one byte of an admitted input" that one full read of every admitted byte always fits, and that a single maximal cross-document comparison therefore sits exactly at the boundary. Both statements are **withdrawn**: the premise bounds nothing, because the charge model re-charges the same bytes and admits a third document, and neither statement was ever exercised against an admitted input through the real accounting path. What remains is the derivation — which keeps the number from drifting from the cap it comes from — and the measured behavior below.
What this limit mainly refuses in practice is amplification: Core conditions have no runtime fan-out, and the one place a small condition can buy large work is the value its pointer selects, once per candidate of an in condition or once per condition over the same large value. "Only amplification is refused" would be false, and it is not claimed either. Against a 100 KB facts document the limit admits roughly two hundred whole-document comparisons; against the values a hand-authored pack compares it is unreachable. Every row of the bundled evaluation corpus charges under 1,000 units, four orders of magnitude inside this limit (TestCoreWorkLimitNeverTripsOnTheCorpus), which is a measurement rather than a bound on inputs no row contains.
const DefaultWorkBudget = 100_000
DefaultWorkBudget is this runtime's §10 evaluation-work limit for one evaluation under the draft grammar, in the work units charged below. The Core path has its own, far larger limit and its own collection-size determination (DefaultCoreWorkLimit, CoreCollectionSizeLimit); this one is smaller because the draft grammar is the only condition form with runtime fan-out, which is what a budget in the tens of thousands has to bound. It is a runtime choice, not a portable one: RFC 0008 leaves the common limit open, so two evaluators may pick different numbers and an above-limit input is not portable between them. The number is set so the RFC's own attack sketch — an aggregate over 10³ elements each carrying an aggregate over 10³ — is refused with room to spare: with the short paths and small values that sketch implies it charges about 6·10⁶ units against this budget of 100,000, a factor of about 60, while every collection a hand-authored pack plausibly quantifies over fits.
It is also this runtime's §10 collection-size limit, which RFC 0008's uplift raises to a MUST alongside the evaluation-work one. No second knob states it, because the work budget already implies it: an aggregate whose where costs c units admits at most (budget - aggregate cost)/c elements, so the cheapest possible where (a one-unit literal) bounds one aggregate at just under 100,000 elements and an ordinary fact predicate over a short pointer, a boolean selected and a boolean authored — six units — bounds it at about 16,000. The document that carries those elements is bounded independently by the carrier layer's byte limit (carrier.HardMaxBytes) and nesting depth (carrier.DefaultLimits), which apply to the facts input before any condition is charged.
const OutcomeValuesExtension = "org.judgmentpack.outcome-values"
OutcomeValuesExtension is the extension name of draft RFC 0016.
Variables ¶
var DecimalReadings atomic.Int64
decimalValue is the one place a value is admitted as a §2.2 decimal and read into a number: the grammar check and the parse together, so every surface that asks about a decimal — the comparison, the identity key — asks the same question of the same tokens. DecimalReadings counts every reading of a value through this grammar -- a comparison reads two, a key or a value one -- so a surface that claims to read each decimal once can be held to it by a test.
Functions ¶
func ArrayIndex ¶ added in v0.17.0
ArrayIndex reads one RFC 6901 array-reference token as an index into a container of the given length — decimal digits with no leading zero other than "0" itself, no sign, no past-the-end "-", and in range — exported for the same single-source-of-truth reason PointerTokens is.
A surface that PLACES a value at a pointer has to admit exactly the tokens the resolution admits. strconv.Atoi is the tempting substitute and it is a wider grammar: it reads "00" and "+0" as zero, so a candidate placed at /items/00 lands at the element the pointer does not address, where this evaluator then resolves nothing and no condition ever reads it (ADR-0024). One rule, stated here, is what keeps a placement and a resolution from disagreeing about which element a pointer names.
func DecimalCompare ¶ added in v0.16.0
DecimalCompare is RFC 0006's pinned ordering as the evaluator applies it, exported so a surface that must judge two decimals the way §7.4 judges them reads this rule rather than restating a second decimal grammar: "5000.0" and "5000" are one value, and a JSON number on either side is not comparable at all. The one-statement discipline ADR-0014 applied to the reason vocabulary, applied to a comparison (ADR-0023).
func DecimalKey ¶ added in v0.16.0
DecimalKey renders one value's §2.2 decimal identity: big.Rat's own canonical string for the number the value denotes, or false for anything §7.4 cannot compare at all — a JSON number, a grouped "5,000", a non-string. Two values share a key exactly when DecimalCompare judges them equal, which is what lets a surface grouping decimals by value fold them in one pass through this grammar (ADR-0023's boundary identity) instead of comparing every pair. It reads the same pattern and the same big.Rat parse DecimalCompare does, so the two cannot drift apart.
func DecimalString ¶ added in v0.17.0
DecimalString renders one exact number as a §2.2 decimal string — -?(0|[1-9][0-9]*)(\.[0-9]+)? — or reports false for a number that grammar cannot spell.
It exists because DecimalKey cannot do this and must not be made to: that function renders big.Rat's canonical form, which is "81/2" for 40.5, so it is an *identity* key and never an emission. A surface emitting a decimal a pack's own literals imply (ADR-0024) needs the emission, and a second decimal writer living outside this package could spell a value the evaluator then declines to compare.
The contract is exactness in both directions. A number whose decimal expansion does not terminate — any value whose denominator in lowest terms keeps a prime factor other than 2 or 5 — is refused rather than rounded, because a rounded value is a different number and a generator that quietly substituted one would emit an input nobody asked for. Where the expansion terminates the rendering carries every digit of it, at the fewest digits that are exact, so the result has no trailing fractional zero and DecimalValue reads back the number this was called with.
Both halves of that contract are about the number and not about the fraction it arrived as, so the value is reduced before it is read. A *big.Rat the arithmetic here produces is already in lowest terms, but this is exported and a caller can hand over one that is not — big.Rat's own Num and Denom return references a caller may set — and 2/4 rendered off an unreduced denominator is "0.50", which states a precision the number does not have, while 3/3 is refused outright for a denominator that clears on sight. Reducing a copy costs one normalization and leaves the caller's value untouched.
func DecimalValue ¶ added in v0.17.0
DecimalValue admits one value as a §2.2 decimal and reads it into an exact number, exported for the same one-statement reason DecimalKey is: a surface deriving new decimals from a pack's own literals — the test-row input generator of ADR-0024 — must admit them through the grammar §7.4 admits, never through big.Rat's own SetString, which reads "1/3", "1e5", and " 5 " as numbers §2.2 has no spelling for. The returned value is this call's own and the caller may do arithmetic on it.
It is the reading half of the pair whose writing half is DecimalString, and the two are exact inverses over every value either admits.
func DecodeDisposition ¶ added in v0.9.0
func DecodeDisposition(raw json.RawMessage) (result.Disposition, error)
DecodeDisposition decodes one §8.3 disposition strictly — unknown members refused — and holds it to every invariant §8.3 states about the disposition alone, through the same canonicalization the row comparator applies. It is exported because the comparator is not the only reader of expected dispositions: the coverage derivation reads them too, and a second, looser decoder could accept as a witness the exact expectation the comparator refuses. One gate, however many readers.
func DecodeHandoffTarget ¶ added in v0.17.0
func DecodeHandoffTarget(raw json.RawMessage) (*result.HandoffTarget, error)
DecodeHandoffTarget decodes one row's expected escalation target: JSON null, which asserts that the evaluation reports no target at all, or an object stating both kind and name. It is exported for the same one-gate reason DecodeDisposition is — the matrix loader and the row comparator must not be able to disagree about what a well-formed expectation is.
It is written against the token stream rather than against a struct, and that is not style. `encoding/json` matches member names **case-insensitively** even under DisallowUnknownFields, so a struct decode reads `{"Kind":"queue"}` as a kind and lets `{"kind":"human-role","Kind":"queue"}` overwrite the exact member with the alias; it also accepts a duplicated member silently, and stops at the end of the first value so trailing JSON rides along unread. A closed shape has to refuse all three, and none of them is expressible as a decoder option.
Both members are required and neither may be empty, because an escalation target §8.1 admits states both and a pack declaring an empty one is refused long before a row could match it. The kind vocabulary is deliberately not restated here: the enumeration belongs to the pack schema, which has already held every evaluated pack to it, and a row naming a kind outside it is reported as the loud mismatch it is rather than by a second copy of a list this package does not own.
func OrderedOperator ¶ added in v0.16.0
OrderedOperator reports whether one authored operator is an ordered comparison. The set is read by a second package and is therefore exported as a function rather than as the map itself: an exported map variable has no write barrier, so any importer could assign into it and silently rewrite §7.4's dispatch — `evaluation.OrderedOperators["equals"] = true` would change what the evaluator compares, from a package that only meant to read the set. A function exports the answer without exporting the ability to change it.
func PointerTokens ¶ added in v0.17.0
PointerTokens splits one authored RFC 6901 pointer into its escape-decoded reference tokens, or reports false for a pointer that is neither empty nor rooted at "/". A surface that must *place* a value at a pointer rather than read one needs the tokens themselves, and ResolvePointer only reads; without this export that surface would carry a second implementation of ~1 and ~0, and a pointer the two escaped differently would place a fact where no condition looks for it (ADR-0024).
The empty pointer selects the whole document and yields no tokens, which is the caller's cue that there is no member to place: a caller placing into a document distinguishes that case rather than treating it as an error here.
func ResolvePointer ¶ added in v0.16.0
ResolvePointer resolves one RFC 6901 pointer against a decoded document exactly as a fact condition resolves its path, exported for the same reason as DecimalCompare: the coverage derivation must locate a fact by the evaluator's own resolution, not by a second pointer walker.
func SameHandoffTarget ¶ added in v0.19.0
func SameHandoffTarget(expected, actual *result.HandoffTarget) bool
SameHandoffTarget is the whole of the target comparison: presence against presence, then each member in full. It is exported because the graph matrix's target assertions (ADR-0032) are decided by this comparison and no other — two comparators would be two chances to disagree.
Comparing the authored strings rather than their capped renderings costs nothing a suite can feel. A row's expectation lives in its matrix document, whose own byte bound — the pack matrix's or the graph rows' — bounds the total length compared across a run by the matrix rather than by the pack, and Go compares a string's length before its bytes, so an assertion of a different length is settled without reading either.
Types ¶
type AdmittedPack ¶ added in v0.16.0
type AdmittedPack struct {
// contains filtered or unexported fields
}
AdmittedPack memoizes the pack half of the §8.2 preflight — full document conformance, the version gate, and the decode — across many evaluations of one pack (issue #78). A suite of ten thousand rows previously re-validated and re-decoded the same bytes ten thousand times; with this, once per distinct capability SET the rows declare: the key is canonical (sorted, deduplicated, collision-free JSON encoding), so two spellings of one set share one admission and no two different inputs can share a key. The admission's outcome, failures included, is replayed identically per row — as a copy, so a caller mutating a returned failure cannot poison the next row — and §8.4's error precedence is unchanged: the pack byte limit is checked per call ahead of the memo exactly where EvaluateWith always checked it, a pack-half failure returns before any facts or evidence document is touched, and unsupported required extensions stay deferred to their §8.2 place.
The decoded packRoot is shared across the rows of one admission; the resolution path reads it and never mutates it, which the corpus and every suite of this repository hold in place. The memo is bounded: past maxAdmissions distinct sets, further sets are admitted without being retained — exactly the old per-row cost, never an unbounded retention of decoded roots. An AdmittedPack is safe for concurrent use.
func (*AdmittedPack) Admits ¶ added in v0.16.0
func (a *AdmittedPack) Admits(supported []string) bool
Admits reports whether this pack reaches §8 under a capability set, from the memo — the raw byte limit stays with EvaluateWith, as it always has.
type Engine ¶
type Engine struct {
// contains filtered or unexported fields
}
Engine evaluates conformant packs. It wraps the validation engine so a pack is always checked to full document conformance before a single condition is interpreted.
func NewEngine ¶
func NewEngine(validator *validation.Engine) *Engine
func (*Engine) AdmitPack ¶ added in v0.16.0
func (e *Engine) AdmitPack(pack []byte) *AdmittedPack
AdmitPack prepares one pack's bytes for evaluation across many rows.
**It takes a snapshot.** The returned value owns a private copy of pack, and every later question — what the pack conforms to, what it decodes to, what its binding digest is — is answered from that copy. A caller may reuse or edit its slice afterwards without changing what was admitted. That is not politeness about aliasing: without it the digest is a promise about bytes somebody else can still rewrite, so a caller could mint a rendering for one pack, admit it, edit the slice into a different pack of the same length, and have a row report the first pack's destination while the evaluation decided from the second — a row passing while its report named a target the evaluated document does not declare, which is exactly what the binding exists to prevent. The copy costs one linear pass, once per pack per run.
The pack byte limit is decided **first**, ahead of the copy and the digest, because it is a bound that existed before any of this and §8.4 makes it a preflight refusal: an oversized pack must be refused rather than copied or hashed. The refusal is recorded once, here, and replayed by every evaluating path, so there is one boundary and no order left to get wrong.
An oversized pack is therefore *not* snapshotted, and the caller's slice is retained instead — the one place the ownership rule above does not apply, and it is safe for the reason the rule exists: no digest is taken for such a pack, so there is no binding for a later edit to defeat, and every evaluating path refuses it before it decodes anything. Retaining it at all is what keeps Admits at its documented semantics, which are conformance and deliberately not the raw byte limit: a conformant pack padded past the limit still Admits, and still refuses to evaluate.
func (*Engine) Admits ¶ added in v0.9.0
Admits reports whether this evaluator would admit the pack itself under one consumer capability set — full document conformance, the one declared specVersion this contract covers, and no required extension outside supported — before any facts or evidence document is consulted. It exists for surfaces that must know whether §8 is reachable at all without running an evaluation: a coverage derivation over a pack the preflight refuses would describe behavior nothing can reach. The capability set is a parameter because admission depends on it: a pack requiring an extension is admitted exactly for the callers that support it. The §8.2 raw byte limit stays with EvaluateWith — every caller of this helper reads through a bounded reader already, and a helper that half-repeated the preflight's resource checks would invite reliance on the half.
func (*Engine) Evaluate ¶
func (e *Engine) Evaluate(pack, facts, evidence []byte, supported []string, command string) (result.Evaluation, *Failure)
Evaluate validates the pack (with the caller's supported extensions), then applies the §§7–8 experiment to the facts and evidence inputs. command names the reporting surface, exactly as in the describe package.
func (*Engine) EvaluateAdmitted ¶ added in v0.16.0
func (e *Engine) EvaluateAdmitted(admitted *AdmittedPack, facts, evidence []byte, options Options) (result.Evaluation, *Failure)
EvaluateAdmitted is EvaluateWith over a pack admitted once for many rows (issue #78). The admission carries the pack half of the preflight — byte limit, conformance, version gate — whose outcome is replayed here before any other input is touched, so the §8.4 precedence a per-row caller saw is byte-identical.
func (*Engine) EvaluateWith ¶ added in v0.3.0
func (e *Engine) EvaluateWith(pack, facts, evidence []byte, options Options) (result.Evaluation, *Failure)
EvaluateWith is Evaluate with the caller's opt-ins stated explicitly.
The §8.2 input preflight runs first and completely: the pack, then the facts document, then the evidence-availability document, then the pack's metadata.requiredExtensions against this caller's supported set. §8.4 makes that order the error precedence, so the first failure here is the class §8.4 requires be reported, and no result can outrace an input error — a pack whose applicability is false, presented with an evidence document carrying an undeclared key, is the malformed-input error and never the not-applicable disposition.
func (*Engine) HandoffTargetRenders ¶ added in v0.17.0
HandoffTargetRenders reports how many escalation targets this engine has rendered. It exists so the bound above is observable rather than asserted: a target name is an authored string §2.1 admits at a megabyte, so the number of times one is canonicalized and hashed is the difference between a bounded run and a matrix that costs gigabytes while staying inside every byte limit.
It counts what this engine minted. A rendering minted elsewhere over the same pack bytes is reported rather than re-minted, and is therefore not counted here; that is the scope of the observation and not a hole in the bound, which is about the loop a run performs.
func (*Engine) PackHandoffTarget ¶ added in v0.17.0
func (e *Engine) PackHandoffTarget(pack []byte) HandoffTargetRendering
PackHandoffTarget renders the escalation target a pack declares. It is the only function in this runtime that renders one, and the only constructor of a non-zero HandoffTargetRendering.
It takes the pack's **bytes** rather than a decoded root for two reasons that are the same reason: the bytes are what the binding digest is over, and taking a caller-supplied map would let a rendering be minted from a document no pack ever was. It applies the shared byte limit before decoding, through the same helper the evaluation path uses, so an oversized pack is refused here as it is everywhere — nothing is scanned that a §8.4 preflight would refuse.
It is a pure function of those bytes, with one exception that is not state the answer depends on: the Engine counts the renderings it has produced, because "once per pack per run" is a claim this record makes and a claim nothing can observe is a claim nobody can hold. Since this is the sole rendering site, that count is complete for renderings this engine minted. See HandoffTargetRenders.
func (*Engine) RunCase ¶ added in v0.4.0
func (e *Engine) RunCase(pack []byte, item MatrixCase, command string) result.EvaluationCorpusCase
RunCase runs one case-carrier row against one pack document and reports the row's result.
A row carrying an expected disposition passes when the disposition this evaluator produces canonicalizes, under RFC 8785, to the same bytes as the row's. A row carrying an expected §8.4 error class instead passes when the evaluation is refused with that class, and with that phase when the row names one. That is the whole comparison, and it is the same one whether the row came from the bundled corpus or from a project's own matrix (ADR-0012): a project gets the byte comparison §8.3 defines rather than a looser one written for it.
A row may state one further expectation, and only one further expectation: expectedHandoffTarget, which is compared against the target reported beside the disposition rather than inside it (ADR-0025). It is optional, it is an assertion and not a report, and a row that omits it is judged exactly as it was before the member existed.
command names the reporting surface, exactly as elsewhere in this package.
func (*Engine) RunCaseAdmitted ¶ added in v0.16.0
func (e *Engine) RunCaseAdmitted(admitted *AdmittedPack, item MatrixCase, declared HandoffTargetRendering, command string) result.EvaluationCorpusCase
RunCaseAdmitted is RunCase over a pack admitted once for the whole suite (issue #78): the same judgment, without re-validating and re-decoding the pack bytes for every row.
func (*Engine) RunCorpus ¶ added in v0.4.0
func (e *Engine) RunCorpus(specVersion, command string) (result.EvaluationCorpus, *Failure)
RunCorpus runs the evaluation corpus bundled for one exact specification version through this evaluator and reports every row's result.
A row carrying an expected disposition passes when the disposition this evaluator produces canonicalizes, under RFC 8785, to the same bytes as the row's — the byte agreement §8.3 requires of two implementations, applied to one implementation and one pinned expectation. A row carrying an expected §8.4 error class instead passes when the evaluation is refused with that class, and with that phase when the row names one.
A failing row decides nothing by itself: §3.4 makes a divergence as likely to be a defect in the row as in the implementation, and neither this runtime nor this function adjudicates that question.
type Failure ¶
Failure reports why an evaluation could not run. It is an operational refusal, never a disposition.
Class and Phase are the JPS Core §8.4 identity of the failure when it is an evaluation error: one of the four Core classes, and the phase it was reached in. Code stays this runtime's finer-grained diagnostic identifier, which §8.4 admits as message detail rather than as the class. A failure that §8.4 does not classify at all — a bad invocation, an internal fault — carries no class.
type HandoffTargetRendering ¶ added in v0.17.0
type HandoffTargetRendering struct {
// contains filtered or unexported fields
}
HandoffTargetRendering is one pack's declared escalation target, rendered for the report, computed **once per pack per run** by whatever owns the row loop and handed to every row of that pack (ADR-0025).
It is a value and not a cache, and that is the whole design. Three earlier drafts tried to make the rendering cheap where it was needed — per row, then memoized on the target's content, then stored on a capability admission — and each was a bound on the wrong thing, because each left the *decision about when to compute* on the row path. A rendering is a function of the pack's bytes alone: §8.1 gives a pack one escalation target, every admission of one pack decodes the identical declared target, and a capability set has nothing to do with it. So it is computed where a pack is loaded, which is a place visited once, and carried down. The row path then has no cache to miss, no admission to re-enter, no mutex, and no atomic — not as a claim about a fast path, but because there is no state there to reach.
**Its members are unexported and it has no constructor but PackHandoffTarget**, which is the one place a target is rendered and the one place the count is taken. That is what makes "the row path renders nothing" and "PackHandoffTarget is the sole rendering site" true by construction rather than by convention: a caller of RunCaseAdmitted cannot hand it a rendering it made itself, cannot hand it stale bytes for a target the pack does not declare, and cannot reach a path that would render one per row behind the counter's back.
The zero value is legal and means exactly "no rendering was computed for this pack". A row that asserts a target and is given it reports the result.HandoffTargetUnavailable convention on the actual side rather than having one rendered for it — the report degrades and says so, and the verdict is unaffected, because the verdict rests on the decoded values (see SameHandoffTarget). No in-tree caller reaches that state: the project runner computes a rendering whenever a row asserts, RunCase computes its own, and no bundled corpus row can assert at all.
**It also carries the SHA-256 of the pack bytes it was minted from, and a row uses it only when that digest is the digest of the pack it evaluated.** Opacity stops a caller *fabricating* a rendering; it does not stop one *transplanting* a genuine rendering minted for another pack. Without the binding, such a rendering would be reported verbatim while the verdict compared the evaluated pack's real target — a row passing while its actualHandoffTarget named a destination the pack never declared, which is the failure this whole record exists to prevent, inverted. With it, a rendering that does not belong degrades to "unavailable": the report says it cannot state the target rather than stating a false one.
The binding is a **digest of the pack's bytes** and not a comparison of the decoded targets, and the difference is the whole of what a review round cost. Comparing decoded targets is exact, but both of its operands come from the *pack*, so a matrix asserting on ten thousand rows scans the pack's target ten thousand times — the resource shape this record has rejected three times over, reappearing in the machinery meant to protect the report. R2-1's per-row comparison is bounded because one operand comes from the *row*, and the matrix bounds the rows; that proof does not transfer to two pack-derived operands, and a draft of this determination claimed it did. A digest comparison is thirty-two bytes whatever the pack weighs.
Same bytes mean the same declared target, so a rendering minted by another engine over the same pack reports honestly rather than degrading. What that costs is observation scope and not correctness: Engine.HandoffTargetRenders counts what *this* engine minted, so a run reusing another's rendering sees a lower count than mints performed. Documented rather than fought — the number is there to hold one run's own loop to its bound.
func (HandoffTargetRendering) Present ¶ added in v0.17.0
func (h HandoffTargetRendering) Present() bool
Present reports whether a rendering was computed — that is, whether the pack declares an escalation target §8.1 admits and rendering it succeeded.
type MatrixCase ¶ added in v0.4.0
type MatrixCase struct {
ID string `json:"id"`
Origin string `json:"origin"`
Facts json.RawMessage `json:"facts"`
EvidenceAvailability json.RawMessage `json:"evidenceAvailability"`
SupportedExtensions []string `json:"supportedExtensions"`
ExpectedDisposition json.RawMessage `json:"expectedDisposition"`
ExpectedHandoffTarget json.RawMessage `json:"expectedHandoffTarget,omitempty"`
ExpectedErrorClass string `json:"expectedErrorClass"`
ExpectedErrorPhase string `json:"expectedErrorPhase"`
Focus string `json:"focus"`
SpecSection string `json:"specSection"`
// Cites is the second project-only extension (ADR-0034), on the same
// footing as ExpectedHandoffTarget: the receipts a row's facts were
// transcribed under, in the gateway's citation shape, reachable only from
// a project matrix declaring matrixVersion 3, held to its grammar by the
// project loader, which carries the values into the row's result. The
// comparator never reads it; the bundled corpus's schema closes its case
// object and never carries it.
Cites json.RawMessage `json:"cites,omitempty"`
}
MatrixCase is one case-carrier row: one pack, one facts document, the tri-state availability of its declared evidence, the consumer's supported extensions, and exactly one expectation — a §8.3 disposition to compare byte for byte, or the §8.4 error class the evaluation must be refused with.
It is exported because the bundled evaluation corpus is not the only carrier of this shape: a project's own instance matrix (ADR-0012) uses the same rows and is compared by the same code, so a matrix a builder writes and the corpus this runtime ships are read, run, and judged identically rather than by two implementations of one comparison.
facts, evidenceAvailability, and expectedDisposition are kept as raw bytes: the evaluator takes its inputs as documents, and re-encoding them here would put a second serializer between the carrier and the evaluation.
ExpectedHandoffTarget is the one optional second expectation (ADR-0025), and it is raw bytes for a reason the other members do not have: the assertion has three states and only raw bytes distinguish them. Absent asserts nothing; the literal null asserts that the evaluation reports no target; an object asserts the target it names. It carries `omitempty` for the same reason — without it, marshaling a row that asserted nothing writes `expectedHandoffTarget: null`, which reloads as the assertion that the evaluation reports no target, so a round trip through this type would invent an expectation nobody wrote.
The member is a **project-only extension** of what the two carriers share, which is the fields this comparator reads — not a claim that a row moves between them untouched. The bundled corpus never carries it, because that manifest's schema closes its case object; the same schema separately requires pack, origin, supportedExtensions, focus, and specSection, which a project row need not declare. So lifting a project row into a corpus means supplying those and dropping any target assertion. The member is reachable only from a project matrix, which is the carrier whose rows are written about a pack the project itself maintains, and only from matrixVersion 2 (project.MatrixVersion).
type Options ¶ added in v0.3.0
type Options struct {
// Command names the reporting surface, exactly as in the describe package.
Command string
// SupportedExtensions is the consumer's extension capability set.
SupportedExtensions []string
// RFC0008Quantifiers admits the draft RFC 0008 aggregates (exists, every,
// uniform). A pack using one is not valid under any published JPS version,
// so every evaluation made under this opt-in is labeled a draft-RFC
// prototype in its output.
RFC0008Quantifiers bool
// RFC0016OutcomeValues admits the value declarations of the draft RFC 0016
// and resolves them once an outcome is produced (ADR-0039). A pack carrying
// one is not valid under any published JPS version, so every evaluation made
// under this opt-in is labeled a draft-RFC prototype in its output. One
// evaluation runs under one draft: a caller that sets this and
// RFC0008Quantifiers is refused before any input is decoded.
RFC0016OutcomeValues bool
// WorkBudget overrides this evaluation's §10 evaluation-work limit. Zero or
// negative selects the default for the path in force: DefaultCoreWorkLimit on
// the Core path, and the draft grammar's smaller DefaultWorkBudget under the
// opt-in. A caller that configures a lower limit has configured its own limit,
// and an input above it is refused with the resource-exhaustion error of §8.4
// rather than evaluated to a disposition.
WorkBudget int
// EvidenceSupplied says the caller supplied an evidence-availability document,
// which is what tells an empty document apart from an omitted one. §8.2 makes an
// omitted document the implicit empty object and not an error; empty bytes are
// not a carrier-conforming JSON text and are the malformed-input error. A caller
// whose surface has no way to supply an empty document leaves this false.
EvidenceSupplied bool
// OversizedInputs names the inputs — "pack", "facts", or "evidence" — the
// caller could not present in full because they exceed this runtime's
// documented byte limit. A caller reports the condition here instead of
// refusing it itself, so the byte limit is reached at that input's own place
// in the §8.2 preflight and §8.4 assigns the class and the order: the pack's
// limit before the facts document's, and a non-conformant pack before either.
OversizedInputs []string
// RequireComparableFacts refuses, once the §8.2 preflight has admitted
// every input and before §8 begins, an evaluation in which a fact some
// comparison of the pack reads is present and of a JSON type that
// comparison can never match (ADR-0046). It is not part of the evaluator
// contract: the refusal carries no §8.4 class, and only a deciding surface
// whose project sets requireComparableFacts asks for it. Off, nothing
// changes.
RequireComparableFacts bool
}
Options carries one evaluation's caller opt-ins. The zero value is the published surface: Core conditions only, no draft grammar, no work budget.