Skip to main content

Glossary

This is the canonical vocabulary for Agent Vigilo documentation. Inline terms throughout the site open the same definitions and link back to their entries here.

Code identifiers and stored values appear in code style. For example, a run can have operational status = completed and policy gate_status = fail at the same time.

Runtime Roles​

Vigilo CLI
Command entry point for evaluator publishing, run management, shard administration, and runtime service modes.
A command can perform work directly or run a coordinator or worker process.
Coordinator
Process that advances durable creation recovery, lease recovery, dispatch, finalization, and outbox publication.
Many coordinators can run concurrently because database claims divide work.
Worker
Process that consumes chunk-ready broker messages, claims chunks, invokes the agent target and evaluators, and persists results.
Many workers can run concurrently; a chunk claim selects one current owner.
Bounded work
Work limited by a configured item count, concurrency count, time budget, or retry budget.
Prevents one cycle, placement, chunk, or evaluator from consuming unlimited resources.
C4 container
Independently runnable process, data store, or infrastructure service in an architecture deployment view.
A C4 container is a deployment boundary, not necessarily an operating-system container.
Command flow
Decisions and side effects produced by one CLI command.
One flow follows one command from input through persistence and external calls.
TOON output
Compact structured CLI encoding intended for agent inspection and language-model tool workflows.
Select with -q -f toon; use -f json when exact standard JSON parsing is required.

Evaluation And Results​

Agent target
HTTP service or workflow whose behavior is being evaluated. It may wrap a model, prompt pipeline, or multi-step agent.
One versioned target configuration per run profile.
Run profile
Versioned configuration for the agent target, evaluator bindings, scoring, persistence, retries, and gate behavior.
One selected profile per run, replicated into immutable shard-local run snapshots; distinct from performance and evaluator artifact build profiles.
Dataset
Versioned collection of evaluation cases.
One dataset version per run.
Test case
One immutable case input, optional expected output, and routing metadata from a dataset.
A test case can be evaluated once in a run.
Case input
Dataset case data sent to the configured agent target.
One required input value per test case; distinct from the complete evaluator input envelope.
Expected output
Optional oracle data available to evaluators but not sent to the agent target.
Zero or one expected value per dataset case.
Agent output
Captured actual result from the agent target, including optional text or structured data, tool calls, trace, raw provider output, and metadata.
Passed to evaluators as input.actual; distinct from expected output.
Evaluator input
Versioned ABI input envelope containing run, execution, and attempt identity, the test case, actual agent output, and evaluator-specific configuration.
One input envelope per evaluator binding invocation.
Evaluator output
Successful ABI output envelope containing evaluator identity, a completed or abstained outcome, diagnostics, and metadata.
Evaluator errors are returned outside the output envelope.
Task type
Case label used by automatic run-profile matching.
One task type per dataset case.
Tags
Case labels used by tags_any and tags_all profile matching rules.
Zero or more tags per dataset case.
Case group
Run-profile rule that selects evaluator bindings and aggregation policy for matching cases.
One explicit group or one or more automatically matched groups per case.
Evaluator
Versioned Wasm component that examines an agent output and returns one measurement or abstention plus optional diagnostics.
Identified by <namespace>/<name>:<version>.
Evaluator binding
Stable profile entry that assigns an evaluator measurement to host-owned normalization, threshold, requiredness, dimension, weight, and blocking policy.
Identified by evaluators[].id; many bindings can apply to one case.
Evaluation plan
Per-case resolution of matching case groups, evaluator bindings and configuration, dimensions, and aggregation policy.
Distinct from the run-wide evaluator execution plan that pins artifacts, ABIs, adapters, runtime, and policy hash.
Evaluator identifier
Immutable published identity in <namespace>/<name>:<version> format.
One identifier selects one versioned evaluator artifact.
Measurement
The single raw observation returned by a completed evaluator invocation.
Binary, numeric, or ordinal; the profile explicitly maps it to utility and judgment.
Normalization policy
Profile-owned mapping from a raw evaluator measurement to a score between 0.0 and 1.0.
Binary, numeric linear, numeric curve, numeric threshold, or ordinal mapping; invalid values are rejected.
Normalized score
Host-derived utility from 0.0 through 1.0 produced by applying one evaluator binding's normalization policy to a completed raw measurement.
One score per completed binding result; errors and abstentions have no score.
Judgment
Host-derived passed or failed result from comparing a normalized score with its evaluator binding threshold.
Only completed measurements receive a judgment.
Evaluator outcome
Persisted invocation state: completed, error, or abstained.
The ABI output carries completed or abstained outcomes; an evaluator-error is returned outside that output and persisted as error.
Diagnostic finding
Non-authoritative evaluator observation with severity, category, reason, evidence, and tags.
Zero or more per invocation; diagnostics cannot score or block.
Evaluator completeness
Execution-level check that every required binding produced exactly one valid, normalized measurement.
Errors, abstentions, missing results, duplicates, or invalid measurements withhold authoritative scores.
Dimension
Profile-owned scoring bucket, such as format or quality.
Host-normalized binding results are grouped into dimensions before total scoring.
Dimension score
The min_score or weighted_mean result for one dimension of one execution.
Zero or one score per configured dimension and execution.
Aggregate score
Weighted total of an execution's dimension scores.
Zero or one total score per execution.
Run scorecard
Authoritative run-wide dimension and evaluator gate results merged from shard-local counters.
One immutable scorecard per completed run; includes coverage, score, error, abstention, and pass-rate metrics.
Scorecard gate
Run-wide rule over one dimension or evaluator binding and an optional case-group or tag slice.
Evaluated from merged shard counters; a violated threshold or required slice with no matches fails the run gate.
Blocking result
Host-derived failed binding result that can fail an execution independently of its aggregate score.
Blocking comes only from evaluator binding or dimension policy.
Execution
Durable evaluation of one dataset case against the agent target.
One expected execution per case in a run.
Attempt
One worker's effort to complete an execution. Retries create later attempts.
Many attempts can belong to one execution; only the current attempt is authoritative.
Current attempt
Attempt ID and number selected by an execution as its current worker effort.
Terminal writes also require the matching worker ID and a live attempt lease; zero or one attempt is authoritative per execution.
Run
One durable evaluation of a dataset version with a run profile and agent target.
A run contains chunks and expected executions.
Cardinality
Exact number of items in a set or the multiplicity of a relationship, such as one run containing many chunks.
In performance workloads, cardinality is the declared input size for a scaling dimension, such as cases, chunks, or events.
Run status
Operational lifecycle state: creating, pending, running, finalizing, completed, failed, or cancelled.
One current value per run.
Gate status
Policy outcome such as unknown, pass, or fail.
One current value per run; distinct from run status.

Work And Routing​

Chunk
Bounded range of dataset cases processed under one worker claim.
Many chunks per run; each chunk belongs to one run shard.
In-flight chunk
Chunk-ready broker delivery currently being processed by one worker process.
Bounded per worker process by max_inflight_chunks.
Prefetch
RabbitMQ limit on unacknowledged deliveries reserved by one consumer.
Configured per worker consumer.
Chunk parallelism
Number of case executions processed concurrently inside one claimed chunk.
Bounded independently from in-flight chunk count.
Run shard
Stable logical segment numbered 0..127 and stored as run_shard; it keeps a chunk and its execution-owned rows together.
A run uses only the shards assigned to its chunks.
Run snapshot
Immutable execution-database copy of the run context required for shard-local worker execution.
One snapshot per used run_id + run_shard.
Run shard summary
Bounded shard-local progress and scorecard rollup used by status, results, and finalization.
One summary per used run shard avoids central scans of execution-owned rows.
Run creation plan
Durable, non-dispatchable control record used to seed exact shard-local chunks and cases and resume interrupted multi-database creation.
One creation plan per creating run until all shard-local materialization is verified.
Control database
PostgreSQL role that owns global run state, placement metadata, dispatch cursors, creation plans, and control outbox records.
Exactly one active control-capable database placement.
Execution database
PostgreSQL role that owns shard-local chunks, snapshots, executions, attempts, results, summaries, and chunk-ready outbox records.
One or more shard-capable placements; the control database may also serve this role.
Database alias
Stable name such as primary or shard_001 used instead of a connection URL.
One alias per database placement.
Database placement
Catalog entry that maps a database alias to a secret environment-variable name, role, and status.
One row per configured PostgreSQL target.
Database placement status
Admission lifecycle for a PostgreSQL target: provisioning is registered but non-routable, active accepts and serves ownership, draining serves existing ownership only, and disabled serves none.
One status per database placement; activation verifies readiness before routing.
Placement drain
Guarded transition that stops new shard ownership before routes are moved away and a database placement is disabled.
The drain does not move rows by itself.
Database router
Process-local DatabaseRouter that reads placement metadata and resolves control or execution pools.
One lazily initialized router per Vigilo process; it does not choose new shard assignments.
Database circuit breaker
Process-local admission guard that temporarily skips one unavailable database alias without changing durable routing.
One independent circuit per contacted execution database alias and process.
Shard placement
Control-plane mapping from run_id + run_shard to a database alias, lifecycle, route version, and write epoch.
One row per used run shard.
Shard placement lifecycle
Route movement phase: active, copying, draining, or moving.
Distinct from database placement status and execution-database local shard admission state.
Route hint
Message-carried database alias and write epoch used as a fast path to an execution database.
Local admission validates the hint, so it is not durable routing authority.
Execution route
Resolved shard placement plus the PostgreSQL pool for its current database alias.
Resolved for one run_id + run_shard.
Route version
Monotonically increasing control-plane CAS generation changed by every route alias or lifecycle update.
One current value per shard placement; not a schema or deployment version.
Write epoch
Monotonically increasing execution-ownership generation carried by routed work and validated in the destination database.
Changes only when ownership moves or is restored.
Local shard admission
Execution-database authority row containing the accepted write epoch and open, draining, prepared, or closed state.
One row per locally known run_id + run_shard; checked in the write transaction.
Dispatch cursor
Control-database progress for dispatching one run shard.
One cursor per used run shard after creation; drained forbids further dispatch.
Chunk dispatch window
Bounded set of pending chunks selected from one run shard in one dispatch operation.
One window can create one run.chunk.ready outbox event record per selected chunk.
Coordinator cycle
Ordered iteration of creation recovery, lease recovery, chunk dispatch, finalization, and outbox publication.
Repeats for coordinator start; runs once for coordinator once.
Coordinator pass
One bounded stage within a coordinator cycle, such as dispatch or outbox publication.
A pass can visit multiple database aliases.
Shard move
Targeted relocation of one run_id + run_shard route and its shard-owned rows to another database alias.
One run shard per move operation.
Rebalance plan
Persisted set of targeted shard moves for a capacity or placement-drain operation.
One plan contains many rebalance items.
Rebalance item
Claimable plan item for moving one specific run_id + run_shard.
One shard move per item; concurrent apply processes can claim different items.

Ownership And Concurrency​

State
Persisted lifecycle value used to determine which transitions are valid.
One current lifecycle value per stateful record.
Transition
Guarded database change from one state to another.
A transition applies only when its authority and current-state predicates hold.
Owner
Process or claim that currently has guarded authority to perform a state transition.
Ownership is temporary unless represented by durable placement state.
Claim
Successful transition that gives a process temporary authority over one work item.
Examples include chunk, dispatch-cursor, outbox-delivery, and rebalance-item claims.
Lease
Time-bounded claim authority that becomes recoverable after its deadline.
Expiry permits recovery but does not alone prevent a stale write.
Claim token
Opaque value issued with a claim and required to settle or renew that exact claim.
A newer claim gets a different token, fencing the previous owner.
Fencing token
Value whose equality proves that an owner or route is still current.
Checked on protected mutations; claim tokens, route versions, and write epochs fence different ownership boundaries.
Row lock
PostgreSQL lock on selected table rows, usually held until the transaction ends.
PostgreSQL enforces conflicting row-lock modes.
Admission lock
Transaction-scoped PostgreSQL advisory lock used cooperatively by shard writers.
Normal writers take the shared form; shard movement takes the exclusive form for one run_id + run_shard.
Route fence
Expected database alias, placement status, and route version validated immediately before a routed write.
One expected fence per resolved execution route.
Route CAS
Compare-and-swap update that changes a route only while its stored alias, lifecycle, and route version still match.
Serializes concurrent control-plane route changes.
Compare-and-swap
Update that succeeds only if stored values still equal expected values.
Used to change a route or settle a claim without overwriting newer state.
Idempotent operation
Operation that can be repeated without duplicating its logical effect.
Retries may execute more than once while producing one durable outcome.
No-op
Valid path that changes nothing because another process already advanced the state.
A no-op is a successful convergence outcome, not necessarily an error.
Recovery
Reassignment or repair after a lease expires or a durable workflow stops mid-operation.
Recovery preserves retry limits and invalidates stale owners.
Stale attempt
Attempt that is no longer authoritative because its lease expired, its chunk was recovered, or a later attempt superseded it.
Retained for history but rejected as a current writer.

Events And Messaging​

Outbox event record
Durable outbox_events row inserted in the same database transaction as the state change it describes.
One logical event per unique dedupe_key.
Outbox delivery row
Temporary publish work in outbox_delivery_queue.
One active delivery row per unpublished outbox event record.
Publish claim
Time-bounded ownership of an outbox delivery row, fenced by claim_token.
One current publisher claim per delivery row.
Broker
RabbitMQ transport that carries published messages from coordinators to workers.
Transport is at-least-once; durable authority remains in PostgreSQL.
Broker circuit breaker
Process-local availability guard that pauses RabbitMQ operations after transport failures and probes for recovery.
Keeps broker outages separate from message retry budgets and durable database authority.
Broker message
RabbitMQ message created from an outbox event record after publication.
May be delivered more than once.
Worker delivery
One broker delivery that identifies a chunk for a worker to claim.
A delivery is acknowledged, delayed, requeued, or quarantined after processing.
Event type
Semantic event name such as run.started, run.chunk.ready, or run.completed.
Stored on the outbox record and used for broker routing.
Dedupe key
Stable identity for one logical outbox event record, also published as the AMQP message ID.
Unique in the outbox ledger.
At-least-once delivery
Delivery guarantee that permits redelivery after uncertain acknowledgement.
Consumers must use database claims and idempotency guards.
Publisher confirm
RabbitMQ acknowledgement that a published message reached the broker.
Required before the outbox event record is marked published.
Message settlement
Worker acknowledgement, delayed redelivery, or requeue decision for one broker delivery.
One settlement outcome per received delivery.
Quarantine
Run-owned holding path for invalid or retry-exhausted worker deliveries that must not re-enter normal processing.
Retains the failed delivery for diagnosis without treating it as completed work.

Performance Testing​

Performance harness
Repository-local cargo perf tooling that builds test subjects, executes versioned workloads, validates correctness, measures supported boundaries, and writes evidence.
Lives in xtask and performance; production Vigilo does not depend on it.
Performance campaign
One bounded cargo perf run or cargo perf compare invocation over a resolved performance profile and one or two test subjects.
Its campaign ID and artifacts are separate from a Vigilo evaluation run.
Performance profile
Versioned campaign composition selecting exact workload tuples, block counts, timing mode, schedule, and resource limits.
Configures the performance harness; distinct from an evaluation run profile and an evaluator artifact build profile.
Performance workload
Versioned contract for one supported behavior, including its measured region, allowed tuples, fixture, exact oracle, metrics, limits, and binary capability.
A profile selects workloads; it does not implement them.
Workload tuple
One named, exact fixture and configuration shape allowed by a performance workload.
Workload ID plus tuple ID identifies one comparison, budget, or model point; tuples are not command instructions or an expanded Cartesian product.
Exact correctness oracle
Required post-execution validation of the expected process, output, service, and durable effects of a performance sample.
A sample contributes timing only after its oracle passes; this is distinct from a dataset case's optional expected output.
Performance test subject
Exact release executable and frozen setup assets produced together by cargo perf build for later measurement.
Created outside the measured interval and reused by run or compare; creating it does not run a performance test.
Performance build manifest
Versioned provenance and compatibility record for a test subject's executable digest, source revision, toolchain, dependencies, setup assets, and workload capabilities.
Stored as build-manifest.json; distinct from Cargo and evaluator manifests.
Measured sample
One scheduled execution with process, resource, and exact external observations plus a validation classification.
Readiness and preconditioning observations are recorded separately and excluded from timing statistics.
Counterbalanced block
Four baseline/candidate executions ordered as ABBA or BAAB to estimate change while controlling position effects.
Opposite orientations are paired; blocks, not executions inside one block, are the independent sampling basis.
Informative timing
Measured performance evidence reported without a numerical regression gate.
Useful for diagnosis and canaries, but it cannot establish that performance is unchanged.
Performance budget
Reviewed practical maximum harmful relative effect for one canonical environment, workload tuple, and metric, with a required minimum block count.
Noise calibration tests whether the host can resolve the budget; calibration does not invent or raise it.
Gating profile
Performance profile whose workloads all use gating timing and resolve exact entries from one published budget policy.
Only compatible canonical evidence can produce its numerical regression verdicts.
Noise calibration
Canonical same-build comparison used to estimate environmental and measurement variation, repeatability, power, and required block counts.
Supports review of practical budgets; it remains separate from capacity calibration.
Canonical environment
Externally validated exclusive host satisfying a versioned performance environment contract.
Only matching canonical evidence may support published blocking performance verdicts.
Performance baseline
Immutable provenance index binding reviewed noise and capacity evidence, a budget policy, a gating profile, and one build manifest.
Distinct from the baseline executable role in an individual A/B comparison.
Performance verdict
Disposition derived from correctness, evidence validity, confidence bounds, and any applicable budget.
Values are pass, regression, improvement, informative, inconclusive, or invalid; a first over-budget interval is inconclusive until an independent matching confirmation.
Capacity staircase
Single-build campaign that holds worker count fixed while increasing offered load through registered steps.
Measures bounded one- and two-worker behavior; it is separate from fixed-load A/B regression testing.
Worker knee
First valid load step where throughput flattens while latency grows materially, or normalized per-worker CPU reaches its reviewed ceiling.
Shared-service saturation invalidates the point; if no knee appears, only the highest observed rate lower bound is reported.
Component scaling model
Accepted fixed-cost-plus-slope or explicit stepped relationship fitted from repeated valid component samples.
Incomplete repetitions, exact-count drift, negative coefficients, nonlinearity, or excessive residuals reject the model.
Model point
Registered mapping from one workload tuple to a positive scaling input and exact external observations.
Every tuple in a modeled workload must have exactly one point.
Capacity projection
Estimate that combines bounded measured capacity with a named deployment's workload, topology, amplification, and independently sourced limits.
Confidence is invalid, directional, planning, or calibrated; it does not claim unmeasured fleet capacity.
Amplification
Ratio of attempts, deliveries, events, acknowledgements, retries, or other external work to intended useful work.
Exact and bounded amplification detects duplicate or runaway work independently from timing.
Reliability run
Single-build operational campaign that checks sustained progress or controlled dependency recovery with resident processes.
Soak and recovery evidence uses safety bounds and does not enter fixed-load A/B statistics.
Performance diagnostics
Post-timing PostgreSQL statement, planning, buffer, and WAL evidence rendered for investigation.
Diagnostics never change correctness, comparison, budget, or component-model verdicts.

Evaluator Packaging​

WIT
Wasm Interface Type language used to define Component Model interfaces and worlds.
Versioned evaluator contracts live under wit/evaluator/<version>/evaluator.wit and become immutable when released.
WIT world
WIT boundary that groups the evaluator's imported and exported interfaces.
The evaluator implements evaluator-world.
Evaluator ABI
Exact versioned WIT binary contract implemented by an evaluator WebAssembly component.
Identified by package, world, interface, version, and immutable contract hash.
Evaluator host adapter
Version-specific host binding that validates, invokes, and maps one supported evaluator ABI.
Selected from the run's immutable execution plan.
Evaluator execution plan
Hashed run snapshot of the exact evaluator ids, artifact hashes, ABI identities, adapters, runtime versions, and scoring policy hash.
Frozen once per run and verified by every worker placement.
WASI 0.2 (Preview 2)
Stable Component Model-based WASI release targeted by evaluator artifacts.
Rust target wasm32-wasip2; Preview 2 is the earlier name for WASI 0.2.
Evaluator artifact
Compiled WebAssembly component stored in the evaluator registry.
One immutable artifact content per evaluator identifier.
Evaluator registry
Durable catalog of published evaluator identities, metadata, contracts, and artifact content.
Contains many versioned evaluators.
Evaluator state
Registry lifecycle value: active, yanked, deprecated, disabled, or removed.
One current state per published evaluator identity.
Package manifest
Cargo.toml file that supplies evaluator crate identity and version.
One Cargo manifest per evaluator crate.
Evaluator manifest
Vigilo.toml file describing artifact paths, WIT expectations, and publish metadata.
One per evaluator package.
WIT contract
Immutable versioned interface definition that a compiled evaluator component must implement.
Validated by declaration, contract hash, and typed component linking.
Build profile
Named artifact selection in Vigilo.toml, such as dev or release.
Selects a build output; not an evaluation run profile.
Wasm store
Fresh Wasmtime execution state created for one evaluator invocation.
One isolated store per invocation.
Fuel
Deterministic Wasmtime instruction budget for one evaluator invocation.
Exhaustion interrupts evaluator execution.
Evaluator semaphore
Process-local cap on concurrently active Wasm evaluator invocations.
One shared semaphore per worker process.
Publish
Validate and insert a versioned evaluator artifact into the evaluator registry.
Publishing is immutable for one evaluator identifier.