Performance Testing
Vigilo's
is an external xtask package. Production Vigilo does not depend on the
harness, while the harness measures only supported CLI, database, RabbitMQ,
HTTP, and evaluator boundaries. Every
must pass its
before it can contribute to statistics.
A resolves one performance profile against one or two test subjects and verifies each subject's before execution.
Component Boundaries
Versioned exercise release-binary run creation, dispatch, HTTP agent, Wasm evaluation, worker persistence, outbox publication, lease recovery, and run finalization. select explicit around page, batch, and concurrency boundaries; the harness never expands dimensions into a Cartesian product. Exact HTTP requests, worker deliveries, and durable rows catch lost batching or before timing is interpreted.
Administration workloads cover cancellation fanout, terminal status/results, JSON and JSONL export, shard move/rebalance, and coordinator placement scaling. They use logical routes backed by isolated databases on one PostgreSQL server, so route growth is measured without pretending to reproduce a database fleet. Creation and HTTP variants also schedule the remaining page/group boundaries and a larger response payload.
Each uses either a sample-level fixed-cost-plus-slope fit or explicit stepped estimates at known batch boundaries. Every requires repeated valid samples, exact observations, and residuals within its registered tolerance.
PostgreSQL normalized statements, planning time/counts, shared/temp buffers,
and WAL counters are captured after the timed process exits. The
cargo perf diagnose command renders them as
.
The model command writes the owned component-model artifact and fails closed on
incomplete or nonlinear evidence.
Large-Data Contracts
Read and export fixtures contain 250 or 251 completed executions, evaluator results, and deterministic diagnostics across one or two routes. Process measurements record total output bytes, time to first and last stdout byte, and peak RSS. JSONL is checked as streamed typed records with a 256 MiB fixture limit; JSON is checked as one materialized document with a separate 512 MiB tested limit.
Shard movement persists the actual rows, bytes, and completed page count for each copied table. A lower bound derived from 1,000 rows or 4 MiB per page is validated, but it is never substituted for the recorded distribution. The move command is then replayed and must report the same verified target without moving it again. Rebalance samples apply one claimed item per resumable pass, require every persisted item to complete verification, and require every selected route to point at the target. Database creation, migrations, route registration, terminal fixture generation, and rebalance planning occur before collector reset.
Criterion remains conditional on a legitimate production library surface. Runtime and database modules are currently binary-private, so the release executable remains the sole component measurement authority.
Calibration Boundaries
calibration-v1performs by comparing one immutable build to itself in 30 for representative CPU, database, broker, Wasm, and end-to-end anchors.- Noise analysis checks confidence bounds and residual orientation effects, then uses the 95% interval width and reviewed power target to recommend an even independent block count for each wall-time budget.
capacity-v1is a single-build . It remains separate from fixed-load A/B regression results and never claims maximum fleet capacity.- A may come from the throughput/latency rule or the reviewed normalized per-worker CPU ceiling.
- Shared PostgreSQL, RabbitMQ, agent, or host pressure invalidates a capacity point instead of being reported as Vigilo's worker knee.
- Publication requires matching build digests and immutable canonical evidence. It binds a reviewed into a versioned policy and .
Deployment Projection
cargo perf project creates a
from
one completed bounded-capacity campaign and a versioned input under
performance/deployments. The input names peak and average traffic, evaluation
run and evaluator mix, payloads, retries, topology, utilization policy, and
usable dependency limits. Each limit retains its measurement, provider, or
operator provenance rather than being inferred from a missing value.
The projection jointly resamples all one/two-worker points and resource-demand coefficients. Its report shows worker, CPU, memory, PostgreSQL, RabbitMQ, HTTP, Wasm, storage, retry, and coordinator demand; formulas and 95% intervals remain visible beside the limiting resource. Missing limits produce unknown overall capacity. Material nonlinearity, shared dependency saturation, operationally unbounded paths, and staging error beyond the declared tolerance prevent a supported label.
Worker counts above the measured two-process range are directional scenarios, not fleet measurements. A small staging observation records model error without requiring maximum scale. The current coordinator broker publish/confirm path is declared unbounded, so the checked-in planning example remains invalid until a separately tested production deadline and settlement contract exists.
Stability And Recovery
Each long-running is separate from fixed-load comparisons. The bounded soak holds one production coordinator and worker process open for at least 30 minutes while small runs make useful progress at fixed observation intervals. It gates exact terminal work, queue drain, process liveness, absolute RSS, file-descriptor growth, retry/delivery amplification, and end-window throughput retention. These safety limits do not become reviewed performance budgets without repeated canonical evidence.
Controlled recovery uses the same isolated topology. It completes work before a real run-owned RabbitMQ application restart, retains the resident processes, then requires new useful work and a drained backlog within the registered deadline. The machine artifact records the fault, recovery time, interval observations, exact durable totals, and every failed condition. A deterministic pure verdict fixture proves that leaks, stalls, lost work, early exits, and missed recovery are not accepted when services are unavailable to unit tests.
Verdict Flow
A comparison using a loads one exact budget for each workload tuple. Results from a different environment ID, too few blocks, excessive orientation bias, or a failed correctness oracle are non-green. A confidence interval within budget passes. The first interval wholly beyond budget is inconclusive and requires one fixed independent confirmation; only the second matching result is a regression.
First Use
Run cargo perf check, create one
,
and begin with the service-free developer-v1 startup workload. When that
passes, run pr-v1 against the same snapshot with Docker available. This
exercises the small real PostgreSQL, RabbitMQ, HTTP-agent, Wasm, and lifecycle
paths before committing to a reference, capacity, recovery, or soak campaign.
Exact-oracle failures take priority over timing; pr-v1 produces
rather
than a regression verdict.
MadSim Boundary
A future MadSim suite should run first with named scenarios, fixed seeds, and exact logical outcomes. It can prove deterministic scheduling, retry, timeout, and fault behavior, but its virtual durations must never enter performance samples, calibration, or budgets. Each simulated scenario should identify the narrow real workload contracts it affects. Those workloads are then compared as immutable baseline and candidate release binaries on the exclusive canonical host using a published gating profile.
An initial over-budget comparison remains inconclusive until the fixed
independent confirmation also exceeds the budget. Fault and sustained-operation
changes additionally run the real recovery or soak profile. A valid conclusion
is that no regression was detected at the published budget and confidence, not
that simulation proved performance unchanged. The complete command sequence is
documented in the repository's performance/README.md.
Shared hosted CI runs service-free contracts and exact correctness checks, not
wall-clock gates. Canonical, nightly, and soak jobs are defined for an exclusive
self-hosted Linux runner but remain inert until repository variables enable the
runner and then its schedules. Manual runs are enabled separately so the host
can be validated before recurring work starts. Reports are appended to the job
summary and retained as bounded artifacts; no performance badge is claimed
before canonical repeatability exists. Activation requires an Actions runner
version 2.327.1 or newer for the workflow's Node 24 actions, the documented
runner labels, and repository variables; it requires no workflow edit.
The comparison on the
defaults to informative reference-v1; after review,
VIGILO_PERF_REFERENCE_PROFILE selects a committed gating profile without a
workflow change.