Skip to content

Add local benchmarking suite for vihaco - #100

Closed
robpatterson13 wants to merge 4 commits into
mainfrom
rob/add-benchmarking
Closed

robpatterson13 wants to merge 4 commits into
mainfrom
rob/add-benchmarking

Conversation

@robpatterson13

Copy link
Copy Markdown
Collaborator

Summary

This branch adds a repository-local benchmark suite for vihaco. It measures VM
execution and composite routing against native Rust and Python implementations,
and supports both local comparisons and CI runs.

What changed

  • Added a separate Cargo workspace under benchmarks/.
  • Added stable benchmark traits that keep the harness independent of vihaco
    types.
  • Added a checkout-specific machine adapter with CPU and composite execution
    routes, SST loading, resolution, and isolated instruction benchmarks.
  • Added seven workloads covering arithmetic, bitwise operations, branching,
    dispatch, large stack frames, recursive Fibonacci, and exponentiation by
    squaring.
  • Added workload contracts with bounded inputs, expected results, and
    independent Rust and Python implementations.
  • Added validation for native Rust, Python, CPU SST, composite SST, and isolated
    instruction routes.
  • Added Criterion measurements for native Rust, SST, CPU, composite, and
    instruction routes.
  • Added smoke and full comparison profiles.
  • Added a Python runner that builds and measures candidate and baseline
    revisions, records provenance, checks source fingerprints, and writes
    Markdown and JSON results.
  • Added support for baseline machine overrides with --base-machine and
    --base-machine-path.
  • Added a native scaling and disassembly audit tool.
  • Added local quality checks and detailed local failure logs.
  • Added pull request smoke benchmarks and manually requested full runs.
  • Added guarded result publication with authorization and provenance checks.
  • Added extensive Python and Node test coverage for contracts, validation,
    reporting, privacy, subprocess handling, CI policy, and publication.
  • Added CLI documentation, including guidance for choosing --base,
    --base-machine, and --base-machine-path.
  • Added contribution guidance requiring API changes to keep the benchmark
    machine adapter compiling and validating.
  • Added ignores for the benchmark Python environment and Ruff cache.

Usage

Run a local smoke comparison:

uv sync --directory benchmarks --locked

uv run --directory benchmarks python -m runner \
  --profile smoke \
  --base main \
  --output ../target/benchmark-runs/example

Run full sampling:

uv run --directory benchmarks python -m runner \
  --profile full \
  --base main \
  --output ../target/benchmark-runs/full

Audit native scaling:

uv run --directory benchmarks python -m runner.audit \
  ../target/benchmark-runs/example

Request a full benchmark from a same-repository pull request:

@github-actions run benchmark

Publish a successful run:

@github-actions commit benchmark <run-id>

Validation

The following checks pass:

cargo test --locked \
  --manifest-path benchmarks/Cargo.toml \
  --workspace --all-targets

uv run --directory benchmarks python -m runner.checks

cargo run --locked --manifest-path benchmarks/Cargo.toml

An end-to-end smoke comparison also completed successfully. It covered candidate
and baseline builds, validation, 55 SST measurements per revision, 26 native
Rust reference measurements, 26 Python reference measurements, and report
generation.

Performance results are advisory. Correctness and pipeline failures are
blocking, while timing changes should be reviewed across repeated runs and,
when needed, with native assembly inspection.

@robpatterson13
robpatterson13 marked this pull request as draft September 14, 2026 20:04
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-06 18:30 UTC

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant