Skip to content

Add benchmarking for vihaco repository - #103

Closed
robpatterson13 wants to merge 1 commit into
mainfrom
benchmark-core
Closed

robpatterson13 wants to merge 1 commit into
mainfrom
benchmark-core

Conversation

@robpatterson13

Copy link
Copy Markdown
Collaborator

Summary

This PR adds a repository-local benchmark suite for vihaco. It measures VM
execution and composite routing against native Rust and Python implementations,
and supports local comparisons between library revisions.

What changed

  • Added a separate Cargo workspace under benchmarks/.
  • Added stable benchmark traits that keep the harness independent of vihaco
    types.
  • Added a checkout-specific machine adapter with CPU and composite execution
    routes, SST loading, resolution, and isolated instruction benchmarks.
  • Added seven workloads covering arithmetic, bitwise operations, branching,
    dispatch, large stack frames, recursive Fibonacci, and exponentiation by
    squaring.
  • Added workload contracts with bounded inputs, expected results, and
    independent Rust and Python implementations.
  • Added validation for native Rust, Python, CPU SST, composite SST, and isolated
    instruction routes.
  • Added Criterion measurements for native Rust, SST, CPU, composite, and
    instruction routes.
  • Added smoke and full comparison profiles.
  • Added a Python runner that builds and measures candidate and baseline
    revisions, records provenance, checks source fingerprints, and writes
    Markdown and JSON results.
  • Added support for baseline machine overrides with --base-machine and
    --base-machine-path.
  • Added a native scaling and disassembly audit tool.
  • Added local quality checks and detailed local failure logs.
  • Added extensive Python test coverage for contracts, validation, reporting,
    privacy, subprocess handling, and result handling.
  • Added CLI documentation for the runner, audit tool, validation commands, and
    baseline machine options.
  • Added contributor guidance requiring the benchmark machine to compile and
    pass validation when vihaco APIs change.
  • Added ignores for the benchmark Python environment and Ruff cache.

Usage

Run a local smoke comparison:

uv sync --directory benchmarks --locked

uv run --directory benchmarks python -m runner \
  --profile smoke \
  --base main \
  --base-machine-path machine \
  --output ../target/benchmark-runs/example

Use --base-machine-path machine while the benchmark machine exists only in
the working tree. After the benchmark machine is committed, a revision-based
override such as --base-machine HEAD can be used.

Run full sampling:

uv run --directory benchmarks python -m runner \
  --profile full \
  --base main \
  --output ../target/benchmark-runs/full

Audit native scaling:

uv run --directory benchmarks python -m runner.audit \
  ../target/benchmark-runs/example

Validation

The following checks pass:

cargo test --locked \
  --manifest-path benchmarks/Cargo.toml \
  --workspace --all-targets

uv run --directory benchmarks python -m runner.checks

cargo run --locked --manifest-path benchmarks/Cargo.toml

An end-to-end smoke comparison completed successfully with:

  • Candidate and baseline builds
  • Candidate and baseline validation
  • 55 SST measurements per revision
  • 26 native Rust reference measurements
  • 26 Python reference measurements
  • Report and publishable artifact generation

Performance results are advisory. Correctness and measurement failures should
be investigated before relying on a comparison.

@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-09-21 18:40 UTC

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant