Repository navigation
Add local benchmarking suite for vihaco - #100
Closed
robpatterson13 wants to merge 4 commits into
Closed
robpatterson13 wants to merge 4 commits into
robpatterson13 wants to merge 4 commits into
Conversation
…CI to smoke test PRs and allow users with write permission to run full benchmarking suite
…eparate APIs to be compared
…co API should update the benchmark machine
Contributor
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This branch adds a repository-local benchmark suite for vihaco. It measures VM
execution and composite routing against native Rust and Python implementations,
and supports both local comparisons and CI runs.
What changed
benchmarks/.types.
routes, SST loading, resolution, and isolated instruction benchmarks.
dispatch, large stack frames, recursive Fibonacci, and exponentiation by
squaring.
independent Rust and Python implementations.
instruction routes.
instruction routes.
revisions, records provenance, checks source fingerprints, and writes
Markdown and JSON results.
--base-machineand--base-machine-path.reporting, privacy, subprocess handling, CI policy, and publication.
--base,--base-machine, and--base-machine-path.machine adapter compiling and validating.
Usage
Run a local smoke comparison:
Run full sampling:
Audit native scaling:
Request a full benchmark from a same-repository pull request:
Publish a successful run:
Validation
The following checks pass:
cargo test --locked \ --manifest-path benchmarks/Cargo.toml \ --workspace --all-targets uv run --directory benchmarks python -m runner.checks cargo run --locked --manifest-path benchmarks/Cargo.tomlAn end-to-end smoke comparison also completed successfully. It covered candidate
and baseline builds, validation, 55 SST measurements per revision, 26 native
Rust reference measurements, 26 Python reference measurements, and report
generation.
Performance results are advisory. Correctness and pipeline failures are
blocking, while timing changes should be reviewed across repeated runs and,
when needed, with native assembly inspection.