Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
8e7310b
bench: carry the harbor-buzz-orchestra harness onto latest main
atishpatel Aug 1, 2026
8e07dab
bench: log what a tool call touched, and pin the effort-wave endpoints
atishpatel Aug 1, 2026
d03fcff
bench: pin the OpenRouter upstream so a cell is one condition
atishpatel Aug 1, 2026
a82d691
bench: add the OpenRouter wave -- kimi-k3 and deepseek-v4-flash cells
atishpatel Aug 1, 2026
affbe81
fix: settle usage on the timeout path, not only after DONE
atishpatel Aug 3, 2026
a362418
bench: add the Meli solo cells -- gmicloud pin (run 1) and baseten/fp…
Aug 7, 2026
8b22413
bench: add the Meli solo cloudflare/fp8 cell (take 3)
Aug 7, 2026
b9fedcc
bench: add Mia/Mie twin-team condition (tb-twins-mia-mie)
Aug 8, 2026
fc8a2d2
bench: twins-2 mia persona — wake-Mie triggers, reread rule, delivera…
Aug 9, 2026
97b85f0
bench: add Luna xhigh persona optimization baselines
Aug 14, 2026
a74ceaf
bench: add persona v0 max baselines
Aug 14, 2026
bb8c50a
moar benchmarks
atishpatel Aug 15, 2026
023ea2d
bench: support prompt transport ablations
atishpatel Aug 18, 2026
ebb0730
bench: add leaderboard-safe DeepSeek max manifest
atishpatel Aug 18, 2026
a075619
fix: preserve DeepSeek max effort on Databricks
atishpatel Aug 18, 2026
1e83943
test: lock DeepSeek max effort passthrough
atishpatel Aug 18, 2026
789900c
test: keep benchmark clock fixture current
atishpatel Aug 20, 2026
a0f6211
bench: emit submittable trajectories and parse current usage
atishpatel Aug 20, 2026
af1a30b
Merge remote-tracking branch 'origin/main' into benchmark/harness-acc…
atishpatel Aug 20, 2026
9b7d06a
test: reconcile merged model corpus count
atishpatel Aug 20, 2026
09bde0b
bench: refresh merged testbed lockfile
atishpatel Aug 20, 2026
f8e1551
fix(bench): preserve task evidence across runtime merge
atishpatel Aug 20, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion Justfile
Original file line number Diff line number Diff line change
Expand Up @@ -335,7 +335,7 @@ test-unit:
# buzz-agent model-capabilities corpus: the Rust half of the
# cross-language drift guard. `model_capabilities.rs` embeds
# scripts/model-capabilities.json + scripts/normative-corpus.json via
# include_str! and replays the full locked corpus as pure in-process tests (no
# include_str! and replays all 104 vectors as pure in-process tests (no
# infra). Enumerated explicitly because nothing in CI runs
# `cargo test --workspace`; without this step a manifest edit that
# diverges Rust from the corpus ships green.
Expand Down
1 change: 1 addition & 0 deletions benchmarks/harbor-buzz-orchestra/.gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -4,3 +4,4 @@ __pycache__/
.pytest_cache/
.benchmark/
jobs/
jobs-archive/
290 changes: 242 additions & 48 deletions benchmarks/harbor-buzz-orchestra/README.md

Large diffs are not rendered by default.

453 changes: 453 additions & 0 deletions benchmarks/harbor-buzz-orchestra/docs/PERSONAS.md

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
# OR3h -- LHTB team on OpenRouter: kimi-k3 lead, deepseek-v4-flash scout,
# deepseek-v4-flash worker. All three seats at reasoning effort `high`.
#
# Structurally this is tb-gt-sol-luna-terra-high with open-weight models
# substituted seat for seat: an expensive lead directing two cheap seats, one
# read-only scout and one that edits and builds. That cell is the reason to run
# this one -- it scored 0.695 against solo Sol's 0.591 on LHTB-46, and the
# open question is whether the long-horizon team win survives a 22x cheaper
# roster or was really about the specific models.
#
# NOTE THE ASYMMETRY WITH ITS OPENAI ANALOGUE. There, scout and worker were
# DIFFERENT models (luna scout, terra worker), which is what let that cell
# separate "cheap in the read-only seat" from "cheap in the editing seat".
# Here both cheap seats are the same model, so this cell cannot make that
# distinction -- it is the two-model analogue of tb-gt-sol-2terra-high, not of
# tb-gt-sol-luna-terra-high. Reading a seat-placement result out of it would be
# reading something it does not contain.
#
# ROSTER IDS ARE PROTOCOL, NOT LABELS. `lead`, `scout` and `worker` are read
# verbatim out of the "Your team" table by the gt personas. Renaming any of them
# breaks addressing silently: the @mention resolves to nobody, the send still
# reports success, and the trial stalls to its timeout.
#
# BOTH UPSTREAMS ARE PINNED (openrouter-live.json): moonshotai/mxfp4 for the
# lead, gmicloud/fp8 for both cheap seats. On a team cell the pin matters more
# than on a solo one -- unpinned, each seat independently lands on a different
# upstream per request, so "the composition" would not even be stable within a
# single trial. It is also what makes prompt caching work: measured 2026-08-01,
# pinned routes served a repeated prefix from cache on every call while the
# unpinned route managed one in three. A team re-sends more history than a solo
# agent does, so this cell is the one that would have been hurt worst.
#
# THIS RUNS UNDER THE PATCHED HARBOR. LHTB needs continue_until_timeout honored
# (patches/apply_continue_until_timeout.py); stock harbor ignores the flag and
# agents self-report DONE minutes into hour-long budgets. Gate the run on
# patches/gatecheck.py -- an unpatched harbor scores this cell far below the
# leaderboard and the run looks merely bad rather than misconfigured.
schema_version: "1"
condition: lhtb-or-kimi-lead-2deepseek-high
roster:
- id: lead
kind: orchestrator
role: lead
count: 1
# The EXPENSIVE model in the directing seat -- 22x the input rate of the
# seats it directs. Whether that is worth it is the cell's question.
endpoint: moonshotai/kimi-k3
model_revision: moonshotai/kimi-k3-20260715
prompt:
# Byte-identical to the file every other gt cell pins.
path: personas/bench/gt/gt-lead.md
sha256: 8a27e833dec0a6d1700c5ea3100f8a81c09e5090ef0a589b9583fac40b24c05c
generation:
thinking_effort: high

- id: scout
kind: worker
role: scout
count: 1
endpoint: deepseek/deepseek-v4-flash-0731
model_revision: deepseek/deepseek-v4-flash-20260731
prompt:
path: personas/bench/gt/gt-scout.md
sha256: 0359a957428dd09e56e57a7fd3fe455d860264d910705f0e4992337dde25e5a5
generation:
thinking_effort: high

- id: worker
kind: worker
role: worker
count: 1
endpoint: deepseek/deepseek-v4-flash-0731
model_revision: deepseek/deepseek-v4-flash-20260731
prompt:
path: personas/bench/gt/gt-worker.md
sha256: 7a529ae58f2a2ec635fdf65fb43284b30c09d2c9dbf31e70951fe021c3b2b766
generation:
thinking_effort: high

prices:
# The rates of the PINNED endpoints, not the model ids' headline rates. Both
# rows are load-bearing: the whole point of the cell is the split between an
# expensive lead and cheap seats, so a wrong row moves the conclusion and not
# just the total.
moonshotai/kimi-k3:
input_per_million_usd: 3.00
cached_input_per_million_usd: 0.30
output_per_million_usd: 15.00
cache_read_rate: 0.0
deepseek/deepseek-v4-flash-0731:
input_per_million_usd: 0.133
cached_input_per_million_usd: 0.0266
output_per_million_usd: 0.266
cache_read_rate: 0.0
trial_budget:
# Identical to the LHTB cells it is meant to be read against. Does not bind
# on LHTB anyway -- Harbor enforces each task's own `[agent] timeout_sec`
# scaled by the run's --timeout-multiplier (3.0).
timeout_seconds: 36000

environment:
# Identical in every condition. override_cpus 4 with -n 8 is exactly 1:1 on a
# 32-vCPU m7a.8xlarge.
override_cpus: 4
override_memory_mb: 8192
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,7 @@ roster:
prompt:
path: personas/orchestrator-tb.md
sha256: 206829331cdd85fb277266901b9489e190ba23097600a8bb69142d044b9a42e6
generation:
max_output_tokens: 4096
context_window_tokens: 200000

- id: worker
kind: worker
role: implementer
Expand All @@ -26,17 +24,30 @@ roster:
prompt:
path: personas/worker-tb.md
sha256: 664a994ae3ba93f02ebbb8f61f781f6a59ea3327b21e26e7ffc2d06827ed3e45
generation:
max_output_tokens: 4096
context_window_tokens: 200000

# Cache-read rates are 0.0 on every Anthropic-route endpoint on purpose:
# buzz-agent never sends `cache_control`, and Anthropic prompt caching is
# opt-in, so no cache reads occur and no discount should be modelled. Revisit
# this the moment buzz-agent starts requesting caching, and set a measured rate
# for OpenAI-route endpoints (which cache automatically) rather than a guess.
prices:
claude-sonnet-4-6:
input_per_million_usd: 3
cached_input_per_million_usd: 0.3
output_per_million_usd: 15
cache_read_rate: 0.0
claude-haiku-4-5:
input_per_million_usd: 1
cached_input_per_million_usd: 0.1
output_per_million_usd: 5
cache_read_rate: 0.0
trial_budget:
timeout_seconds: 900
# 36000 = the dataset's longest task (12000s) at the 3x timeout multiplier the
# study runs with (DEFAULT_TIMEOUT_MULTIPLIER in scripts/benchmark.py). The two
# are one setting in two files: benchmark.py refuses to start if this is the
# smaller of the pair, because the harness would then cut a task short of the
# deadline Harbor granted it and record a timeout the agent never hit. A flat
# 900s did exactly that before the multiplier existed — 39 of the 89 tasks
# allow more than 900s, and one live trial died at 905s against a task budget
# of 1800s.
timeout_seconds: 36000
109 changes: 109 additions & 0 deletions benchmarks/harbor-buzz-orchestra/manifests/tb-gt-3luna.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
# G0 -- the goosetown persona set at the cheap floor: a luna lead coordinating
# one read-only luna scout and one luna worker, all three direct from OpenAI.
#
# This is the CONTROL for the goosetown wave, and without it a G1s win is
# uninterpretable. G1s/G2s put an expensive model in the lead seat at the same
# time as they change the persona bytes, so a score move there could be either
# thing. G0 holds the personas fixed and puts the cheap model in the lead seat,
# so:
#
# G0 -> G1s varies the lead's model only (luna -> sol)
# G0 vs C1 varies the persona scheme (and the route -- see below)
#
# It is also the only cell in this wave that serves the study's actual headline.
# G1s/G2s/G1/G2 are all hybrid cells: they ask whether a better-orchestrated
# expensive lead earns its keep, which is doc 02 §1's *rival*, not its thesis.
# "Three cheap models talking to each other match one expensive model" needs a
# cheap-only team, and after B1 (0.536) and C1 (0.494) both came in *below* solo
# luna (0.545) this is the study's remaining shot at it.
#
# What is being tested is not prose. The old lead persona
# (personas/bench/lead-delegate.md:40) told the lead "do the work directly when
# that is the shorter path", which describes a solo agent that occasionally pays
# a round trip -- and that is exactly what C1/C3 measured (score at or below the
# matching solo, input tokens ~1.9x). personas/bench/gt/gt-lead.md inverts it
# into a boundary: reads are the lead's, every byte written is a teammate's, and
# there is no trivial-write exception. The scout seat is new to the study
# entirely -- a read-only role cannot collide with anyone, which is what makes
# parallel recon safe and is why gt-lead is allowed to send several assignments
# in one turn.
#
# ROUTE CAVEAT. This runs direct against api.openai.com, while A1/B1/C1 ran
# through the Databricks AI Gateway. A G0-vs-C1 comparison therefore varies the
# persona scheme *and* the route. Within the goosetown wave, G0/G1s/G2s share
# the route exactly, so those three comparisons are clean. Report the route with
# the number.
#
# Endpoint names resolve via testbed/endpoints/openai-live.json.
schema_version: "1"
condition: tb-gt-3luna
roster:
- id: lead
kind: orchestrator
# Read verbatim out of the "Your team" table by both delegate personas
# ("Your lead is the teammate whose Role column reads `lead`"). Renaming it
# means every report @mentions nobody, the sends all report success, and the
# trial stalls to its timeout with no error anywhere.
role: lead
count: 1
endpoint: gpt-5.6-luna
model_revision: gpt-5.6-luna
prompt:
path: personas/bench/gt/gt-lead.md
sha256: 8a27e833dec0a6d1700c5ea3100f8a81c09e5090ef0a589b9583fac40b24c05c

- id: scout
kind: worker
# `scout` is read by gt-lead.md, which routes recon and verification here
# and edits away from here. A seat labelled anything else gets sent edits it
# is forbidden to make, and answers "this needs a worker" every time.
role: scout
count: 1
endpoint: gpt-5.6-luna
model_revision: gpt-5.6-luna
prompt:
path: personas/bench/gt/gt-scout.md
sha256: 0359a957428dd09e56e57a7fd3fe455d860264d910705f0e4992337dde25e5a5

- id: worker
kind: worker
role: worker
count: 1
endpoint: gpt-5.6-luna
model_revision: gpt-5.6-luna
prompt:
path: personas/bench/gt/gt-worker.md
sha256: 7a529ae58f2a2ec635fdf65fb43284b30c09d2c9dbf31e70951fe021c3b2b766

prices:
# Repriced 2026-07-30: luna 1.0/0.1/6.0 -> 0.20/0.02/1.20 (-80%).
# Sol, Opus 5, Sonnet and Haiku did not move. This changes the
# condition hash, so a cell that ran before this date carries the old
# hash and the old rates in its receipts -- that mismatch is expected,
# not corruption. Restate a completed cell's cost with
# benchmark-runs/tools/reprice.py, which re-prices the measured tokens
# at a new sheet. Never edit a receipt: the tokens are the
# measurement, the price sheet is only an overlay on them.
# Identical to every other luna cell, deliberately: holding the rates equal is
# what makes the cost *ratio* against A1 and C1 sound.
gpt-5.6-luna:
input_per_million_usd: 0.2
cached_input_per_million_usd: 0.02
output_per_million_usd: 1.2
# Fallback only; superseded by the provider's measured split (doc 02 §2.1).
# Caching matters more for a team than for a solo agent, because a team's
# bill is driven by re-sent context and that is exactly what a prefix cache
# serves.
cache_read_rate: 0.0
trial_budget:
# 12000s longest task x the study's 3x multiplier, identical to every cell in
# the comparison set. benchmark.py's check_budget_clears_clock refuses to
# start below that product.
timeout_seconds: 36000

environment:
# Identical in every condition, solos included -- see tb-solo-luna.yaml for
# the full reasoning. Note this roster is 3 agents in one container, the same
# headcount C1 ran at.
override_cpus: 4
override_memory_mb: 8192
111 changes: 111 additions & 0 deletions benchmarks/harbor-buzz-orchestra/manifests/tb-gt-opus-2luna.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
# G1 -- goosetown personas with a claude-opus-5 lead over one read-only luna
# scout and one luna worker, through the Databricks AI Gateway.
#
# THIS IS THE CLEAN PERSONA A/B, and it is the only cell in the wave that is
# one. Against tb-team-opus-luna (C3, condition tb-team-opus-2luna) it holds
# constant every single thing the study can hold:
#
# same lead model databricks-claude-opus-5
# same delegate model databricks-gpt-5-6-luna, 2 seats
# same headcount 3 agents in one container
# same route opus -> anthropic-messages, luna -> /responses
# same price sheet, same clock, same container, same effort (medium, unpinned)
#
# The only difference is the persona bytes. So the G1-vs-C3 delta is
# attributable to the prompt scheme and to nothing else -- which is what makes
# this cell worth running on the two-route gateway rather than moving it to the
# clean OpenAI path like G1s. Moving it would have bought route cleanliness and
# lost the controlled comparison, and G1s already provides the clean route.
#
# What C3 measured, and what this is trying to fix:
#
# C3 score 0.831 cost $68.00 input 51.6M $/solved 0.92
# A3 score 0.795 cost $48.17 input 28.5M $/solved 0.69 (solo opus)
# A2 score 0.843 cost $41.52 input 25.5M $/solved 0.55 (solo sol)
#
# C3 bought +0.036 over solo opus for 1.81x its input tokens, and still lost to
# solo sol on both axes. The diagnosis is in the old persona's own words:
# personas/bench/lead-delegate.md:40 told the lead "do the work directly when
# that is the shorter path" and framed delegation as something that has to earn
# its round trip. That describes a solo agent with an expensive habit, and C3's
# numbers are what that produces -- score at solo-opus level, tokens at team
# level.
#
# personas/bench/gt/gt-lead.md replaces the trade-off with a boundary: reads are
# the lead's, every byte written is a teammate's, and there is deliberately no
# trivial-write exception, because the exception is the loophole the old persona
# fell through. Whether a lead that cannot type is better or worse than one that
# can is the question; a null here with C3's cost profile intact would be a
# genuine finding about hierarchy, not a failed rewrite.
#
# Endpoint names are exact Databricks serving-endpoint names and resolve via
# testbed/endpoints/databricks-example.json. buzz-agent's `databricks_v2` provider
# picks the route per model from the endpoint name; see tb-solo-opus.yaml for the
# detail.
schema_version: "1"
condition: tb-gt-opus-2luna
roster:
- id: lead
kind: orchestrator
role: lead
count: 1
endpoint: databricks-claude-opus-5
model_revision: claude-opus-5
prompt:
path: personas/bench/gt/gt-lead.md
sha256: 8a27e833dec0a6d1700c5ea3100f8a81c09e5090ef0a589b9583fac40b24c05c

- id: scout
kind: worker
role: scout
count: 1
endpoint: databricks-gpt-5-6-luna
model_revision: gpt-5.6-luna
prompt:
path: personas/bench/gt/gt-scout.md
sha256: 0359a957428dd09e56e57a7fd3fe455d860264d910705f0e4992337dde25e5a5

- id: worker
kind: worker
role: worker
count: 1
endpoint: databricks-gpt-5-6-luna
model_revision: gpt-5.6-luna
prompt:
path: personas/bench/gt/gt-worker.md
sha256: 7a529ae58f2a2ec635fdf65fb43284b30c09d2c9dbf31e70951fe021c3b2b766

prices:
# Repriced 2026-07-30: luna 1.0/0.1/6.0 -> 0.20/0.02/1.20 (-80%).
# Sol, Opus 5, Sonnet and Haiku did not move. This changes the
# condition hash, so a cell that ran before this date carries the old
# hash and the old rates in its receipts -- that mismatch is expected,
# not corruption. Restate a completed cell's cost with
# benchmark-runs/tools/reprice.py, which re-prices the measured tokens
# at a new sheet. Never edit a receipt: the tokens are the
# measurement, the price sheet is only an overlay on them.
# Byte-identical to tb-team-opus-luna.yaml. The G1-vs-C3 comparison is a cost
# comparison as much as a score one, so the rate sheet has to be the same
# sheet or the delta is partly a pricing edit.
#
# Opus is 5x luna on input and cache-read but ~4.17x on output; see
# tb-solo-opus.yaml for the cache-write (cache_creation) surcharge caveat,
# which lands hardest on the lead's high-context seat. See tb-solo-luna.yaml
# for why cache_read_rate is 0.0 (DB10).
databricks-claude-opus-5:
input_per_million_usd: 5.0
cached_input_per_million_usd: 0.5
output_per_million_usd: 25.0
cache_read_rate: 0.0
databricks-gpt-5-6-luna:
input_per_million_usd: 0.2
cached_input_per_million_usd: 0.02
output_per_million_usd: 1.2
cache_read_rate: 0.0
trial_budget:
timeout_seconds: 36000

environment:
# Identical in every condition -- see tb-solo-luna.yaml for the full reasoning.
override_cpus: 4
override_memory_mb: 8192
Loading
Loading