Skip to content

Databricks run collection, experiments, and dashboard - #25

Merged
jeffbrennan merged 11 commits into
mainfrom
feat/databricks-runs
Sep 13, 2026
Merged

jeffbrennan merged 11 commits into
mainfrom
feat/databricks-runs

Conversation

@jeffbrennan

Copy link
Copy Markdown
Owner

Summary

Adds offline Databricks run collection and an explicit performance-experiment workflow, then hardens it against the review feedback.

  • Collection (databricks.py): page-bounded Query History and run collection that preserves partial data on budget/transport exhaustion, resolves the latest terminal run across pages, and normalizes run metadata.
  • Report (runreport.py): query aggregates are complete only when every matched query is final and attribution is certain; scoped Git evidence (run vs task); run environments/performance target/job parameters captured for configuration identity with credential exclusion.
  • Experiments (experiments.py): trials keyed by run (relabeling no longer duplicates), config fingerprints include execution parameters and distinguish derived vs asserted, partial observations retained as raw deltas with caveats, task comparisons use medians with spread, and stable trial ordering.
  • Dashboard (experiments_dashboard.py): variant/commit/eligibility filters, task trend with per-variant spread, richer run detail, partial values marked distinctly, live disk-backed refresh, and manifest-pinned baseline.

Testing

  • just ci locally: ruff check, ruff format check, pyrefly (0 errors), pytest tests/ — 561 passed, 2 skipped.
  • New unit coverage for pagination retention, query finality/attribution, config identity, credential redaction, task coverage propagation, and dashboard helpers.
  • Integration tests (test.py, test_capture.py) are excluded from just ci and require a JVM.

Preserve pages already fetched when a later Query History request or run
pagination exceeds the collection budget, and paginate job runs/list so the
latest terminal run is found past the first page.
…g identity

Query aggregates are complete only when every matched query is final and no
query in the window is unattributed. Revision evidence records whether a commit
came from the run or a single task, and per-task commits are retained. Run
environments and the run-level effective performance target are normalized so
configuration identity is not silently dropped.
…edians

Trials are keyed by run so relabeling no longer duplicates an execution, and
recollection preserves stored metadata. Configuration identity is derived from
run facts and a variant mixing configurations is rejected instead of pooled.
Metric samples exclude partial subtotals, and per-task comparisons use medians
with retained sample spread.
The experiment dashboard gains variant/commit/eligibility filters, a task
selector with per-task trend and per-variant spread, and run detail with top
queries, failure excerpts and compute context. Missing context metrics render as
gaps instead of zero, refresh reloads comparisons and rows from disk, and the
default baseline honors the manifest's pinned trial.
Plot partial query aggregates with open markers and exclude them from median
lines and CSV/JSON exports so a subtotal is never read as a complete total.
…deltas

Configuration fingerprints now include run job parameters, task parameters and
cluster Spark configuration, so runs that differ only in settings such as
shuffle partitions are separate configurations. Partially observed metrics stay
in the comparison as raw deltas with a coverage caveat instead of disappearing,
while medians still use complete observations. Recollection no longer rewrites
the first-collected timestamp, so trial ordering is stable.
…ports

Task spread now uses only eligible trials of the same configuration. The context
chart uses all runs for its x-axis so partial observations stay aligned, and CSV
export flattens list-valued fields.
Configuration capture now drops credential-like keys (tokens, secrets, access
keys, passwords) from job parameters, task parameters and Spark configuration
before they are stored or fingerprinted. A page-limited task list marks summed
task time as a partial subtotal with a caveat and stops missing tasks from being
reported as added/removed. Recollection recomputes derived configuration
fingerprints while preserving explicitly asserted ones, so late enrichment
cannot leave a stale identity attached to an updated report.
trial_rows() now lists task_execution_ms in partial_metrics when the collected
task list was page-limited, so dashboard plots and CSV/JSON exports qualify the
subtotal instead of presenting it as complete.
Argument lists are joined with the value following a credential-like flag
replaced by <redacted> (including --flag=value form), while opaque positional
values are left untouched.
@gitguardian

gitguardian Bot commented Sep 13, 2026

Copy link
Copy Markdown

⚠️ GitGuardian has uncovered 1 secret following the scan of your pull request.

Please consider investigating the findings and remediating the incidents. Failure to do so may lead to compromising the associated services or software components.

🔎 Detected hardcoded secret in your pull request
GitGuardian id GitGuardian status Secret Commit Filename
37241604 Triggered Generic High Entropy Secret 760b57c tests/data/databricks/jobs_runs_list.json View secret
🛠 Guidelines to remediate hardcoded secrets
  1. Understand the implications of revoking this secret by investigating where it is used in your code.
  2. Replace and store your secret safely. Learn here the best practices.
  3. Revoke and rotate this secret.
  4. If possible, rewrite git history. Rewriting git history is not a trivial act. You might completely break other contributing developers' workflow and you risk accidentally deleting legitimate data.

To avoid such incidents in the future consider


🦉 GitGuardian detects secrets in your source code to help developers and security teams secure the modern development process. You are seeing this because you or someone else with access to this repository has authorized GitGuardian to scan your pull request.

@jeffbrennan
jeffbrennan merged commit 236d13a into main Sep 13, 2026
4 checks passed
@jeffbrennan
jeffbrennan deleted the feat/databricks-runs branch September 13, 2026 20:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant