Skip to content

ENH: restrict CPU-specific code paths via REPROSEED_CPU - #9

Open
yarikoptic-gitmate wants to merge 5 commits into
masterfrom
claude/ecstatic-brahmagupta-jsqg3p
Open

yarikoptic-gitmate wants to merge 5 commits into
masterfrom
claude/ecstatic-brahmagupta-jsqg3p

Conversation

@yarikoptic-gitmate

@yarikoptic-gitmate yarikoptic-gitmate commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Numerical libraries choose SIMD code paths at run time based on the host CPU, so the same computation with the same seed can give results that differ in the last bits from one machine to another. This PR adds REPROSEED_CPU, which defaults to x86-64-v3 (AVX2+FMA) on x86_64. Elsewhere it defaults to native, which leaves everything alone. On x86_64 it exports the following variables, unless the caller has already set them (even to an empty value):

  • NPY_DISABLE_CPU_FEATURES
  • OPENBLAS_CORETYPE=Haswell
  • MKL_CBWR=AVX2
  • ATEN_CPU_CAPABILITY=avx2
  • ONEDNN_MAX_CPU_ISA/DNNL_MAX_CPU_ISA=AVX2

README.md gets a short section listing these variables; the details and caveats are in the new docs/cpu-features.md.

This PR also moves reproseed's own messages to stderr, so stdout carries only the wrapped command's output.

Closes #7

Rationale, verification

Each commit message carries the details. In brief:

  • What changes without the restriction (outputs compared by md5 on an AVX512 host):
    • NumPy: AVX512 exp/log results and argsort tie order
    • OpenBLAS: SkylakeX vs Haswell kernels change A@A/inv
    • MKL: torch exp/mm
    • PyTorch: softmax
    • oneDNN: conv2d
  • NumPy list: AVX512F AVX512_SKX AVX512_ICL AVX512_SPR X86_V4.
    • These are the AVX512 targets that NumPy 1.20–2.3 actually dispatch, plus the 2.4+ names (2.5 uses the same ones). Disabling one target does not disable the others.
    • Elementwise and BLAS results are identical on NumPy 1.24, 1.26 and 2.4.
    • Not set if /proc/cpuinfo shows no AVX512. Otherwise NumPy 1.20–1.24 would print a visible RuntimeWarning on every import.
    • Not set if NPY_ENABLE_CPU_FEATURES is set, because NumPy refuses to import with both.
  • Variables that force kernels:
    • OPENBLAS_CORETYPE and ATEN_CPU_CAPABILITY select kernels without checking the CPU, so they are set only when /proc/cpuinfo confirms AVX2 and FMA.
    • The other variables only cap what the library may use.
  • MKL across vendors: oneMKL branches other than AUTO and COMPATIBLE are Intel-only and silently fall back to AUTO elsewhere.
    • With the default MKL_CBWR=AVX2, MKL results can still differ between Intel and AMD, as @leej3 measured.
    • The docs point to MKL_CBWR=COMPATIBLE (slower) for mixed-vendor use.
    • Whether COMPATIBLE should be the default is open.
  • Caller's values: values already set by the caller are kept, which covers an explicit MKL_CBWR=COMPATIBLE or AVX2,STRICT and NumPy exclusions such as X86_V3 (Preserve explicit MKL reproducibility settings and NumPy exclusions #11).
  • Not set:
    • GLIBC_TUNABLES: libm ifuncs depend on FMA/AVX2 but not AVX512.
    • Thread counts: they also affect BLAS results; this is documented as a caveat in docs/cpu-features.md.
  • Tests: shellcheck and the repo's pre-commit lint are clean, and 18 bats tests pass.
    • Fake uname, and grep reading a fake cpuinfo file, cover every branch on any host.
    • Mutation-checked: removing any one guard, or any >&2, fails a test.
  • History: master (CI from ENH: tox, GitHub Actions CI, pre-commit with shellcheck and codespell #10) is merged in, not rebased, since Preserve explicit MKL reproducibility settings and NumPy exclusions #11 is based on this branch.
  • Pre-existing issues, not addressed here:
    • Sourcing runs the caller's "$@".
    • set -u fails when REPROSEED is unset.
    • bash is required for $RANDOM.

🤖 Generated with Claude Code

https://claude.ai/code/session_0142mrYgtL7zA9zCpBtEgcxn

reproseed.sh is used as a wrapper of a command (or sourced into a
script), so its own messages must not get mixed into stdout of the
command.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0142mrYgtL7zA9zCpBtEgcxn
@yarikoptic-gitmate
yarikoptic-gitmate force-pushed the claude/ecstatic-brahmagupta-jsqg3p branch 7 times, most recently from dffac41 to 182383e Compare October 2, 2026 17:32
Numerical libraries pick SIMD code paths at run time based on the
detected CPU, so the same computation (same seed, same software) can
give results differing in the last bits on different machines.
Verified on an AVX512 host: NumPy AVX512 exp/log and argsort tie
order, OpenBLAS SkylakeX vs Haswell GEMM kernels, MKL (also within
PyTorch exp/mm), PyTorch softmax, oneDNN conv2d all change.

New REPROSEED_CPU (default x86-64-v3 on x86_64, native elsewhere;
"native" to not touch anything) exports on x86_64:

- NPY_DISABLE_CPU_FEATURES: AVX512 dispatch targets used by NumPy
  (outside of its _simd test module): AVX512F, AVX512_SKX, AVX512_ICL,
  AVX512_SPR in NumPy < 2.4, and X86_V4, AVX512_ICL, AVX512_SPR in
  >= 2.4, since disabling one target does not disable the others.
  Not set if /proc/cpuinfo shows no AVX512 (NumPy warns about unknown
  names, visibly in 1.20-1.24), or if NPY_ENABLE_CPU_FEATURES is set
  (NumPy refuses both).
- MKL_CBWR=AVX2, ONEDNN_MAX_CPU_ISA=AVX2 and DNNL_MAX_CPU_ISA=AVX2
  (oneDNN < 2.5): only cap code paths.
- OPENBLAS_CORETYPE=Haswell, ATEN_CPU_CAPABILITY=avx2: force code for
  the level without checking CPU support, so set only if /proc/cpuinfo
  confirms AVX2 and FMA.

A different preset value is overridden with a warning.  Only warnings
are printed, to stderr.

README.md gets a brief section, with details and caveats in
docs/cpu-features.md; mentioned projects and variables link to their
documentation (or source where undocumented).

Not set: GLIBC_TUNABLES, since libm ifuncs depend on FMA/AVX2 but not
AVX512, so libm is consistent across x86-64-v3 CPUs; numbers of
threads (also affect BLAS results), documented in docs/cpu-features.md
instead.

Closes #7

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0142mrYgtL7zA9zCpBtEgcxn
@yarikoptic-gitmate
yarikoptic-gitmate force-pushed the claude/ecstatic-brahmagupta-jsqg3p branch from 182383e to 5f82297 Compare October 2, 2026 17:36
@yarikoptic

Copy link
Copy Markdown
Member

@CodyCBakerPhD and @leej3 I think this might be of interest to you too: @asmacdo pointed to his fundings today on what is needed to make "trivial" data interpolation reproducible, and it required some extra env vars setting to disable some CPU features. Which reminded me about this little project from the past. Please have a peak -- it might be of interest to your computes or even may be to mention within STAMPED framework to live to the promise of "reproducible" in

Any scientific workflow that adheres to the following principles is guaranteed to be rigorous and reproducible!

@leej3 leej3 left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI draft - not reviewed by John

PR 11 proposes focused changes to PR 9: preserve explicit MKL reproducibility settings, merge NumPy exclusions, and clarify supported-library scope and backend conditions. PR 9 replaces explicit MKL settings with AVX2; PR 11 preserves an explicit COMPATIBLE choice while retaining AVX2 as the default when MKL_CBWR is unset.

Experiments and DataLad reruns cover Typhon, Smaug, and a Unity AMD node, with code, a pinned environment, and downloadable arrays. The tested candidate shell script matches PR 11 commit 51d4c3578fe3bb43aeae87a29c5768a2cc53c4fd byte for byte. A separate follow-up could consider a mixed-vendor profile and numerical coverage for PyTorch/oneDNN; those are not changes proposed by PR 11.

Comment thread reproseed.sh
done << EOF
limit NPY_DISABLE_CPU_FEATURES AVX512F AVX512_SKX AVX512_ICL AVX512_SPR X86_V4
force OPENBLAS_CORETYPE Haswell
limit MKL_CBWR AVX2

@leej3 leej3 Oct 3, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI draft - not reviewed by John

PR 9's default x86-64-v3 profile sets MKL_CBWR=AVX2, including replacing an explicit COMPATIBLE setting. With MKL 2025.2.0, the pr benchmark using PR 9's default reports AVX2 (10) on Typhon and Smaug (Intel), but AUTO (2) on Unity cpu053 (AMD EPYC 7763). At each fixed thread count (1, 2, or 4), all ten tested arrays agree between the two Intel hosts. Comparing either Intel host with Unity at the same thread count, all four MKL GEMMs differ while the six NumPy/OpenBLAS arrays agree.

PR 11 preserves an explicitly supplied MKL_CBWR=COMPATIBLE. In the compatible_fixed benchmark, that configuration makes all ten tested arrays agree across the three hosts at each fixed thread count. The benchmark candidate shell script matches PR 11 commit 51d4c3578fe3bb43aeae87a29c5768a2cc53c4fd byte for byte. PR 11 still defaults to AVX2 when MKL_CBWR is unset; it does not switch the default to COMPATIBLE.

Recorded metadata and comparisons; AMD diagnostic for PR 9's default; Intel branch semantics.

These comparisons support preserving the explicit COMPATIBLE choice for the tested workload and library version. They do not establish agreement across different thread counts or untested workflows. A default profile for mixed Intel/AMD use is a separate question; no performance tradeoff was measured.

@yarikoptic yarikoptic Oct 5, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI draft - not reviewed by John

please chime in back when it is ready/reviewed by John?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two commits address this:

  • 216b22d: a value the caller has already set, such as MKL_CBWR=COMPATIBLE or AVX2,STRICT, is now kept rather than overridden. This applies to every CPU variable, even when set to an empty value.
  • da85ebd: the docs now say that oneMKL branches other than AUTO/COMPATIBLE are Intel-only, so the default MKL_CBWR=AVX2 does not align MKL results on AMD CPUs, and they point to MKL_CBWR=COMPATIBLE for mixed Intel/AMD use.

Whether the default itself should become COMPATIBLE (consistent across vendors, but slower) is left open for @yarikoptic.


Generated by Claude Code

claude added 3 commits October 5, 2026 19:25
…hmagupta-jsqg3p

Conflict in tests/test_basic.sh "source check output": master switched
to $(...) with quoted "$out", this branch added 2>&1 since messages now
go to stderr; kept both.  Merged rather than rebased, since
#11 is based on this branch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0142mrYgtL7zA9zCpBtEgcxn
Previously a different preset value was overridden with a warning,
which e.g. replaced an explicit MKL_CBWR=COMPATIBLE or AVX2,STRICT, or
NumPy exclusions such as NPY_DISABLE_CPU_FEATURES=X86_V3 (found by
@leej3 in #11).  Now any of these variables already
set, even to an empty value, is left as is, silently, as is
NPY_DISABLE_CPU_FEATURES when NPY_ENABLE_CPU_FEATURES is set.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0142mrYgtL7zA9zCpBtEgcxn
oneMKL documents branches other than AUTO and COMPATIBLE as available
only on Intel processors, and silently uses AUTO otherwise.  So with
the default MKL_CBWR=AVX2, MKL results can still differ between Intel
and AMD CPUs, as measured by @leej3 (#9 review).
Point to MKL_CBWR=COMPATIBLE for that case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0142mrYgtL7zA9zCpBtEgcxn
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add env vars to disable CPU specific features affecting basic algebraic operations

5 participants