Skip to content

Mpi - #221

Open
rodrigoaramendes wants to merge 9 commits into
mainfrom
mpi
Open

Mpi#221
rodrigoaramendes wants to merge 9 commits into
mainfrom
mpi

Conversation

@rodrigoaramendes

Copy link
Copy Markdown
Contributor

Summary

Adds vectorized overlap evaluation, chunked and MPI-parallelized one-sided projected energy, cache-preserving parameter assignment, and coupled-cluster derivative-overlap caching to the objective/equation layer. Motivated by profiling a PCCD/STO-6G N2 calculation, where get_energy_one_proj and repeated scalar overlap/Hamiltonian-integration calls dominated wall-clock time (~3.5x10^11 total calls across a Powell-minimized run).

What's implemented

Vectorized overlap dispatch (fanpy/eqn/base.py)

  • New wrapped_get_overlaps(sds, deriv=False) on BaseSchrodinger. Uses a wavefunction's native get_overlaps batch method when available (hasattr check), otherwise falls back to looping the existing scalar get_overlap. Handles both the plain-wavefunction and composite-wavefunction (ProductWavefunction/LinearCombinationWavefunction) cases.
  • Note: as of this PR, no wavefunction class in the tree actually implements get_overlaps yet (need to work on that), so this currently always takes the scalar-fallback path. The dispatcher is correctly wired and tested against the scalar path; the actual speedup via vectorization requires implementing a batched get_overlaps on whichever wavefunction class(es) profiling identifies as hot (e.g. the PCCD/AP1roG overlap evaluation flagged in profiling). Tracked as follow-up, not blocking this PR.

Chunked and MPI-parallelized one-sided projected energy (fanpy/eqn/base.py, fanpy/eqn/energy_oneside.py)

  • EnergyOneSideProjection gains an mpi_comm=None constructor parameter, stored as self.mpi_comm and threaded through objective()/gradient() for root-only printing/logging.
  • assign_refwfn accepts a chunked projection-space object (anything exposing iter_chunks()), in addition to the existing list/tuple/CIWavefunction forms.
  • get_energy_one_proj accumulates the projected-energy numerator/denominator (and derivatives) chunk-by-chunk when given a chunked refwfn, and additionally splits work by MPI rank (comm.Get_rank()/Get_size()) with allgather-based reductions when mpi_comm is supplied and Get_size() > 1. Both mechanisms compose (chunking + MPI together).
  • Verified: MPI-reduced results match the serial computation bit-for-bit-equivalent (within np.allclose) for both the flat SD-list and chunked-refwfn cases, tested via a real-multithreaded fake communicator (exercises the actual allgather collective without requiring mpi4py/mpiexec in CI).

Cache-preserving parameter assignment (fanpy/eqn/base.py)

  • assign_params now compares each component's proposed parameter array against its current one before calling component.assign_params, skipping the call (and whatever cache-clearing it triggers) when unchanged. Verified via a spy test confirming components are only reassigned when their slice of the parameter vector actually changes.

Coupled-cluster derivative-overlap caching (fanpy/wfn/cc/base.py)

  • BaseCC gains its own load_cache(), setting up functools.lru_cache-backed _olp and _olp_deriv, plus an uncached (maxsize=0) _olp_double_deriv. clear_cache() (inherited from BaseWavefunction) works against these unchanged, since it operates generically over self._cache_fns.

Optional basin-hopping solver wrapper (fanpy/solver/equation.py)

Explicitly out of scope for this PR

  • HDF5-backed projection spaces (fanpy/tools/hdf5_space.py, HDF5SlaterSpace): still under active development per the accompanying writeup; no tests added here, should land in a separate PR once stable.
  • Two-sided and variational energy: do not have MPI implementations; unaffected by this PR.

Bugs found during review/tests and fixed here

All three share the same root cause: np.array([... for sd in sds]) collapses to shape (0,) instead of the expected 2-D (0, k) when sds is empty, which happens whenever a chunk's size doesn't divide evenly across MPI ranks, leaving one or more ranks with an empty local determinant slice.

  • wrapped_get_overlaps, non-composite deriv branch: fixed with .reshape(sds.size, inds_component.size).
  • wrapped_get_overlaps, composite-wavefunction deriv branch: same fix, same reasoning.
  • get_energy_one_proj, chunked+MPI branch, d_integrals construction: fixed with .reshape(local_ref_sds.size, self.active_nparams). (The flat SD-list MPI branch already guarded this case correctly with an explicit if local_ref_sds.size: ... else: np.zeros(...); the chunked+MPI branch didn't, until this fix.)

Known follow-up (not fixed in this PR, filed separately)

  • BaseCC.load_cache's memory-budget formula for the lru_cache maxsize is missing a factor of 8 (byte-to-element-count conversion), so the CC overlap-derivative cache can grow to ~8x the configured memory= budget. Filed as issue #.

Tests added

  • test_objective_schrodinger_base_new_features.py: vectorized-overlap/scalar consistency (with and without a native batch method), cache-preserving assign_params, MPI-gated save_params, chunked-vs-unchunked energy agreement, and MPI-reduction correctness (flat list and chunked, using a real-threaded fake communicator).
  • test_objective_schrodinger_onesided_energy_new_features.py: chunked refwfn assignment, MPI-gated root/non-root printing in objective()/gradient(), and the gradient() default-argument changes (normalize now defaults to False, save now respects step_save, matching objective()).

For real multi-rank MPI (requires mpi4py):

mpiexec -n 4 python your_script.py

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant