Minimal reproducible case is:
import cupy as cp
cderi = cp.zeros((159, 937, 937))
_mo1 = cp.zeros((2, 117, 937, 108))
rhok1 = cp.zeros((2, 116, 937, 108, 159))
i0 = 0
i1 = 116
print(f"cderi.shape = {cderi.shape}, _mo1.shape = {_mo1.shape}, _mo1[:,i0:i1].shape = {_mo1[:,i0:i1].shape}, rhok1.shape = {rhok1.shape}")
print(f"cderi.dtype = {cderi.dtype}, _mo1.dtype = {_mo1.dtype}, _mo1[:,i0:i1].dtype = {_mo1[:,i0:i1].dtype}, rhok1.dtype = {rhok1.dtype}")
from gpu4pyscf.lib.cupy_helper import contract
contract('Lpq,snqi->snpiL', cderi, _mo1[:,i0:i1], out=rhok1)
print("pass")
And when running on one A100 card, with native CUDA 12.9, it will crash with the following output:
Floating point exception (core dumped)
This is encountered at
|
contract('Lpq,snqi->snpiL', cderi, _mo1[:,i0:i1], out=rhok1) |
when running UHF/UKS + density fitting + hessian calculation.
To reproduce, it requires cutensor-cu12==2.7.0, the current tip release of cutensor. 2.6.0 will return CUTENSOR_STATUS_NOT_SUPPORTED. 2.5.0 or earlier versions are fine (at least it doesn't crash).
Minimal reproducible case is:
And when running on one A100 card, with native CUDA 12.9, it will crash with the following output:
This is encountered at
gpu4pyscf/gpu4pyscf/df/hessian/rhf.py
Line 1319 in db6bb7f
when running UHF/UKS + density fitting + hessian calculation.
To reproduce, it requires
cutensor-cu12==2.7.0, the current tip release of cutensor.2.6.0will returnCUTENSOR_STATUS_NOT_SUPPORTED.2.5.0or earlier versions are fine (at least it doesn't crash).