Repository navigation
Conversation
… last cases Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
|
| if (!is_single_tensor) { | ||
| TensorMapStorage *ptr = nullptr; | ||
| NVTE_CHECK_CUDA( | ||
| cudaMallocAsync(reinterpret_cast<void **>(&ptr), sizeof(TensorMapStorage), stream)); | ||
| tensor_maps.reset(ptr); |
There was a problem hiding this comment.
Repeated workspace allocations Every varying-last-dimension grouped MXFP8 call now allocates and frees a workspace of roughly 34 KiB, even when calls are serialized. That adds allocator work to a quantization hot path that previously used static storage. Please measure the latency of repeated small grouped calls and consider a reusable, concurrency-safe workspace if the cost is material.
Knowledge Base Used: Native GEMM and quantization kernels
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Description
TODO: the impact to CPU overhead is huge: quantize + sync goes from ~50 µs to ~300 µs. I need to come up with something else. This naive fix costs too much
Fixes #3630
Type of change
Changes
Checklist: