Skip to content

Add bf16 gemm pingpong with num_stages=3 - #818

Open
jungpark-mlir wants to merge 2 commits into
shared/triton-gfx950-launchfrom
jungpark/pp-ns3-merge
Open

jungpark-mlir wants to merge 2 commits into
shared/triton-gfx950-launchfrom
jungpark/pp-ns3-merge

Conversation

@jungpark-mlir

Copy link
Copy Markdown

Transforms 'BLOCK_SIZE_M': 256, 'BLOCK_SIZE_N': 256, 'BLOCK_SIZE_K': 32 with num_stages=3.
With tripke buffering, it can deal hazard between waves more easily. e.g., wait/copy/dot can be freely ordered.
This requires carrying wait token over the loop (TRITON_HIP_ENABLE_F16_ASYNC_PINGPONG=1).
Waits are merged in the pingpong pass for now, since pipeliner is not able to merge it correctly.

Transforms 'BLOCK_SIZE_M': 256, 'BLOCK_SIZE_N': 256, 'BLOCK_SIZE_K': 32
with num_stages=3, it can deal hazard between waves more easily.
Requires carrying wait token over the loop, waits are merged in the
pingpong pass.
return;
}


Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: extra newline

return success();
}

LogicalResult Pingponger::transformNS3(OpBuilder &builder, Location loc) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see there's already a comment for documentation; but it would be great to capture what these transform functions do; e.g. what is the assumed state of the IR before tranformation, how/when is conditional barrier inserted, what the execution mode looks like before/after the transformation.

Comment thread third_party/amd/lib/TritonAMDGPUTransforms/StreamPipeline.cpp
amd.passes.ttgpuir.add_reorder_instructions(pm)
use_block_pingpong = is_pingpong_schedule_enabled(options.arch)
if use_block_pingpong and options.num_stages in [2, 4]:
if use_block_pingpong and options.num_stages in [2, 3, 4]:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we only allow pingpong to be enabled for pre-vetted scenarios where we know that pingpong is faster than non-pingpong? Or do we allow the user to enable pingpong even when it's slower (e.g. for testing or analysis...)?
I ask because I'm wondering if we distinguish between "supported" cases being numerically correct or performant also.

Comment thread third_party/amd/lib/TritonAMDGPUTransforms/BlockPingpong.cpp Outdated
Try to reduce salu/valu stalls by
using sched.group.barrier to interleave MFMA and salu/buffer_load
using sched.barrier to split ds_read and interleave address calcs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants