Is your feature request related to a problem? Please describe.
The current Triton AOTI deployment runs a single stateful model instance and disables Triton scheduler-level batching:
The exported forward embeds lookup, allocation, onboarding launch/wait, and HSTU compute in one AOTI call: ExportKVCachedInferenceRankingGR.forward. Offload/reap operators exist in the C++ runtime, but they are not part of this serving forward path. If they are added directly to the same AOTI call, FlexKV RPC, GPU-to-host transfer, and task reaping can extend the request critical path and block the only model instance.
The broad scheduling problem is already tracked by #461. This issue should be a concrete implementation subtask focused on separating the KV-cache lifecycle from the compute graph.
Describe the solution you'd like
A clear and concise description of what you want to happen.
Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.
Additional context
Add any other context, code examples, or references to existing implementations about the feature request here.
By submitting this issue, you agree to follow our code of conduct and our contributing guidelines.
Is your feature request related to a problem? Please describe.
The current Triton AOTI deployment runs a single stateful model instance and disables Triton scheduler-level batching:
max_batch_size: 0instance_group count: 1The exported forward embeds lookup, allocation, onboarding launch/wait, and HSTU compute in one AOTI call:
ExportKVCachedInferenceRankingGR.forward. Offload/reap operators exist in the C++ runtime, but they are not part of this serving forward path. If they are added directly to the same AOTI call, FlexKV RPC, GPU-to-host transfer, and task reaping can extend the request critical path and block the only model instance.The broad scheduling problem is already tracked by #461. This issue should be a concrete implementation subtask focused on separating the KV-cache lifecycle from the compute graph.
Describe the solution you'd like
A clear and concise description of what you want to happen.
Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.
Additional context
Add any other context, code examples, or references to existing implementations about the feature request here.
By submitting this issue, you agree to follow our code of conduct and our contributing guidelines.