Documentation · Installation · Quick Start · Compatibility
A PyTorch device plugin for the FlagOS software stack. torch-fl exposes a single flagos device that routes operators among reusable native kernels, portable compiler kernels, vendor-native implementations, and explicit CPU fallback.
Accelerator vendors provide different runtimes, compiler stacks, and PyTorch integration strategies. torch-fl offers a PyTorch-native device interface that isolates these differences behind a unified runtime and operator-routing layer.
Users program against standard PyTorch APIs and the flagos device. The plugin selects kernel implementations per operator based on platform capabilities and configuration, without exposing vendor differences as distinct device names or requiring workload changes when moving between accelerators.
torch-fl is built on five principles:
-
PyTorch-native interface — Standard PyTorch APIs work unchanged; users target the
flagosdevice rather than vendor-specific extensions. -
One logical device — A single device name (
flagos) abstracts vendor differences. Platform-specific routing happens transparently at the operator level. -
Layered operator backends — Each operation may dispatch to a different implementation. Routing decisions are per-operator, not per-device or per-model.
-
Reuse before reimplementation — Established kernels and compiler stacks are integrated where their dispatch and ABI boundaries permit, rather than rewriting functionality that already exists.
-
Explicit capability boundaries — Unsupported operations and CPU fallback paths are documented rather than presented as complete native coverage. Status levels distinguish validated support from experimental integrations.
torch-fl implements three principal operator execution strategies:
-
Native vendor kernels: Direct calls into vendor runtime and operator libraries (ACLNN, mudnn, topsaten). The plugin generates bindings to each vendor's C/C++ API.
-
Compatibility boxing: Zero-copy metadata conversion into an independent PyTorch dispatch key when the vendor stack exposes one that can coexist with PrivateUse1. CUDA boxing reuses NVIDIA kernels via an external
libtorch_cuda.so. -
Portable compiler kernels: FlagGems kernels generated through Triton or a compatible compiler backend. These kernels target multiple accelerator families without per-platform rewrites.
A platform may combine execution paths. These are implementation strategies, not user-selectable product tiers.
PyTorch API
|
flagos device (PrivateUse1)
|
device runtime + per-operator routing
|
FlagGems/compiler kernels | compatibility boxing | native vendor kernels | CPU fallback
|
accelerator runtime
This diagram is conceptual. Detailed component architecture, dispatch internals, distributed collectives, compilation integration, and profiler design are documented under docs/architecture/.
torch-fl provides project-level support for:
- PyTorch eager tensor operations and device management
- Autograd and model training
torch.compileintegration- Distributed collectives and DDP through
ProcessGroupFlagOS torch.profilerintegration- FlagGems/Triton operator integration
- Explicit CPU fallback for uncovered operations where PyTorch semantics permit
Capability availability and validation status vary by platform. A feature existing in the torch-fl codebase does not imply that every platform implements or validates it. See the Compatibility Matrix for per-platform detail.
| Platform | Execution path | Validated capabilities | Status | Guide |
|---|---|---|---|---|
| NVIDIA CUDA | CUDA boxing over external libtorch_cuda.so |
Eager, autograd, distributed (FlagCX/NCCL), profiler (CUPTI), FlagGems (Python + C++) | Stable | CUDA |
| MetaX | CUDA boxing via cu-bridge, or native MetaX kernels | Eager, autograd | Stable | MetaX |
| Ascend | Native ACLNN backend, optional FlagGems via triton-ascend | Eager, autograd, RNG suite | Beta | Ascend |
| PPU | CUDA boxing against PPU's CUDA-13-compatible SDK | Eager, autograd | Experimental | PPU |
| Hygon DCU | CUDA boxing over hipified DTK torch | Eager, autograd, profiler | Beta | DCU |
| Enflame GCU | Native topsaten backend, CPU fallback for unrouted/int64 ops | Eager | Experimental | GCU |
| Moore Threads MUSA | Native mudnn backend, CPU fallback for unrouted ops | Eager | Experimental | MUSA |
| D-Robotics BPU | No eager kernels; torch.compile graph path via hbdk4 |
Graph compilation only | Runtime only | BPU |
| TsingMicro | Runtime build selector exists | Setup not documented | Runtime only | — |
Status definitions:
- Stable: Critical paths are continuously tested and the supported version combination is documented.
- Beta: The primary path is validated, but coverage, packaging, or release procedures are not yet stable.
- Experimental: Validation exists for a specific setup, model, or hardware environment; interfaces or build procedures may change.
- Runtime only: Device runtime support exists, but the platform is not a general eager operator backend.
Detailed capability breakdowns (eager execution, training, compile, distributed, profiler, FlagGems) and on-hardware validation evidence are in the Compatibility Matrix.
| Component | Supported range | Notes |
|---|---|---|
| Python | 3.8 or later | Platform SDKs and available wheels may impose a narrower range. |
| PyTorch | 2.10.x (>=2.10,<2.11) |
Generated ATen bindings are tied to this minor line. |
| FlagGems | Platform dependent | Installed from PyPI or a vendor-compatible build only where the platform route uses it. |
| Triton/compiler | Platform dependent | Use the compiler distribution required by the selected accelerator. |
torch-fl generates native bindings to PyTorch's internal ATen operator registry. These bindings are sensitive to C++ ABI and operator schema changes, so the project pins to a PyTorch minor line. The current pinned line is 2.10.x.
Using a different PyTorch minor version (e.g., 2.11.x) will result in build or runtime failures. Patch versions within the same minor line (e.g., 2.10.0 to 2.10.1) are compatible.
Vendor SDK versions, CUDA toolkit versions, and other platform-specific requirements are documented in each platform's installation guide.
Choose your platform in the Installation Guide, then try the following:
import torch
import torch_fl
# Create a tensor on the flagos device
x = torch.randn(4, 4, device="flagos:0")
# Operations route to platform-appropriate kernels
y = torch.relu(x @ x)
# Move result back to CPU
print(y.cpu())Operator routing (FlagGems, vendor kernels, compatibility boxing, or CPU fallback) is determined by platform detection and runtime configuration. The code above works unchanged across all supported accelerators.
For device queries, synchronization, and multi-device usage patterns, see the Quick Start Guide.
- Installation — Platform selection and source builds
- Quick Start — Basic usage patterns
- Compatibility Matrix — Per-platform capability validation and status
- Operator Support — Measured per-hardware FlagGems overload coverage
- Environment Variables — Build and runtime configuration
- Distributed Collectives — ProcessGroupFlagOS, FlagCX, and vendor fallbacks
- Profiler Integration — torch.profiler parity and CUPTI integration
- torch.compile Integration — Inductor GPU device registration
Contributions are welcome in these areas:
- Operators: Add missing operators or optimize existing implementations for specific backends.
- Runtime and platform integration: Port torch-fl to new accelerators or improve existing backend support.
- Compiler integration: Extend torch.compile support, improve Triton codegen, or add new compilation backends.
- Distributed and profiler work: Enhance collective communication backends or profiling integration.
- Tests: Add operator correctness tests, model integration tests, or performance benchmarks.
- Documentation: Improve setup guides, troubleshooting docs, or vendor-specific integration notes.
See CONTRIBUTING.md for development workflow, code generation procedures, testing requirements, and PR guidelines.
All GitHub-facing text must be written in English — PR titles, PR descriptions, commit messages, issue text, and code review comments. This repository's code, comments, and history are in English, and PRs are reviewed by contributors who do not read other languages.
torch-fl builds on several upstream projects:
- PyTorch — The PrivateUse1 extension mechanism and ATen operator surface
- FlagGems — Triton-based portable operator kernels
- Triton (via FlagTree and vendor distributions) — GPU kernel compilation infrastructure
- FlagCX — Heterogeneous collective communication library
- Vendor runtime and operator libraries — ACLNN (Ascend), mudnn (MUSA), topsaten (GCU), cu-bridge (MetaX), DTK (DCU), and others
These acknowledgements do not imply endorsement by the named projects or vendors.
torch-fl is licensed under the Apache License 2.0.