A simple design that shows how to use the packet switch feature of the AIE. It uses two cores, one to multiply and one to add two vectors together. The dataflow diagram is shown below:

There are three data paths in the design:
- Data path from the shim tile to the memory tile. It is assigned with two possible packet IDs: 0 and 1.
- Data path from the memory tile to the add compute tile (CT0). It is assigned with packet ID: 0
- Data path from the memory tile to the multiply compute tile (CT1). It is assigned with packet ID: 1
The initial packet ID is generated by the shim tile (runtime_sequence part) and the packet ID is passed to the memory tile. The packet id is kept in the memory tile therefore the total packet length is increased by 4 bytes. Then the packet is read out from the memory tile the AIE automatically uses the first 4 bytes as the packet ID. Therefore, the packet ID initiated by the shim tile is used again to determine where the packet should be sent (CT0 or CT1). Finally, the processed data is sent back to the host for verification.
During the compilation, only one xclbin is generated and is shared by both add and multiply operations. Two instruction files are generated with respect to the add and multiply configurations in the shim tile.
-
aie_add_placed.py: A Python file that describes the AIE design. The shim tile initializes the packet ID as 0 and therefore the packet will be sent to CT0 (doing add operation). -
aie_mul_placed.py: A Python file that describes the AIE design. The shim tile initializes the packet ID as 1 and therefore the packet will be sent to CT1 (doing multiply operation). -
host/host.cpp: This C++ code is a testbench for the design example targetting Ryzen™ AI (AIE-ML). It receives a flag 0 or 1 as an input argument to determine which operation to perform (0 to add and 1 to multiply).
To compile all files required for the design:
make allTo compile the design on npu1:
make all DEVICE=npu1To compile the design on npu2:
make all DEVICE=npu2To compile the C++ testbench:
make hostTo run the design (both add and multiply):
make runTo run the design with add operation:
make run_addTo run the design with multiply operation:
make run_mul