Tensor Engine clockwork: one matmul on Trainium, cycle by cycle
An interactive walk through what the Trainium/Inferentia Tensor Engine does on every clock of one nc_matmul: weights parked in a systolic array, inputs streamed in, partial sums falling into PSUM. Plus how the same picture shows up in a Neuron Explorer profile.
I’ve been working through a roofline curriculum for AWS Trainium: take an op and a chip, work out from the docs how fast the op can run, measure how fast it does run, and explain the gap.
The first number in a roofline is P, peak compute: the most floating-point operations per second (FLOPS) the hardware can do. For the Tensor Engine, the part of a NeuronCore that runs matrix multiplies, P comes from three facts in the architecture guides:
- The engine is a systolic array with 128 rows and 128 columns of processing elements (PEs).1 That’s 128 × 128 = 16,384 PEs.
- Each PE does one multiply-accumulate (MAC) per cycle: multiply two numbers and add the result to a running sum. One MAC is 2 FLOPs.
- The Tensor Engine clock is 2.8 GHz on Trainium/Inferentia22 and 2.4 GHz on Trainium2.3
P = 16,384 PEs × 2 FLOPs per MAC × clock
= 91.75 TFLOPS bf16 on Trn1 / Inf2 (2.8 GHz)
= 78.6 TFLOPS bf16 on Trn2 (2.4 GHz)
The docs round these to 92 TFLOPS4 and 79 TFLOPS.5
That formula assumes every PE does a MAC on every cycle. I wanted to see when that’s true and when it isn’t, so I built a toy version of the Tensor Engine and stepped through it one clock at a time.
The setup
In NKI, a matmul on the Tensor Engine is one instruction, nisa.nc_matmul(stationary[K,M], moving[K,N]), which computes stationary.T @ moving6 and writes an [M, N] result to PSUM, the Tensor Engine’s output buffer.
- stationary
[K, M]is loaded into the array first (LoadStationary) and stays put. PE(k, m) holdsS[k, m]for the whole instruction. - moving
[K, N]is streamed through it (MultiplyMoving), one column per cycle. - K, the dimension being summed over, sits on the SBUF partitions, which feed the PE rows. M maps to the PE columns, which is why the stationary free size is capped at 128.7 Each column drains into one PSUM partition, so the PSUM partition dimension is rotated 90 degrees compared to SBUF.8 N is time: one column of the moving tensor enters per cycle. N is capped at 512 because each
nc_matmulwrites to a single PSUM bank, which holds 512 FP32 elements per partition.9 In the widget below, the PSUM grid is that one bank: one column per partition m, one cell per n.
Each PE does one MAC per cycle: take the partial sum from the PE above, add its weight times the value arriving from the left, and pass the result down. Row k gets its inputs k cycles late (the skew), so the partial sum falling into a row and the matching input arrive at the same moment. Moving value X[k, n] reaches PE(k, m) at cycle t = n + k + m.
Step through it
A 3×3×3 matmul does 27 MACs in 7 cycles on 9 PEs: 43% utilization. The array spends cycles 0–2 filling up and cycles 4–6 draining, and only peaks at 7 of 9 PEs in the middle. One instruction takes N + K + M − 2 cycles end to end.
Why real kernels fall short of peak
Scale it to 128 × 128 with N = 512 and the fill and drain shrink to a sliver, and back-to-back instructions overlap one’s drain with the next one’s fill. That’s the regime where P is real. Every way a kernel falls below it is a part of this picture going dark:
- K < 128 leaves PE rows idle. K = 64 uses half the array: 50% of P.
- M < 128 leaves PE columns idle, same story.
- Short N. The time between back-to-back matmuls is roughly
max(N, MM_INIT_LATENCY), whereMM_INIT_LATENCYis 64 cycles on NeuronCore-v2.10 A moving tile with N = 32 still pays for 64: 50% of P. - fp32 inputs. The instruction costs roughly 4× more than bf16.11
- Stalls between instructions: waiting on data, on the next stationary load, or on another engine to read PSUM.
On a roofline plot, each of those is a lower ceiling under the flat compute roof.
In a real profile
Neuron Explorer shows the Tensor Engine’s two instructions as two rows: LDWEIGHTS on the Tensor row and MATMUL on the TensorMatrix row.12 Here they are for a 2048 × 2048 × 2048 bf16 matmul on inf2. Every gap in the TensorMatrix row is time when the whole array in the widget is dark. The full hand count and profile for this matmul are in this notebook, and a SwiGLU MLP gets the same treatment.

Caveats
The mapping (what’s stationary, what streams, where K, M and N live, the issue rate) comes from the NKI Trainium/Inferentia2 architecture guide.13 The exact per-PE cycle timing is the textbook weight-stationary model. AWS doesn’t publish the array’s internals at that level, so treat the clock-by-clock part as a model of the hardware, not a trace of it.
William Pan’s How Systolic Arrays Work is a great companion: same dataflow, a 2×2 example with a step-through simulator.
Sources
Footnotes
-
Trainium/Inferentia2 architecture guide: “a systolic array with 128 rows and 128 columns of processing elements” ↩
-
Trainium/Inferentia2 architecture guide, engine table: Tensor, Frequency (GHz): 2.8 ↩
-
Trainium2 architecture guide, engine table: Tensor, Frequency (GHz): 2.4 ↩
-
Trainium/Inferentia2 architecture guide: “a maximum throughput of 92 TFLOPS” ↩
-
Trainium2 architecture guide: “158 FP8, 79 BF16/FP16/TF32 and 20 FP32 dense TFLOPS” ↩
-
Trainium/Inferentia2 architecture guide: “nki.isa.nc_matmul(stationary[K,M], moving[K,N]) performs a stationary.T @ moving calculation” ↩
-
Trainium/Inferentia2 architecture guide: “stationary tensor free axis size (stationary_fsize) must never exceed 128, due to the number of PE columns in TensorE” ↩
-
Trainium/Inferentia2 architecture guide: “PSUM partition dimension is purposely rotated 90 degrees compared to SBUF partition dimension due to systolic array data flow” ↩
-
Trainium/Inferentia2 architecture guide: “moving tensor free axis size (moving_fsize) must never exceed 512, due to the fact that each nc_matmul can only write to a single PSUM bank, which can only hold 512 FP32 elements per PSUM partition” ↩
-
Trainium/Inferentia2 architecture guide: “roughly max(N, MM_INIT_LATENCY), where MM_INIT_LATENCY is 64 TensorE cycles on NeuronCore-v2” ↩
-
Trainium/Inferentia2 architecture guide: “For FP32 input data type, the instruction cost is roughly 4x higher than BF16/FP16/TF32/cFP8” ↩
-
Neuron Explorer glossary, opcode table: “LDWEIGHTS: Load weight data for matmul” ↩