· 23 min read

Hardware entitlement to roofline: a 7-day curriculum

Take an op and a chip, work out from the documentation how fast the op can run, measure how fast it does run, and explain the difference. Seven days, built on AWS Trainium and Inferentia2, with a check-yourself problem for each.

trainiumnkirooflinehardware

By the end of the week you can take an op and a chip, work out from the documentation how fast the op can run, measure how fast it does run, and explain the difference.

  • Entitlement is the fastest the hardware allows for this op, as written. It is a number: a rate or a time.
  • The roofline is the plot of that number against the hardware’s two limits.

This assumes you know what a matmul, HBM and SBUF are, and how to run a kernel.

How the pages fit together

  • This page: the plan. Each day has a goal, how it works, what to read, what to do, and a problem to check yourself, with the answer folded underneath. Day 7’s check is your own kernel.
  • Worked examples: each day’s idea derived and worked through as one problem, with a plot.
  • How-to guide: the full procedure as 12 steps, with a plotting script. Keep it open on Day 7.
  • Case study: V-JEPA 2 MLP on Trainium2: all seven days applied to one MLP, checked against a hardware profile.
  • Two notebooks on real inf2 hardware: a 2048³ matmul and a SwiGLU MLP, each counted by hand, profiled and plotted.
  • Tensor Engine clockwork: an interactive, cycle-by-cycle look at where the Tensor Engine’s peak comes from (Day 3).

Notation

Symbol Meaning Units
P Peak compute for one dtype FLOP/s
BW Bandwidth of the slow link (HBM into on-chip memory) bytes/s
W Work. On Day 2, W is also the weight matrix in y = xW. FLOPs
Q Traffic across the slow link, counting every reload bytes
T Run time s
I Arithmetic intensity, W ÷ Q FLOPs/byte
I* Ridge point, P ÷ BW FLOPs/byte
f Clock frequency Hz
p Bytes per element (2 for bf16) bytes
B Batch: rows of activations (tokens) per matmul count
D, F A layer’s input and output widths count
M, K, N C = A·B with A of shape M×K and B of shape K×N. K is summed over. count
r_A, r_B Times each byte of A or B crosses the slow link count
LNC Physical cores per logical NeuronCore (2 on Trn2 by default) count
k Times a B×F intermediate crosses HBM: 0 fused, 4 when up-projection, GELU and down-projection run separately (write h, read h, write g, read g) count

The model in brief

A kernel does W FLOPs and moves Q bytes across the slow link. Computing takes at least W/P. Moving takes at least Q/BW.

T_best  = max(W/P, Q/BW)
T_worst = W/P + Q/BW

The best case assumes compute and data movement overlap fully. The worst case assumes no overlap. Dividing W by T_best gives the entitlement as a rate:

W/T ≤ min(P, I × BW),   I = W/Q

The two terms are equal at I* = P/BW. Below the ridge the kernel is memory-bound. Above it, compute-bound.

  • P and BW come from the hardware docs.
  • W comes from the op’s shapes.
  • Q comes from the code: how many times each byte crosses the link. Most mistakes are in Q.

Day 1: The roofline model

Goal. Draw the roofline from memory and place a kernel on it from W and Q alone.

How it works

  • On log-log axes, y = I × BW is a straight line of slope 1, because log y = log I + log BW. y = P is a flat line. Log axes are used because both quantities span orders of magnitude.
  • The ridge I* = P/BW is the number of FLOPs the machine can do in the time it takes to move one byte.
  • Q counts only bytes fetched from slow memory into local memory (cache on a CPU, SBUF on Trainium). A value still in local memory when it is needed again adds nothing to Q.
  • Blocking a matmul reorders its loops so values are reused while still local. W stays the same and Q goes down: same op, different code, different intensity.
  • Moving a point right means more FLOPs per byte, by reusing data. Moving it up means using more of the hardware at the same intensity.
  • No overlap at most doubles the time, since the sum of two numbers is at most twice the larger. A kernel more than 2× under its roof is losing time to something else, such as unused lanes or idle cores. Each can be drawn as a lower ceiling.

Read

  • Williams, Waterman & Patterson, “Roofline” (CACM 2009). Read for the bound, the log axes and lower ceilings.

Do

  • Select an op you use, such as a matmul, an elementwise add or a layernorm. Draw the roofline for the Day 1 worked-example machine, mark the ridge, and place the op with W and Q counted by hand. Done when you can say what the ridge means in words and which side your op is on.
  • Derive the intensity of a dot product, sparse matrix-vector multiply, and matmul untiled and blocked, and plot them. Done when they match the Day 1 worked example.

Check yourself

A machine has P = 1 TFLOPS and BW = 100 GB/s. A kernel does W = 2 GFLOPs and moves Q = 400 MB. Find I, the ridge, which roof binds, T_best and T_worst.

I = 2e9 ÷ 400e6 = 5 FLOPs/byte. Ridge = 1e12 ÷ 100e9 = 10. 5 < 10, so memory-bound.

W/P = 2e9 ÷ 1e12 = 2 ms. Q/BW = 400e6 ÷ 100e9 = 4 ms.

T_best = max(2, 4) = 4 ms. T_worst = 2 + 4 = 6 ms. Entitled rate = 2e9 ÷ 4e-3 = 0.5 TFLOPS, which is I × BW = 5 × 100e9.

Why is the slanted roof slope 1, and what are its units?

log y = log I + log BW, a line of slope 1. Units: FLOPs/byte × bytes/s = FLOP/s, the same as the flat roof.

Day 2: Rooflines for ML ops

Goal. Write FLOPs and bytes for transformer ops as formulas in their shapes, and say whether each is compute- or memory-bound before running it.

The setup in words

A linear layer computes y = xW.

  • x is the activations: B tokens, each a row of D numbers. Shape B × D.
  • W is the weight matrix: D rows, F columns. Shape D × F.
  • y is the output: B rows of F numbers. Shape B × F.
  • p is bytes per number (2 for bf16).

How it works

  • FLOPs. y has B × F numbers. Each is a dot product of length D: D multiplies plus D adds = 2D FLOPs. So W = 2BDF.
  • Bytes. Read x, read W, write y: Q = p(BD + DF + BF).
  • Intensity. I = W ÷ Q. When B is small, the weights dominate Q. Each weight is fetched once (p bytes) and used once per token (2B FLOPs), so in bf16 I ≈ B.

Make it real: one 4096 × 4096 projection on one physical Trn2 core

P = 78.6 TFLOPS, BW = 0.375 TB/s and ridge ≈ 210, derived on Day 6. D = F = 4096, bf16 (p = 2).

Decode, B = 1 (one token at a time):

  • W = 2 × 1 × 4096 × 4096 = 33,554,432 FLOPs
  • Q = 2 × (1×4096 + 4096×4096 + 1×4096) = 2 × 16,785,408 = 33,570,816 bytes
  • I = 33,554,432 ÷ 33,570,816 ≈ 1.0 FLOPs/byte. Ridge is 210, so memory-bound.
  • W/P = 33.55e6 ÷ 78.6e12 = 0.43 µs. Q/BW = 33.57e6 ÷ 0.375e12 = 89.5 µs.
  • T_best = max(0.43, 89.5) = 89.5 µs. The Tensor Engine is busy 0.43 of those 89.5 µs, under 0.5%.

Prefill, B = 2048 (a whole prompt at once):

  • W = 2 × 2048 × 4096 × 4096 = 68.72e9 FLOPs
  • Q = 2 × (2048×4096 + 4096×4096 + 2048×4096) = 2 × 33,554,432 = 67,108,864 bytes
  • I = 68.72e9 ÷ 67.11e6 = 1024, so compute-bound.
  • W/P = 68.72e9 ÷ 78.6e12 = 874 µs. Q/BW = 67.11e6 ÷ 0.375e12 = 179 µs. T_best = 874 µs.

Where it crosses the ridge. With D = F, I = BD ÷ (2B + D). Set I = 210:

  • B × 4096 = 210 × (2B + 4096)
  • 4096B − 420B = 860,160
  • B = 860,160 ÷ 3676 ≈ 234

Below about 234 tokens per call this layer is memory-bound. Above it, compute-bound.

What changes it:

  • Batching decode requests raises B, so each fetched weight is used more times.
  • int8 weights halve the weight bytes. Decode I goes from about 1 to about 2. Still memory-bound, but T_best halves.
  • The KV cache in decode attention is read once per generated token and used once, so I ≈ 1 regardless of batch. That is why long-context decode is memory-bound.
  • A link between chips is another roof with its own bandwidth and bytes. Entitlement is the minimum over all roofs.

Read

Do

  • Redo the three calculations above (B = 1, B = 2048, the crossover) without looking. Done when your numbers match.
  • Repeat for an MLP up-projection from a model you work with (D = model width, F = MLP width). Done when you have its crossover B.
  • Count decode attention with a KV cache of length L. Done when you can say which tensor dominates Q and why.
  • Add the interconnect as a roof for a matmul split across two chips. Done when you can state when the transfer sets the time.

Check yourself

On one physical Trn2 core (ridge 210), at what batch does a 2048 × 2048 bf16 projection become compute-bound? Is that more or less than the 4096 × 4096 case, and why?

I = BD ÷ (2B + D) with D = 2048. Set I = 210: 2048B = 210(2B + 2048), so 2048B − 420B = 430,080 and B = 430,080 ÷ 1628 ≈ 264.

More than 234. A smaller D means the input and output rows are a bigger share of Q relative to the weights, so it takes a bigger batch to reach the same intensity.

Day 3: P and BW from the docs

Goal. Derive both roofs for one Trainium1 NeuronCore from its architecture guide, and name why real kernels fall short of them. Day 6 repeats this for Trainium2.

Terms

  • PE (processing element): one multiply-accumulate unit.
  • MAC (multiply-accumulate): a × b added to a running sum. 2 FLOPs.
  • Systolic array: a grid of PEs. Each does one MAC per cycle and passes values to its neighbor. The Tensor Engine is a 128 × 128 systolic array. Tensor Engine clockwork steps through one matmul on it, cycle by cycle.
  • Partition: one of SBUF’s 128 rows. Each engine lane reads one partition.
  • Data-path width: how many elements an engine processes per cycle.
  • Stationary and moving tiles: a matmul instruction loads a stationary tile (up to 128 × 128) into the array and streams a moving tile (up to 128 × 512) through it. N is the moving tile’s free size.

How it works

  • Tensor Engine P = PEs × 2 FLOPs × clock = 128 × 128 × 2 × 2.8e9 = 91.75 TFLOPS bf16. The guide documents 92. fp32 is documented at 23, a quarter of that.
  • Vector Engine = 128 elements/cycle × 1.12e9 = 143e9 elements/s. A step that runs here, such as a cast or an add, has this roof, not the Tensor Engine’s.
  • BW. The device has 820 GB/s shared by 2 cores, so 820 ÷ 2 = 410 GB/s per core.
  • Ridge = 91.75e12 ÷ 0.41e12 = 224 FLOPs/byte. The profiler’s Overall Summary shows its own P ÷ BW as Peak Flops Bandwidth Ratio: 223.78 on inf2. The NKI matmul tutorial gives 222.
  • Why real kernels fall short:
    • Small stationary tiles. A matmul contracting over K = 64 fills 64 of the 128 PE rows, so half the array does nothing.
    • Small moving tiles. The guide gives the time between back-to-back matmuls as about max(N, 64) cycles on NeuronCore-v2. A moving tile with N < 64 still costs 64 cycles.
    • Wrong dtype. fp32 inputs run at a quarter of the bf16 rate.
    • Small transfers. Many small DMA transfers reach less of the 410 GB/s than a few large ones. Day 4 measures this.

Read

Do

  • Derive Tensor Engine P for bf16 and fp32. Done when you get 92 and 23.
  • Compute the Vector Engine’s elements/s. Done when you can say how many times slower one Vector op per element is than the Tensor Engine.
  • Compute per-core BW and the ridge. Done when you get 224.

Check yourself

On Trn1, what is the effective Tensor Engine rate for back-to-back bf16 matmuls with K = 64 and N = 512? With K = 128 and N = 32?

K = 64: 64 of 128 PE rows in use, so 91.75 × 64/128 = 45.9 TFLOPS. N = 512 is above 64, so there is no extra loss.

K = 128, N = 32: each instruction costs max(32, 64) = 64 cycles but does 32 cycles of work, so 91.75 × 32/64 = 45.9 TFLOPS. Each instruction takes 64 ÷ 2.8e9 = 22.9 ns.

Day 4: Spec vs measured

Goal. Measure both roofs on real hardware and explain each difference from the spec numbers.

This day assumes you can run NKI kernels on a Trn instance and capture a profile with Neuron Explorer.

How it works

  • A spec roof is what the parts allow. A measured roof is what a simple, well-written kernel reaches. Label which one you use.
  • Bandwidth roof. Write a kernel that only copies a large tensor HBM → SBUF → HBM. It does no FLOPs, so its time is all data movement. BW = bytes moved ÷ time from the profile.
  • Compute roof. Load a large bf16 matmul’s operands into SBUF once, then run the matmul many times in a loop. Time only the loop, so no HBM traffic is in the timed region. P = 2MNK × loop count ÷ time.
  • Transfer size. Small copies reach less bandwidth than large ones. Run the copy at several sizes and plot BW against size. The level where it flattens is the roof to use.
  • Shared HBM. Cores on one device share HBM. Run the copy on one core, then on all cores that share it. If one core alone gets more than its even share, the per-core figure is pessimistic for one-core kernels.
  • If a measurement comes out above spec, a spec input is wrong, usually the clock.

Read

  • Neuron Explorer docs. Read for how to capture a profile and read the engine and DMA timelines in Neuron Explorer.

Do

  • Write the copy kernel, profile it and compute BW. Done when you have GB/s and the transfer size.
  • Sweep the transfer size and plot BW against it. Done when you can name the size where it flattens.
  • Write the resident-matmul loop and compute P. Done when you have TFLOPS and the % of spec.
  • Run the copy on one core, then on all sharing cores. Done when you know whether the even split holds.
  • Name a cause for each gap. Done when no gap is labeled only “overhead”.

Check yourself

A copy kernel on one Trn2 core reads 64 MiB from HBM and writes it back. The profile shows 400 µs. What is the measured BW, and what percent of the 375 GB/s per-core share is it?

Bytes moved = 64 MiB in + 64 MiB out = 134,217,728 bytes.

BW = 134,217,728 ÷ 400e-6 = 335.5 GB/s. 335.5 ÷ 375 = 89.5% of the spec share.

Day 5: Buffers, reuse and fusion

Goal. Count Q for any mapping of an op onto the hardware, and see how fusing ops changes Q.

How it works: reuse inside one op

  • For a matmul, W = 2MNK and is fixed by the shapes.
  • Each element of A is used N times and each element of B is used M times. If an element stays in SBUF across its uses, it crosses from HBM once. If it is evicted, it crosses again each time it is reloaded.
  • Q = p(r_A·MK + r_B·KN + MN). Loop order, tile sizes and SBUF capacity set r_A and r_B. Same op, different code, different Q.
  • Best case: every input crosses once. For a 1024³ bf16 matmul that gives I = 2·1024³ ÷ (2 × 3 × 1024²) ≈ 341.

How it works: reuse across ops (fusion)

When one op’s output feeds the next, that intermediate is either kept on chip or written to HBM and read back. Each write plus read is 2 crossings of its full size. Fusing the ops keeps it on chip. W does not change. Q does.

Make it real: an MLP on one physical Trn2 core

The setup in words. A transformer MLP block does three steps, one after the other:

  1. Up-projection. Multiply the activations x (B tokens × D) by weight matrix W_up (D × F). The output is h (B × F): each token widened from D to F numbers.
  2. GELU. Apply the GELU activation to every element of h. The output is g (B × F), the same shape.
  3. Down-projection. Multiply g by weight matrix W_down (F × D). The output is y (B × D), back to the original width.

h and g are the intermediates: they exist only to feed the next step. B = 2048 tokens, D = 1024, F = 4096, bf16 (p = 2). P = 78.6 TFLOPS, BW = 0.375 TB/s, ridge ≈ 210.

Work, the same either way. The up-projection is 2BDF and the down-projection is 2BFD, so W = 4BDF = 4 × 2048 × 1024 × 4096 = 34,359,738,368 FLOPs. GELU’s FLOPs are small enough to ignore.

Unfused, as three separate ops:

Step Reads Writes
Up-projection x (BD), W_up (DF) h (BF)
GELU h (BF) g (BF)
Down-projection g (BF), W_down (FD) y (BD)
  • Total elements = 2BD + 2DF + 4BF. The 4BF is h and g, each written once and read once (k = 4).
  • Q = 2 × (4,194,304 + 8,388,608 + 33,554,432) = 92,274,688 bytes
  • I = 34,359,738,368 ÷ 92,274,688 = 372

Fused, as one kernel:

  • Both weight matrices are loaded into SBUF once: 2 × 1024 × 4096 × 2 B = 16 MiB, which fits in the 28 MiB SBUF.
  • For each block of 128 tokens the kernel runs the up-projection, GELU and down-projection before moving to the next block, so only one block of h is on chip at a time and h and g never go to HBM (k = 0).
  • Total elements = 2BD + 2DF
  • Q = 2 × (4,194,304 + 8,388,608) = 25,165,824 bytes
  • I = 34,359,738,368 ÷ 25,165,824 = 1365

Time:

  • W/P = 34.36e9 ÷ 78.6e12 = 437.1 µs for both
  • Fused: Q/BW = 25.17e6 ÷ 0.375e12 = 67.1 µs. T_best = 437.1 µs. T_worst = 437.1 + 67.1 = 504.2 µs.
  • Unfused: Q/BW = 92.27e6 ÷ 0.375e12 = 246.1 µs. T_best = 437.1 µs. T_worst = 437.1 + 246.1 = 683.2 µs.
  • The GELU op alone moves 2BF × 2 = 33.55 MB and does almost no math: 33.55e6 ÷ 0.375e12 = 89.5 µs of data movement.

Both are right of the ridge, so both are entitled to 437.1 µs. Fusion removes 67.1 MB of Q, the round trips of h and g. Day 7 shows what that changes. The case study repeats this on an LNC=2 logical core.

Read

  • Sze, Chen, Yang & Emer, Efficient Processing of DNNs, chapters on kernel computation and dataflow. Read for which operand stays in the buffer and why.

Do

  • Write a matmul loop nest under two tilings. Done when they differ in what stays in SBUF.
  • Count r_A, r_B and I for each. Done when W is identical and Q is not.
  • Redo the MLP numbers above without looking. Done when they match.
  • Count Q fused and unfused for an op chain from a model you work with. Done when you can name each intermediate and its k.

Check yourself

One attention head: L = 4096 tokens, head size d = 128, bf16. Compute S = QKᵀ (L × L), A = softmax(S), O = A·V. Find W, then Q and I unfused (S and A each written and read) and fused (S and A stay on chip). Which side of the 210 ridge is each?

W: QKᵀ is 2L²d and A·V is 2L²d, so W = 4L²d = 4 × 4096² × 128 = 8,589,934,592 FLOPs (softmax ignored).

Q, K, V and O are each L × d × 2 B = 1 MiB. S and A are each L × L × 2 B = 32 MiB.

Unfused: 4 MiB + 4 × 32 MiB = 132 MiB = 138,412,032 bytes. I = 8.59e9 ÷ 138.4e6 = 62. Memory-bound.

Fused: 4 MiB = 4,194,304 bytes. I = 8.59e9 ÷ 4.19e6 = 2048. Compute-bound.

Day 6: The Trainium entitlement sheet

Goal. One page of Trn2 roofs, with every number derived or cited.

How it works

  • P. 128 × 128 PEs × 2 FLOPs × 2.4e9 = 78.6 TFLOPS bf16 (documented 79). The Trn2 guide documents 20 fp32, 79 bf16 and 158 fp8. Roofs and ridges are per dtype.
  • Matmul issue rate. The Trn1 guide gives max(N, 64) cycles per back-to-back matmul on NeuronCore-v2. The Trn2 guide does not restate it, so check it in your profile.
  • Layout. The Tensor Engine contracts over partitions, so K must be on SBUF’s 128 partitions. A matrix stored M × K must be pre-transposed or loaded with dma_transpose. That costs DMA time but no extra HBM bytes.
  • PSUM. 128 partitions × 8 banks × 512 fp32 values = 2 MiB. One 128 × 512 matmul output fills one bank. Only the Vector and Scalar engines can read PSUM. The Tensor Engine only writes to it.
  • Other engines. The Trn2 guide documents the Vector Engine at 1.0 TFLOPS fp32. A step there has that roof.
  • BW. 3 TB/s per device shared by 8 cores: 3 ÷ 8 = 375 GB/s per core. Ridge = 78.6 ÷ 0.375 ≈ 210. This is the weakest number on the sheet: the guide’s memory-hierarchy figure gives ~0.5 TB/s per core, which would put the ridge near 157. Use your Day 4 measurement next to it.

LNC=2

  • Two physical cores act as one logical core. They share an HBM bank and address space. Each has its own SBUF, PSUM and Tensor Engine.
  • kernel[2] runs one instance per physical core. Anything loaded unconditionally is loaded on both cores, so it crosses HBM twice.
  • Scope: P = 2 × 78.6 = 157.3 TFLOPS and BW = 0.75 TB/s, so the ridge is still ≈ 210. Using whole-chip numbers (8 cores) for one logical core (2 cores) overstates the roof by 4×.

Read

Do

  • Build the ceiling table: throughput per engine and dtype, SBUF and PSUM sizes, HBM BW, DMA limits. Done when every number has a source.
  • Compute the bf16 and fp32 ridges. Done when you have recorded which BW figure you used.
  • Write the tile limits: partitions, free size, PSUM bank. Done when you can say why one output tile fills one bank.
  • State your scope: one physical core or one LNC=2 core. Done when P, BW and the per-core loads all use the same scope.

Check yourself

An LNC=2 kernel computes y = xW with B = 2048, D = 1024, F = 4096, bf16. Each core handles half the tokens but loads all of W unconditionally. Find Q and I. What I would you get if you counted W once by mistake?

x: 2048 × 1024 × 2 B = 4 MiB, split between cores, so read once. y: 2048 × 4096 × 2 B = 16 MiB, written once. W: 1024 × 4096 × 2 B = 8 MiB, loaded by each core, so 16 MiB.

Q = 4 + 16 + 16 = 36 MiB = 37,748,736 bytes. W = 2 × 2048 × 1024 × 4096 = 17,179,869,184 FLOPs. I = 17.18e9 ÷ 37.75e6 = 455.

Counting W once: Q = 28 MiB = 29,360,128 bytes, I = 585. That overstates I by 29%.

What does a PSUM eviction cost in HBM bytes?

Zero. It stays on chip. It costs Vector or Scalar engine time. One full tile is 128 × 512 × 4 B = 256 KiB.

Day 7: Capstone on your own kernel

Goal. Entitle one of your NKI kernels by hand, measure it, plot it, and explain the gap. Then change one thing and predict where the point moves.

The two inf2 notebooks are this day done end to end: a matmul whose prediction failed because the compiler moved fewer bytes than the source implies, and a SwiGLU MLP whose prediction held but still runs well under peak.

How it works

  • Entitlement = min(P, I × BW), with P and BW from the Day 6 sheet and I from your hand count.
  • Gap = T_measured − T_best. Don’t infer its cause. Find it on the profile timeline. For a compute-bound kernel, W ÷ T_measured ÷ P is the kernel’s MFU.
  • If entitlement is low, the mapping moves too many bytes: fix Q. If entitlement is high and measured is far below it, the waste is in execution.

Same W and same P give the same entitlement. Two mappings of one op that are both right of the ridge are entitled to the same time, W/P.

Where they differ, using the Day 5 MLP:

  • Bandwidth needed to stay compute-bound is P/I. Fused: 78.6e12 ÷ 1365 = 58 GB/s. Unfused: 78.6e12 ÷ 372 = 211 GB/s. Higher I needs less.
  • With no overlap, time is W/P + Q/BW. Fused: 437.1 + 67.1 = 504.2 µs. Unfused: 437.1 + 246.1 = 683.2 µs. Both have a best case of 437.1 µs.

Causes of a gap, and where to see each in the profile:

  1. Tensor Engine idle: gaps between matmul instructions on the Tensor Engine row.
  2. No overlap: DMA and matmuls take turns instead of running at the same time.
  3. Extra bytes: the profiler’s DMA bytes are higher than your Q.
  4. Slower engine: a Vector or Scalar instruction runs while the Tensor Engine waits on it. The Trn2 Vector Engine is documented at 1.0 TFLOPS fp32, against 79 TFLOPS bf16 for the Tensor Engine.
  5. Fixed cost: time before the first matmul starts and after the last one ends.
  6. Tensor Engine busy but slow: the Tensor Engine row is solid, but W ÷ busy time is under P. Check tile sizes and the Throttling row. The profiler’s summary can’t separate small tiles, instruction placement, fast weight load and throttling.

Do

  • Pick a kernel. Done when shapes and dtype are written as numbers.
  • Count W from the shapes. Done when it matches matmul instruction count × size.
  • Count Q as a table of transfers, with size and count, including stores. Done when the table shows the resident data fits in SBUF.
  • Compute I and entitlement. Done when you have T_best = max(W/P, Q/BW) and T_worst = W/P + Q/BW.
  • Profile with Neuron Explorer, skipping the first run. Done when you have T_measured from the device timeline.
  • Compare your W, Q and I with the profiler’s Overall Summary: Raw Flops, HBM read + write bytes, and Mm Arithmetic Intensity. If the kernel transposes on the Tensor Engine, Raw Flops includes the Transpose Flops and Mm Arithmetic Intensity leaves them out. Done when each difference has a reason.
  • Plot the roof, the ridge, the entitlement point and the measured point. Done when the measured point sits directly below the entitlement point.
  • Find the gap on the profile timeline. Done when every part of it points to a region of the timeline, and the parts add up to the whole gap.
  • Change the tiling once. Predict the new point, then measure. Done when you can say whether the prediction held, and if not, what Q missed.

Common mistakes

Mistake Effect Catch it
Counting tensor sizes, not bytes moved Ignores reloads, so overstates I Walk the loop nest and count each transfer × times
Spec BW the access pattern can’t reach Overstates entitlement left of the ridge Measure BW with a copy kernel (Day 4)
Whole-model latency vs one op’s roof Includes other ops and host time Entitle and measure one op from the device timeline
Reading a gap as wasted compute when another engine binds The binding roof isn’t on the plot Entitle each step against its own engine
Wrong LNC scope Wrong roof, and per-core loads counted once Check the LNC setting and count per-core loads once per physical core
Sources