This tutorial takes the stock MobileNetV2 QC-Ultra kernel (raw CGC output, 484,072 cycles in sdk-cli 26.09) and makes it 13.7% faster in four steps. Every step comes from reading the GPNPU timeline, and none of them changes arithmetic. For each step we zoom into the trace, say what it shows, make the smallest code change that addresses it, and re-trace the same window to check the change did what we expected. Every kernel in this document is byte-exact under both sdk-cli 26.09 (traced) and 26.08. Everything needed to reproduce it is in examples/timeline_tutorial/ in the SDK (see How to reproduce).
| step | what the timeline showed | change | 26.09 cycles | 26.08 cycles |
|---|---|---|---|---|
| stock | – | – | 484,072 | 512,612 |
| A | R7 waits 43.6K for a DMA issued in R6; DDR idle in R2–R3 | issue the FC-weight memCpys at R2 start | 440,675 (−43,397) | 482,301 (−30,311) |
| B trap | core blocked on the 9th memCpy (8-deep ELS queue) | all 48 later weight memCpys at R2 start | 524,920 (+84,245) | 566,281 |
| A+B | same idle DDR, but the queue has 8 slots | 4 coalesced group DMAs + setFlowId | 438,883 (−1,792) | 480,313 (−1,988) |
| A+B+C | R1 quantize waits for the last byte of a 588 KB DMA | 7 row bands, one DMA and one read flow each | 429,728 (−9,155) | 471,489 (−8,824) |
| D trap | R6 +11.4K: its own loads queue behind the chunks | 4 FC chunks issued at R6 start | 429,233 (−495) | 470,997 |
| A+B+C+D | R7 is a 1.25 MB OCM→LRM load with nothing to overlap | 4 no-wait chunks spread over R6, barrier in R7 | 417,680 (−12,048) | 460,774 (−10,715) |
| total | −66,392 (−13.7%) | −51,838 (−10.1%) |
Rows in italics are measured traps: changes that look right and are slower. They are part of the lesson. Deltas are against the previous non-trap row.
1. Why a timeline
The ISS profile (profile.json) gives per-region counters: cycles, stall reasons and buckets such as data_external or mac. Counters say how much time went where. They cannot say what happened at the same time, and on a GPNPU that is usually the question. The array, the DDR↔OCM engine (ELS), the OCM↔LRM engine (PLS), the queue that feeds LRM (QLS) and the instruction fetch unit all run concurrently. A region with 40K cycles of DDR traffic and 40K cycles of compute can take 40K cycles (perfect overlap) or 80K (fully serialized), and its counters look the same.
The timeline records every engine's busy intervals cycle by cycle. From it you can read:
- Execution sequence: what the kernel actually did, in order, including DMAs that were issued long before they ran.
- Overlap: whether compute and data movement happen at the same time, or one waits for the other.
- Coordination between hardware components: the core blocking on a queue, a PLS load waiting for an ELS load, an in-order queue holding a small transfer behind a big one.
- Dependencies: every DMA slice carries its
flow_idand thedependencyit waited on. - Payload: bytes and achieved bytes-per-cycle for every transfer.
Every step below starts from a picture that no counter shows.
2. Reading the trace
Capturing
sdk-cli 26.09 added the timeline to sdk source (and to sdk graph compile --run). The flags are --timeline PATH, --timeline-level coarse|standard|full (default coarse) and --timeline-max-bytes (per-core size cap, default 256M); sdk source --help lists them:
sdk source ... --timeline OUT.json --timeline-level coarse|standard|full
examples/timeline_tutorial/run_kernel.py builds and runs a kernel through sdk source with this tutorial's fixed QC-U configuration, checks the output byte-exact against the golden, and prints the cycle count. --timeline adds the two flags above, so you get the byte-exact check and the cycle count plus a trace:
## in examples/timeline_tutorial; ./run.sh starts an sdk-cli:26.09 container
./run.sh 26.09 python3 /tut/run_kernel.py --workdir /ws/work/stock \
--timeline /ws/work/tl_standard.json --timeline-level standard
## {"pass": true, "cycles": 484072, "exact": true}
Tracing does not perturb the simulation: stock measures 484,072 at every level and untraced.
If you run the host program yourself instead of through sdk source, attach the timeline producer (the ISS iss_ingestors.timeline ingestor) to the ISS event stream over a FIFO. The reader must be attached before the ISS starts, and it must not exit early:
mkfifo /tmp/iss_events
python3 -m iss_ingestors.timeline --fifo /tmp/iss_events -o timeline.json &
./<proj>_host ... --eventStreamTarget stream:/tmp/iss_events
wait; rm -f /tmp/iss_events
Levels and cost (stock MobileNetV2, 484K cycles)
| level | trace | wall (untraced 26.4 s) | what you get |
|---|---|---|---|
| coarse | 3.4 MB | 31.8 s | ELS / PLS lanes, one compute lane, data-hazard stalls, regions |
| standard | 26.8 MB | 35.1 s | + MAC, ALU, Other/blocked, QLS, IFU (and MLS / RMLS tracks) |
| full | 40.5 MB | 34.8 s | + ALU split into multiply / bitwise / reg-move / branch / flow-setup |
Use standard for this kind of work: coarse folds blocked cycles into "compute" (see caveats). All images in this tutorial are standard-level traces.
The lanes
| lane (image row) | trace track | what a slice means |
|---|---|---|
| DDR→OCM ELS load | DDR→OCM (ELS Load) | one memCpy DDR→OCM: bytes, bandwidth_B_per_cycle, flow_id |
| OCM→DDR ELS store | OCM→DDR (ELS Store) | a memCpy OCM→DDR (here: only the 4 KB output) |
| OCM→LRM PLS load | OCM→LRM (PLS Load) | a read flow into the PE array's local memory; dependency = the flow it waited for |
| LRM→OCM PLS store | LRM→OCM (PLS Store) | a write flow from LRM back to OCM |
| QLS FIFO↔LRM | QLS (FIFO↔LRM) | the load/store queue moving PLS data into / out of LRM; in order |
| (MLS, RMLS) | MLS (core→core) | core-to-core moves (empty in this kernel) |
| IFU I-fetch | IFU (instr fetch) | instruction-cache line fills (4 KB, 354 cycles each) |
| MAC | MAC | MAC-array instructions (slice args: pc, inst, PE usage) |
| ALU / reg / branch | ALU / Compute (+ full-level splits) | everything else the core executes |
| Blocked (Other: dbar, nrb…) | Other / Unclassified | the core is issuing but waiting: DMA barriers, flow waits, a full queue |
| Data-hazard stalls | Stalls (data hazard) | pipeline stalls with a reason (SPRF_RAW, DEREF_REGPTR_CONFLICT, …) |
| region labels | Profile Regions | the kernel's cgc::profileStart/Stop markers |
Caveats
- An ELS slice spans enqueue → end, so it includes queue time. A 4 KB bias load shows up as a 90K-cycle slice because it sat behind the 1.25 MB weights. For transfer time, use
bytes / peak: this configuration measures 19.08 B/cycle on the 256-bit bus this tutorial uses (the stock kernel's 602,112 B input copy takes 31,553 cycles). The slice'sbandwidth_B_per_cycleis bytes / slice length, so a low value means queued, not slow. - PLS slices include dependency wait too. A PLS load issued before its source DMA finished starts at issue and ends when its data lands.
- The coarse compute lane counts blocked cycles as busy. At standard level, waiting moves to Other / Unclassified ("Blocked" in the images), so "compute 74%" at coarse can be "Blocked 74%, ALU 0%" at standard (region 1 below).
- The ELS queue holds 8 descriptors. The 9th
memCpyblocks the core until one retires (step B). - At overview zoom, slices closer than a pixel merge, so a lane can look solid. Zoom in before drawing conclusions.
Opening a trace
Perfetto UI: open https://ui.perfetto.dev and drag the JSON onto the page. Each lane is a thread track; click a slice for its args (bytes, flow id, dependency, pc, instruction). Time is shown in µs at 1.7 GHz;
start_cycle/end_cycleare in the args.examples/timeline_tutorial/analyze.pyTRACE: a text summary of a trace: per-region busy percentage of every lane, DDR bytes completed per region, and the largest compute-idle gaps with what the other lanes were doing during them. The trace is plain Chrome-trace JSON (ph: "X"slices withstart_cycle/end_cycleargs), so it is easy to script against.The images in this tutorial come from
examples/timeline_tutorial/tlplot.py(matplotlib, one panel per trace, shared cycle window, fixed lane colors, annotation helpers), driven bymake_figures.py.
Locating a region
Every Rn in this tutorial is a slice you can click. The kernel's cgc::profileStart/Stop markers become a Profile Regions track with one slice per region, named region: n / 7 — so R6 is the slice region: 6 / 7. The track is present at every detail level, coarse included, so you do not need a bigger trace just to navigate one.
In the Perfetto UI, search for region: 6, select the slice and zoom to the selection. Its args carry the cycle window, which is where the numbers quoted throughout this tutorial come from:
| arg | region: 6 / 7, stock trace |
|---|---|
start_cycle | 316,440 |
end_cycle | 428,024 |
duration_cycles | 111,584 |
Reading the regions in order is also the quickest way to find where a kernel spends its time before zooming anywhere: R6 is 111,584 of the stock kernel's 484,072 cycles, or 23%. analyze.py prints exactly this, keyed by the same region: n / 7 names, so a label you read in the text, a slice you click in the UI and a row in the summary all refer to the same thing.
Regions come from the kernel source, so they survive your edits: after a change, the same region index still brackets the same code, which is what makes the before/after comparisons below meaningful.
3. The stock kernel's big picture

A per-region summary of the same trace:
DDR floor: 4.69 MB read (245,835 cycles), 0.00 MB written (210 cycles); reads and writes overlap,
so the run cannot go below ~245,835 cycles. It is at 2.0x that floor.
| region | cycles | kind | compute | DDR in | DDR-eq | comp∩DDR | DDR only | idle |
|---|---|---|---|---|---|---|---|---|
| 6 / 7 | 111,585 | balanced | 86% | 86% | 118% | 78% | 8% | 6% |
| 5 / 7 | 79,859 | balanced | 83% | 79% | 86% | 67% | 12% | 5% |
| 2 / 7 | 75,971 | mixed | 60% | 0% | 0% | 0% | 0% | 3% |
| 3 / 7 | 73,872 | mixed | 65% | 2% | 2% | 0% | 2% | 1% |
| 7 / 7 | 55,747 | mixed | 79% | 78% | 0% | 78% | 0% | 0% |
| 4 / 7 | 43,283 | mixed | 75% | 31% | 27% | 18% | 13% | 2% |
| 1 / 7 | 43,052 | balanced | 74% | 77% | 76% | 73% | 4% | 1% |
("compute" here includes blocked cycles. For R1 and R7 the image shows it is almost all Blocked, not work.)
What the picture says before any counter does:
- The floor. The kernel reads 4.69 MB from DDR. At 19.08 B/cycle that is 245.8K cycles, and no schedule can go below it without moving fewer bytes. Stock is at 2.0x the floor, so more than half the run is not DDR-limited.
- DDR is idle for ~86K cycles (R2–R3, 33.7K → 119.5K): the compiler issues each layer's weights just before the layer, and R2–R3 need only 24 KB.
- Just-in-time DMAs where the core then waits: R1 waits for the 588 KB input, and R7 waits 43.6K for the 1.25 MB FC weights that R6 requested.
- R5–R6 are balanced: DDR and compute are both busy and overlapped. Cutting either one alone buys little there.
The steps below go after the waits in order of size and use the idle DDR window to hide them.
4. Optimization steps
All steps are cumulative: A → A+B → A+B+C → A+B+C+D. Each shows the code change as an excerpt of the diff against the stock mnv2_qcu.cpp (full diffs in examples/timeline_tutorial/patches/; long template argument lists are shortened to … here).
Step A — overlap DMA with compute
Before. Zoom into R6–R7 (316K–486K):

What it tells you. In the top panel the ELS lane shows the 1,280 KB FC-weight load (const_tensor_53) and its 4 KB bias (const_tensor_54). The compiler issues them in R6, behind R6's own weights, so they are still downloading when R7 starts. R7's first PLS read (the bias, note the long PLS slice: dependency wait) cannot start until the DMA lands. The core sits Blocked for 43.6K cycles, and only then does the 1.25 MB OCM→LRM load (11.9K) run. Meanwhile the overview shows DDR idle for 86K cycles in R2–R3. The data is constant, so nothing forces a late issue: the only cost of an early issue is OCM space to hold it.
Change. Move the two memCpys to the start of R2 and give them OCM addresses above everything else the kernel places (the stock kernel's OCM use stays below 2.5 MiB, and the FC weights are the last consumer, so nothing overwrites them):
cgc::profileStart("region: 2 / 7");
+ /* issue the dense (FC) weights now; DDR is idle in regions 2-3 */
+ DdrTensor const_tensor_53_ext(pconstDdrTensor + 2773696);
+ OcmTensor const_tensor_53_ocm;
+ const_tensor_53_ocm.ptr = 2621440; /* re-homed above peak OCM use */
+ const_tensor_53_ocm.bound = TensorOffset::fromTensor();
+ memCpy(const_tensor_53_ext, const_tensor_53_ocm, 0, 0, 0, 0);
+ DdrTensor const_tensor_54_ext(pconstDdrTensor + 4084416);
+ OcmTensor const_tensor_54_ocm;
+ const_tensor_54_ocm.ptr = 3932160; /* right after the FC weights */
+ ...
+ memCpy(const_tensor_54_ext, const_tensor_54_ocm, 0, 0, 0, 0);
...
(region 6)
- DdrTensor const_tensor_53_ext(pconstDdrTensor + 2773696);
- OcmTensor const_tensor_53_ocm;
- const_tensor_53_ocm.ptr = 9280;
- ...
- memCpy(const_tensor_53_ext, const_tensor_53_ocm, 0, 0, 0, 0);
(same for const_tensor_54)
Correctness needs nothing extra: region 7's read flows depend on the tensors' flow ids, which the two memCpys set, so R7 still waits if the data had not arrived.
After (bottom panel, same window). No FC traffic is left in R6, and R7 drops from 55.7K to 12.2K. What remains is the 1.25 MB OCM→LRM load, with nothing overlapping it (step D).
Measured: 484,072 → 440,675 (26.09), 512,612 → 482,301 (26.08). Byte-exact on both.
Step B — dependencies and queue depth
Step A used the idle DDR window for 1.28 MB. If early issue works, why not issue every later weight early and get all remaining DDR traffic off the critical path?
Before and the trap. Zoom into R2 (40K–210K). Here the ELS lane is drawn as one row per outstanding descriptor:

- Top (after A): DDR carries the FC weights during R2, then is idle again from ~112K until R4 (~80K cycles). 46 weight memCpys (tensors 7–52, 2,753,984 B = 2.63 MiB) still wait for their regions.
- Middle (naive): all 48 weight memCpys (tensors 7–54: those 46 plus the two FC tensors) moved to the R2 start, each into its own OCM slot (
mk_prefetch_all.py). The descriptor rows fill to exactly 8. The ELS queue is 8 deep, so the 9thmemCpyblocks the core until the oldest transfer retires, and so does every one after it. The Blocked lane is solid from 44K to 130K and R2 compute cannot start until the last descriptor is queued. R2 goes from 76K to 162K, and the total to 524,920, worse than stock.
What the timeline taught: a memCpy is asynchronous only while the queue has room.
Change. In the blob, tensors 7–54 are contiguous in consumption order. Issue one memCpy per consumer group (4 descriptors: R3–R4 weights, R5, R6, FC), and point each member tensor into its group's OCM copy. The important line is setFlowId: each member's old memCpy is replaced by inheriting the group DMA's flow id, so every consumer (BroadcastFlow / read flow) still carries a hardware dependency on the DMA that actually fills it. Without it, the consumer would see a flow id of 0 (no dependency) and read OCM before the data lands, a race the timeline would not even show until outputs mismatch.
cgc::profileStart("region: 2 / 7");
+ /* coalesced weight prefetch, one DMA per consumer region, issued while DDR is idle */
+ DdrTensor gW34_ext(pconstDdrTensor + 19712);
+ OcmTensor gW34_ocm;
+ gW34_ocm.ptr = 2621440;
+ gW34_ocm.bound = TensorOffset::fromTensor();
+ memCpy(gW34_ext, gW34_ocm, 0, 0, 0, 0);
+ ... gW5 (1,308,928 B), gW6 (1,195,520 B), gWfc (1,314,720 B: const_tensor_53 + 54) ...
...
(region 3, and likewise for every tensor 7..54)
OcmTensor const_tensor_7_ocm;
- const_tensor_7_ocm.ptr = 1670592;
+ const_tensor_7_ocm.ptr = 2621440; /* inside gW34 */
const_tensor_7_ocm.bound = TensorOffset::fromTensor();
- memCpy(const_tensor_7_ext, const_tensor_7_ocm, 0, 0, 0, 0);
+ const_tensor_7_ocm.setFlowId(gW34_ocm.getFlowId()); /* filled by gW34 */
(mk_prefetch_groups.py applies it to the stock kernel. The FC tensors are the 4th group, so this step includes step A.)
After (bottom panel). 4 descriptors, at most 8 outstanding, so the core never blocks on issue and R2 is unchanged at 76K. From R3 on, every weight is already in OCM, R4 and R5 get a little faster, and the ELS lane goes quiet after 256K.
Measured: 440,675 → 438,883 (26.09), 482,301 → 480,313 (26.08). Byte-exact on both. The gain is small because R5–R6 were already balanced (their DMAs overlapped their compute). The lessons are in the trap, and in keeping the dependency correct. From here on DDR never limits the kernel, and the remaining waits are all local.
Step C — chunk a large transfer
Before. Zoom into R1 (0–50K), with ELS descriptors on separate rows:

What it tells you. R1 copies the 588 KB FP32 input (602,112 B) in one memCpy and runs one read flow over the whole tensor. The first PLS read (14 KB, the first tile) depends on the input's flow id, so it waits for the last byte: 31.6K cycles of Blocked at full DDR rate (19.08 B/cycle, the transfer itself is efficient). Only then do the 49 quantize tiles run (~11K cycles), with DDR idle. The work is a stream (tile rows are independent), but the single flow id makes it store-and-forward.
Change. Cut the input into 7 row bands of 32 rows (84 KB each). Each band is its own memCpy, so it gets its own flow id, and its own BufferedViewReadFlow, so quantizing band b waits only for band b's DMA. The loop body (same write flow, same tile order, same arithmetic) is unchanged. The seven small R2 weights are coalesced into one descriptor (gW0) so R1 issues 7 + 1 = 8 descriptors, exactly the queue depth, and the core never blocks on issue (step B's lesson).
- OcmTensor, 1, 3, 224, 224> input_input_ocm;
- input_input_ocm.ptr = 150528;
- memCpy(input_input, input_input_ocm, 0, 0, 0, 0);
+ using InBand = OcmTensor, 1, 3, 32, 224>;
+ constexpr int32_t IN_BAND_BYTES = 3 * 32 * 896; /* planar band: 3 channel planes of 32 rows, 896 B row pitch */
+ InBand in_band0;
+ in_band0.ptr = 150528 + 0 * IN_BAND_BYTES;
+ in_band0.bound = TensorOffset::fromTensor();
+ memCpy(input_input, in_band0, 0, 0, 0, 0); /* rows 0-31: memCpy(src, dst, b, c, h, w) offsets */
+ ... in_band1 .. in_band6 at rows 32, 64, ..., 192 ...
+ DdrTensor gW0_ext(pconstDdrTensor + 0); /* R2 weights, one descriptor */
+ ... const_tensor_0..6: ptr inside gW0, setFlowId(gW0_ocm.getFlowId()) ...
...
- BufferedViewReadFlow<…, OcmTensor, 1, 3, 224, 224>> input_input_ocm_flow_0(input_input_ocm, …);
- for (int32_t tb_y = 0; tb_y < 7; ++tb_y) {
- for (int32_t tb_x = 0; tb_x < 7; ++tb_x) {
- input_input_ocm_flow_0.read(input_input_ocm_lrm);
+#define TUTORIAL_QUANT_BAND(band) { \
+ BufferedViewReadFlow<…, decltype(band)> band_flow(band, cgc_core_stack_pointer, 0, 0, 0, 0, 0); \
+ for (int32_t tb_x = 0; tb_x < 7; ++tb_x) { \
+ band_flow.read(input_input_ocm_lrm); \
... quantize + ocm_tensor_0_flow_0.write(...), unchanged ...
+ TUTORIAL_QUANT_BAND(in_band0) ... TUTORIAL_QUANT_BAND(in_band6)
After (bottom panel). Band 0 lands at 5.0K and its tiles are quantized and written back (PLS store bursts) while band 1 is in flight. The pattern repeats per band, and only the last band's quantize is exposed after the DMA. R1 goes from 43.0K to 33.9K. The input's DDR time is the same 31.6K: we hid the work behind it rather than making the transfer faster.
Measured: 438,883 → 429,728 (26.09), 480,313 → 471,489 (26.08). Byte-exact on both.
Step D — overlap the tail
Before. After A+B+C, zoom into R6–R7 (300K–432K):

What it tells you (top panel). R7 is now 12.2K cycles of one thing: the FC layer's blocking UnbufferedReadFlow pulls the 320 × 1024 v4i8 weight matrix (1,280 B per PE) from OCM into LRM. PLS and QLS are busy, and MAC and ALU are idle. The data has been in OCM since step A (~256K), so the load could run during R6, which has gaps in its PLS and QLS lanes. Two things must hold:
- An LRM home that R6 does not touch. The compiler's core stack is 1,584 B and R6 uses all of it. We grow the stack allocation to 2,864 B and put the FC weights at offset 1,584, so nothing in R6 or R7 reuses it.
- A non-blocking issue. A flow
read()waits for its data. We useplsReadIssueNoWait(added toattention_stubs.hpp, the stub header CGC emits next to the kernel): it programs the same PLS read a flow would, starts it, records its flow id on the tensor and returns.barrier(tensor)waits for it later.
Change. Split the matrix into 4 chunks of 80 rows (320 B per PE each) and issue one chunk after each of four R6 all-to-alls (chunk 0 after the 2nd, then the 4th, 6th, 8th). Put a barrier on each chunk before the next all-to-all, and barriers on all four in R7 instead of the read. The bytes land at the same per-PE LRM offsets the stock read used (row r of chunk k is row 80k+r), and the FC arithmetic is untouched.
- qVar_t cgc_core_stack_pointer = allocate_core_memory(1584);
+ qVar_t cgc_core_stack_pointer = allocate_core_memory(2864); /* +1280 B LRM for the FC weights */
...
cgc::profileStart("region: 6 / 7");
+ using FcChunkOcm = OcmTensor;
+ using FcWAccess = TensorAccessor, …>;
+ using FcWFlow = UnbufferedReadFlow;
+ FcChunkOcm fcw0, fcw1, fcw2, fcw3;
+ fcw0.ptr = 5375424 + 0 * 80 * 1024 * 4; fcw0.setFlowId(gWfc_ocm.getFlowId()); /* ... fcw1..3 */
+ container::NDArrayView, 320> fc_w_lrm0(cgc_core_stack_pointer, 1584); /* 1904, 2224, 2544 */
...
allToAllPartitionBcast<2, 40>(T_qlinear_conv2d_lrm23, T_vap_all_to_all_bcast_lrm15);
+ plsReadIssueNoWait(fcw0, fc_w_lrm0); /* chunk 0, no wait */
...
+ barrier(fcw0); /* chunk 0 has landed before the next all-to-all */
allToAllPartitionBcast<2, 240>(T_qlinear_conv2d_lrm24, T_vap_all_to_all_bcast_lrm16);
... (chunks 1, 2, 3 the same way)
cgc::profileStart("region: 7 / 7");
- UnbufferedReadFlow<…, OcmTensor> const_tensor_53_ocm_flow_0(const_tensor_53_ocm, 0, 0, 0, 0);
...
- container::NDArrayView, 1280> const_tensor_53_ocm_lrm(cgc_core_stack_pointer, 32);
- const_tensor_53_ocm_flow_0.read(const_tensor_53_ocm_lrm);
+ container::NDArrayView, 1280> const_tensor_53_ocm_lrm(cgc_core_stack_pointer, 1584);
+ barrier(fcw0); barrier(fcw1); barrier(fcw2); barrier(fcw3);
qVar_t _116 = nn::fullyConnectedTileBlock(const_tensor_53_ocm_lrm);
Trap 1 — issuing everything at once (middle panel). All four chunks issued at the start of R6 took the PLS and QLS lanes for ~12K cycles. The QLS serves requests in order, so R6's own first weight loads queued behind the chunks. R6 grows by +11.4K, which eats almost all of the 12K saved in R7 (429,233, −495). The timeline shows the push directly: the same work, shifted right. Spreading the chunks puts each ~2.9K transfer where the queue is otherwise quiet.
Trap 2 — issuing a chunk before an all-to-all (no image; measured). Placing each chunk right before the same four all-to-alls was byte-exact under 26.08 but not under 26.09 (417,689 cycles, wrong output). The timelines look alike, so this is where the byte-exact check earns its keep. Our reading: 26.09 rewrote allToAllPartitionBcast for rep=2 to drive all four neighbour ports at once (qAllNeighbors, sdk_install/src/core/local_data_movement.hpp), and a PLS stream still entering the array during it is not safe. This is a hypothesis from the source diff, not verified in the ISS. What we ship issues each chunk after an all-to-all and barriers it before the next one, so no chunk is ever in flight during an all-to-all. The guard costs nothing measurable (417,680 guarded vs 417,689 unguarded). Issuing all four chunks after the last all-to-all measured 417,623, so for this kernel the placement, not the chunking itself, is what wins.
After (bottom panel). The four chunk transfers (~2.9K each) sit in R6's PLS and QLS gaps. R6 is unchanged (111.2K), and R7 drops from 12.2K to 299 cycles (the bias read, the FC MACs and the 4 KB output DMA).
Measured: 429,728 → 417,680 (26.09), 471,489 → 460,774 (26.08). Byte-exact on both.
Where the total stands. 417,680 is 1.70x the 245.8K DDR floor, and DDR is now busy only during R1–R5. The final kernel against the stock one at the same scale:

5. Sidebars
Payload: count the bytes, not just the transfers
Every DMA slice carries bytes, so the timeline also answers "is this data necessary?". The FC weight matrix is 320 × 1024 v4i8, but the layer has 1,000 output classes. The last 24 columns are padding to the 1,024-PE array: 320 × 24 × 4 = 30,720 B (30.7 KB) of zeros, moved twice (DDR→OCM, then OCM→LRM). That is ~1.6K cycles of DDR time. After step A it is off the critical path, but on a DDR-bound kernel it would not be.
Bus width shows up the same way. run_kernel.py passes --ddr-axi-width 256; sdk source defaults to 128. A diagnostic-only run of the stock kernel at 128 bits (run_kernel.py --ddr-axi-width 128; it changes the simulated hardware, so its cycle count is not comparable with the others):

The 602,112 B input drops from 19.08 to 15.90 B/cycle: 128 bits × 1.7 GHz supplies 16 B/cycle, less than the 18.8 B/cycle that 32 GB/s needs. R1 grows 43.1K → 49.4K, and the whole run 484,072 → 516,785 (+6.8%). When a lane's achieved B/cycle sits at a round number below the configured peak, check the configuration before the kernel.
Instruction-fetch holes

Every launch starts with a cold instruction cache. The IFU lane shows 20 fills of 4 KB lines, 354 cycles each (7,080 cycles, 1.5% of stock). During a fill nothing issues, so each is a hole in every compute lane. The fills land where the program first crosses into a new line, often in the middle of a region.
A hand-tuned kernel can remove them: it places a jumped-over dret stub in each line and calls the stubs (djal with link save/restore) during region 1, while the core would otherwise wait for input DMA. The lines are then warm when execution reaches them, and IFU stalls in R2–R7 drop to zero. It is pure control flow (no data, LRM, OCM or flow state touched), so it is exact. We do not implement it here: it needs inline assembly padded to keep the backend's size estimate, which is beyond this tutorial. The point is that you can see all 7K cycles and where each one falls.
Below the DDR floor needs fewer bytes
After step B, R2–R6 are either compute-limited or balanced, and the overview's DDR floor (245.8K) becomes the relevant bound for further scheduling work: once every transfer overlaps compute, a better schedule cannot beat bytes / (19.08 B/cycle). Going lower means moving fewer bytes, for example lossless bit-packing of the weights in the blob, unpacked on the array.
6. The workflow
capture → overview → find the longest serialization or idle lane → zoom → hypothesize → change → re-trace the same window → verify byte-exact
- Capture at standard level (
run_kernel.py --timeline …, or--timelineonsdk source).- Overview: the whole run, plus the DDR floor (bytes read / peak B/cycle) and a per-region summary. Where is the kernel relative to its floor?
- Find the longest serialization (one engine waits for another with nothing overlapping: Blocked spans, DMA-only stretches) or the longest idle lane (DDR idle while later regions are DDR-heavy).
- Zoom into that window. Read bytes, B/cycle, flow ids and dependencies on the slices. Remember that ELS/PLS slices include queue and dependency time.
- Hypothesize what would move it: issue earlier, coalesce, chunk, move off an in-order queue, fewer bytes. Check the resources the move needs (OCM/LRM lifetime, queue depth, flow dependencies).
- Change one thing. Scheduling, placement and flow changes only; keep every consumer's dependency on the DMA that fills it (
setFlowId,barrier).- Re-trace the same window and compare panels. Did the gap close, or did the work just move somewhere else (step B's queue, step D's first trap)?
- Verify byte-exact on every SDK you care about. Step D's second trap passed on one SDK and failed on the other.
7. How to reproduce
Everything is in examples/timeline_tutorial/ in the SDK; its README.md lists the files. Besides that folder you need only docker, the ghcr.io/quadric-io/sdk-cli:26.09 and :26.08 images, and a license. The binary files are stored with Git LFS.
| file | purpose |
|---|---|
stock/mnv2_qcu.cpp | the stock kernel: the whole network as CGC emitted it |
stock/attention_stubs.hpp | the stub header CGC emits next to the kernel (included by kernel and host) |
stock/mnv2_qcu_host.cpp | the host program: reads input.bin and the weight blob, runs the kernel, writes output.bin |
stock/const_tensor_data.bin | the packed weight / constant blob (4,088,416 B) |
stock/input.bin, stock/golden_output.bin | one input tensor and the stock kernel's output for it |
run_kernel.py | stage, build and run one kernel with sdk source (QC-U: 16 MACs/PE, 8 MB OCM, 4 kB LRM, 32 GB/s read and write, 1.7 GHz, 256-bit DDR bus), check byte-exact, print cycles; --timeline OUT also captures a trace, --ddr-axi-width 128 is the bus-width diagnostic |
run.sh TAG CMD | run a command in ghcr.io/quadric-io/sdk-cli:TAG with the folder mounted at /tut and the current directory at /ws |
reproduce.sh WORKDIR | builds every step's kernel from stock/, verifies each under 26.09 (traced) and 26.08, runs the bus-width diagnostic, writes the diffs and renders all figures |
mk_step_a.py | step A: FC memCpys to region-2 start |
mk_prefetch_all.py | step B trap: 48 memCpys at region-2 start |
mk_prefetch_groups.py | step A+B: 4 group DMAs with setFlowId |
mk_step_c.py | step C: 7 input bands + gW0 |
mk_step_d.py SRC DST PLAN [--guard] | step D: FC chunks; PLAN says where they go ("2a:0 4a:1 6a:2 8a:3" = after the 2nd/4th/6th/8th all-to-all of R6; start:… = the trap) |
attention_stubs_step_d.hpp | stock attention_stubs.hpp + plsReadIssueNoWait |
patches/*.diff | the resulting diff of each step |
tlplot.py, make_figures.py | the annotated Gantt renderer and the figure definitions |
analyze.py TRACE | per-region lane-busy table, DDR bytes per region, largest compute-idle gaps |
cd examples/timeline_tutorial
./reproduce.sh /tmp/tltut
## stock 26.09 {"pass": true, "cycles": 484072, "exact": true} | 26.08 {"pass": true, "cycles": 512612, "exact": true}
## ...
## D 26.09 {"pass": true, "cycles": 417680, "exact": true} | 26.08 {"pass": true, "cycles": 460774, "exact": true}
The docker invocation always passes -u $(id -u):$(id -g) so the work tree stays yours, and mounts ~/.quadric (override with QUADRIC_HOME) for the license. Each step takes about a minute per SDK. Figures render in a few seconds from the traces with matplotlib on the host (python3 make_figures.py OUTDIR [figure ...], run from the work directory).
The stock sources come from mobilenet_v2_opt_asym_int8_q.onnx compiled with sdk graph compile --target QC-U --macs-per-pe 16 --ocm-size 8MB --ext-read-bw 32GBps --ext-write-bw 32GBps under sdk-cli 26.08, with the generated C++ split into the kernel (mnv2_qcu.cpp) and its host program (mnv2_qcu_host.cpp) so the kernel can be edited on its own. The step scripts match that exact source text and assert on it. For a kernel generated by another SDK version, or your own network, the ideas carry over, but apply the changes by hand (the diffs in patches/ show them).
