Unlock this feature

This feature isn’t part of your plan yet

Contact sales to get upgraded to the full DevStudio experience.

Unlock this feature

This feature isn't part of your plan yet.

What's New


NOTE: The Jupyter Notebook below is included in the Chimera SDK and can be run interactively by running the following CLI command:

$ quadric sdk notebook

From the Jupyter Notebook window in your browser, select the notebook named /quadric/sdk-cli/examples/doc_template/whats_new.ipynb.


Quadric Chimera SDK 26.09 – What's New

A Chimera at the harvest — the quarter's work gathered in.

26.09 is the quarterly MAJOR release. This rolls up the full quarter — everything since 26.06, spanning the 26.07, 26.08 and 26.09 monthly cuts. One through-line is more of the work staying where it is: post-processing that used to return to the host now runs on the array, K/V and layout staging that used to bounce through DDR now stays in L2, and attention blocks that used to arrive with stubs to hand-fill now arrive complete. The other opens new routes onto the hardware: llama.cpp ships inside the SDK image, built against the Chimera GPNPU backend, and PyTorch models now have a direct path through the stack — ResNet-18 runs end to end today, with more to follow in future releases. Mixture-of-Experts and Gated DeltaNet kernels arrived this quarter. Six new end-to-end demos shipped: DETR-R101 detection, SegFormer-B0 segmentation, WaveFormer sEMG gesture classification, NMS-free YOLO26, a Sarashina 2.2-3B LLM pipeline, and ModernBERT-base. Stay tuned for an MoE demo in 26.10! Multi-core ViT-B/16 at 1024×1024 now matches a single-core run bit for bit, inside an 8 MB budget. An ISS run can now be opened as a Perfetto timeline and read cycle by cycle.

Summary

  • Six New End-to-End Demos: DETR-R101 and SegFormer-B0 in 26.07, WaveFormer in 26.08, YOLO26, Sarashina 2.2-3B and ModernBERT-base in 26.09. Learn More
  • llama.cpp Ships Inside the SDK Image: Built against the Chimera GPNPU backend — llama-cli, llama-server, llama-bench and the device kernels — instead of a separate runtime image. Learn More
  • Attention Lowers Natively; No Stubs to Hand-Fill: DETR's decoder custom op is gone and all 17 attention cores lower through the compiler; a 12-block ViT went from 48 hand-filled attention stubs to none and an 8-block SegFormer from 16 to none. Learn More
  • ViT-B/16 at 1024×1024, Across Cores: the model compiles inside an 8 MB budget with activations split through external memory and LayerNorm streamed from DDR, and a multi-core spatial split now matches single-core bit-exactly — two-core output had drifted to 0.37 correlation across twelve blocks. Learn More
  • LLM Kernels: MoE, Gated DeltaNet, Quantized Decode: First Mixture-of-Experts and Gated DeltaNet kernels in 26.07, a hybrid attention + DeltaNet decoder block in 26.08, and in 26.09 a complete MoE block on real W4A16 weights plus quantized DeltaNet decode — 55% fewer DDR bytes per token, autoregressive decode 27.7–44.6% faster. Learn More
  • INT4 Weights on INT8 MACs: The DoubleINT8 W4A16 path runs INT4 weights on INT8 MACs at ~1.7× the FP16 path, and flat activations now tile across all 1024 cores instead of 32. Learn More
  • Quantization State Survives Export: Quantization state travels with the model instead of being reconstructed after a format conversion — operand scales are read off the quantized module, and opaque modules stay opaque through export. quadric_pyquant ships with the SDK and quantizes thirty-one tracked checkpoints — ResNet-18/50, VGG-16, MobileNetV2 and V3, EfficientNet-B0/B1/B2, ConvNeXt-tiny, Swin-T and ViT-B/16 for vision; Qwen3, Qwen2, Llama-3.2, Mistral-7B, Gemma2, GPT-OSS-20B and Sarashina-2.2-3B for language; Whisper and Moonshine for speech; plus Qwen3-VL, ModernBERT, pi0.5 and BEVFormer-tiny. Learn More
  • PyTorch Front End Support: quantize with quadric_pyquant, export with torch.export, compile. ResNet-18 is the first PyTorch model to run the whole path with no intermediate ONNX step; more PyTorch models will be supported in future releases. Learn More
  • quadric_pyquant v10: Quantizable K/V Caches: every attention site exposes its key and value caches as quantizable tensors, so an int8 or int4 KV cache is a config setting rather than a code change. Asymmetric weight grids widen the W4A16 recipes, and per-layer sensitivity ranking, a measured upgrade curve and a hardware cost model price a precision assignment in DDR bytes or cycles. Learn More
  • Custom Ops from a Plugin Manifest: a plugin.yaml and a kernel.hpp in a plugins directory describe a subgraph and the custom op that replaces it, and the swap happens before the model reaches CGC. Two bundles ship, with four tutorial notebooks. Learn More
  • Prefill Takes Runtime Extents: FP16 flash prefill, gate+up projection and qLinearMul take runtime extents, closing out a quarter of multicore and runtime-shape work. Learn More
  • ISS ~1.3–1.5× Faster; Three Power-Profiler Defects Fixed: ISS ~1.3–1.5× on ResNet-18 and Qwen, and three profiler defects that had a 500 MHz run reporting twice the power it drew. Learn More
  • A Perfetto Timeline for an ISS Run: sdk source --timeline and sdk graph compile --run --timeline turn an ISS run into a Chrome-Trace file you open in ui.perfetto.dev — per-engine memory movement, array activity, stalls, and the instruction each slice belongs to. Stage 1, schema 1.0, and the ISS gains no new flags. Learn More

Detailed Changes

Detailed Changes

Dev Studio

The quarter added six end-to-end demos, and finished pulling the vision and audio post-processing onto the array. 26.07 brought DETR-R101 detection and SegFormer-B0 segmentation; 26.08 added WaveFormer, an INT8 sEMG gesture classifier on QC-N; 26.09 adds YOLO26, whose NMS-free decode never returns to the host, Sarashina 2.2-3B, a Japanese LLM taken from Hugging Face checkpoint to one w4a16 Chimera kernel, and ModernBERT-base, a masked language model validated against ONNX Runtime.

New Notebooks

CategoryNotebook
DetectionYOLO26 N/M/L — end-to-end NMS-free detector on QC-U (26.09)
DetectionDETR-R101 — ResNet-101 backbone, 17 attention blocks (26.07)
SegmentationSegFormer-B0 — 512×512 (26.07)
Audio / wearablesWaveFormer — INT8 sEMG gesture classifier on QC-N (26.08)
Masked LMModernBERT-base — FP32 to validated INT8, then lowered and checked against ONNX Runtime (26.09)
Language modelingSarashina 2.2-3B — Japanese Llama-family 3B, Hugging Face checkpoint to one w4a16 Chimera kernel (26.09)
Custom opsPlugin tutorial — quickstart, consuming a bundle, the API guide, and pattern debugging. Ships in the repo under examples/custom_op/plugin_tutorial/, not on the docs site (26.09)
  • YOLO26, N/M/L. Backbone, PSA attention, neck, head and the NMS-free decode compile into one program, taking an image to 300 scored boxes. CGC compiles the convolutional graph; two CCL kernels cover the rest. One source pair serves all three variants, with every shape derived from the ONNX edges at compile time. yolo26Attention maps one token per PE and keeps the softmax PE-local in ~2.6 kB of the 4 kB LRM; yolo26Decode collapses the two-stage top-K into a single global top-300, so 900 per-level survivors merge and 300 boxes are decoded. On QC-U at 1.7 GHz, 1 core, 16 MB L2, 4 kB LRM, 16 MACs/PE, 128 GB/s: 2.14M cycles, 1.26 ms, 796 FPS for N; 10.3M cycles, 6.06 ms for L. Detection counts match the ORT int8 reference on every test image. External writes are 0.01 MB per inference, against 2.69 MB when the decode runs on the host.
  • llama.cpp in the SDK image. The Chimera GPNPU backend compiles against the SDK inside the image, shipping llama-cli, llama-completion, llama-server, llama-bench, llama-perplexity and the device kernels. The build fails if no kernels are produced. Earlier releases shipped a separate runtime image carrying binaries only; this one can be run and rebuilt against the release installed.
  • ModernBERT-base, in two notebooks. modernbert_quant_pipeline.ipynb takes answerdotai/ModernBERT-base from FP32 to a static INT8 ONNX model in its own venv, with no local checkpoint, dataset or API token: WikiText-2 serves as both the SmoothQuant calibration corpus and the masked-LM perplexity set, at 4.37 quantized against 4.06 FP32 over the same windows and masks. modernbert_gpnpu.ipynb lowers that output through ChimeraJob and checks the ISS against ONNX Runtime. On a single-core QC-N with 1 MB OCM the two agree on all 8 masked tokens of the validation paragraph, logits cosine 0.989 and rms/std 0.169, at 567.9 M cycles per 512-token sequence.
  • Sarashina 2.2-3B. A staged pipeline takes sbintuitions/sarashina2.2-3b-instruct-v0.1 from its Hugging Face card to one w4a16 Chimera kernel, runs it on the archsim for QC-N, QC-P or QC-U, or publishes a handoff directory for the HAPS-200 board. Each stage writes a STAGE.json recording its inputs, knobs and toolchain, and rebuilds only when that record stops matching. Decode cycles, token rate and rel_rmse_affine per target are in the notebook under examples/models/sarashina2.2-3b/, each row beside the toolchain it was measured on.
  • Hardware performance estimator in quadric_pyquant. Given a quantized model and a target HardwareProfile, it estimates per-layer roofline time as max(compute_time, memory_time) from the quantization config already on each layer. Profiles are declared in YAML: array geometry, native INT and float MAC support, DRAM bandwidth. It covers all five MAC-unit classes; QConv3d is unmodelled except in the non-overlapping case.

New Guides

GuideWhat it covers
llama.cpp on Chimera GPNPURunning and rebuilding llama-cli / llama-server / llama-bench against the GPNPU backend shipped in the image (26.09)
PyTorch model ingestionTaking a PyTorch checkpoint through quadric_pyquant and torch.export to a compiled program, and where the layer boundary sits (26.09)

Notebook Updates

  • DETR-R101 dropped its decoder custom op in 26.08 — all 17 attention cores lower through the compiler, and the model runs 25.9M → 18.5M cycles on QC-U.
  • pi0.5 gained a fused gate+up MLP in 26.08: 14.17M → 12.16M cycles at seq=968 on 8 cores, after 26.07 brought the pipeline to multicore at production sequence length.
  • WaveFormer moves onto the merged int8 attention kernel in 26.09. The kernel stages K/V in L2 and no longer takes KCache/VCache, so the adapter reduces to an argument reorder. 5,648,820 cycles on QC-N, about 18% fewer than the previous int8 P·V measurement, argmax match retained, and max |GPNPU − ORT| of 0.423 against a 6-LSB tolerance.
  • Whisper got faster and more accurate in 26.09. Attention DMAs move as head stripes instead of ~24k tiny transfers per layer, V is written directly into the VCache backing, and attention-adjacent edges travel as fp16 with no FX32 conversion passes. End to end 40.8 → 32.2 ms on QC-U, 1 core, 16 MB L2, 128 GB/s — 35.1 ms with the DMA work alone — with decoder PSNR 46.98 → 49.34 dB and the top-3 tokens exact against ORT. QC-P 4-core, the tutorial target, runs 25.07 ms at the same accuracy.
  • The MoE block is complete, on real W4A16 weights. The ChiPy example computed only the routed path; the real block is routed + sigmoid(x·W_sg) · shared, a dense shared-expert SwiGLU over every token scaled by a sigmoid output gate. Both now run, fused: one matmul over [Wg | Wu | Wsg | Wr] yields the shared gate/up, the output-gate logit and every router logit at once, and the accumulator is seeded with the gated shared output before the routed experts accumulate into it, so only the final hidden reaches DDR.
  • One notebook per Qwen model. Qwen3-8B had four notebooks and a separate LM-head component; decode and prefill now live in a single qwen3_8b/qwen3_8b.ipynb. The Qwen 1.5B GQA and decoder component tutorials are retired, and every GPNPU result in the set is checked against a host reference.
  • DeltaNet ops take the quantized signature: integer weight codes plus one symmetric scale per output channel for both projections, with FX32 activations rather than fp16 bit patterns.
  • Calibration observes an output range on every layer, fp32_mode included — the four ops that cannot derive a fixed-point format from their input (layer_norm, rms_norm, gelu, group_norm) previously had no observed span in that mode.

Core

CGC

  • Attention stubs are pre-provided. The two op sequences that had no compiler helper — a bare requantize after each of the Q, K and V projections, and the 1/sqrt(d_k) scaling after the QK matmul — are now implemented and registered in the attention replacer. A 12-block ViT went from 4 stubs per block to 0; an 8-block SegFormer from 2 per block to 0; a 17-block DETR was already covered. All three emit an attention_stubs.hpp with nothing to fill in.
  • The torch.export front end covers ResNet-18. Four activation Q-layers convert natively — QSigmoid, QReLU, QHardsigmoid, QHardswish — joining QGELU and QSiLU, with the ATen glue ResNet-18 needs: max_pool2d, adaptive_avg_pool2d, select.int, expand and cat. QAdd/QMul read operand scales off the quantized module instead of walking the graph back to reconstruct them; the walk-back and its two "missing upstream scale" failure branches are deleted, and an operand is requantized at its calibrated scale rather than passed through. Export applies no decompositions, so ops keep their granularity and opaque modules stay opaque — what reaches the compiler is the module the customer quantized, not a flattened approximation. A published-but-unusable output scale on a stats-only activation fails loudly rather than being silently dropped, and int64 compile-time constants narrow to int32 with a clear error when a value will not fit.
  • PyTorch model coverage. quadric_pyquant quantizes, and ChiPy exports, straight from PyTorch. The two paths have different reach, so they are stated separately:
    • Quantizes today — vision classification: ResNet-18/50, VGG-16, MobileNetV2, MobileNetV3-small/large, EfficientNet-B0/B1/B2, ConvNeXt-tiny, Swin-T, ViT-B/16. Language: Qwen3-0.6B/4B/8B, Qwen2-1.5B, Llama-3.2-1B/3B, Mistral-7B-v0.3, Gemma2-2B, GPT-OSS-20B, Sarashina-2.2-3B. Speech: Whisper tiny/small, Moonshine tiny/base. Vision-language: Qwen3-VL-2B/4B. Masked LM: ModernBERT-base. Robotics and automotive perception: pi0.5, BEVFormer-tiny.
    • Exports through torch.export to a compiled program today — ResNet-18. Additional models follow in future releases.
    • The boundary is the layer, not the model. A model whose quantized layers all have module converters exports; one leaning on a layer without a converter quantizes but does not yet export. The PyTorch ingestion guide carries the current converter tables.
  • Vision pipelines sped up in 26.08: SegFormer ~1.6× at 512×512 from tile-block layout throughout, and HRNet-family resize −36.9% (34.15M → 21.55M cycles) once native on-array bilinear extended from 2×-only to exact-4× INT8 half_pixel. The resize figure is measured with synthetic calibration.
  • ViT-B/16 at 1024×1024, on 8 MB and across cores. 26.08 got the model to compile inside an 8 MB budget: activations split through external memory, LayerNorm streamed from DDR, and io_ext liveness tracked in EXT rather than held in OCM. 26.09 makes it correct on more than one core. A multi-core spatial split declared its OCM output at the full anchor height while each core filled only its own band, so the spill memCpy over-copied core 0's unfilled rows into the next core's DDR region — ViT-1024 on two cores drifted from one, correlation 0.92 at a single block down to 0.37 across twelve. Banding the OCM output to the per-core ROI makes two cores bit-exact against one, and a producer whose consumer runs single-core is no longer split at all.
  • New operators across the quarter: ONNX GatherND as an opaque cgc::gatherND, ScatterElements as cgc::scatterElements, concat of graph inputs over every axis with multi-core support, cast-wrapped FP16 MatMul, and — from 26.07 — ONNX LSTM in forward, reverse and bidirectional forms. 26.08 added ScatterND, GatherElements, 2-D LayerNorm and quantized patch merge. 26.09 adds a quantized int8 embedding table — the first ONNX Gather over an int8 initializer to lower through the compiler.
  • Matmul takes a bias as a real third input. contrib.epu.qlinear_matmul gains bias as a dataflow argument rather than an attribute, mirroring the dense op: a matmul bias is per-PE, one value per output column, and layout adaptation, L2 allocation and liveness only see dataflow. A no-bias matmul carries an all-zero sentinel so the op always has three inputs, and both lowering paths — the kernel-library qGemm and the compiler's own accumulator bias-add — are covered.
  • Rank-only reshapes alias instead of flowing through L2, broadcast_to is handled as a shape transform in codegen, and float-constant concatenate inputs are converted at canonicalization.
  • MatMulNBits takes the exact int8 fold by default, and the per-K-tile recombine rounds.
  • Custom-op substitution is declarative. A plugin bundle — a plugin.yaml manifest plus a kernel.hpp — describes a subgraph and the custom op that replaces it, and the swap happens before the model reaches CGC. Bundles are discovered from the shipped library, a ./plugins folder, plugins_dir= or $QUADRIC_PLUGINS_PATH, with your own folder taking precedence and priority ordering within one; two bundles claiming the same nodes is an error naming both rather than a silent loss. Kernel template parameters and frac bits are derived per match from tranges and shapes instead of hardcoded in the manifest. job.plugin_result reports what fired, and plugins=False or QUADRIC_PLUGINS_OFF=1 turns it off — with nothing matched, the compile is byte-for-byte the one tvm's ChimeraJob would have produced. Two bundles ship (MatMulNBits dequant, and an ultraface bounding-box tail pattern); the entry point is the plugin's own ChimeraJob, so there is no second API to learn.

Quantization Library

26.09 ships quadric_pyquant v10. The theme is deciding where to spend precision rather than applying one setting everywhere: the library measures per-layer sensitivity on every workload, plots task accuracy against the number of layers upgraded, and prices the result against a hardware profile.

  • K/V cache quantization. Every attention site exposes its key and value caches as quantizable tensors, so an int8 or int4 cache is a config choice rather than a code change. The cache dominates memory traffic on long-context decode.
  • Asymmetric weight grids, and the W4A16 recipes that use them. Blockwise weights already existed; what is new is an asymmetric grid on the K-blocked linear path, on embedding tables and on batched-linear expert banks — all paths where the weight reaches the multiplier dequantized. Where a weight instead reaches an integer accumulator as raw codes it is still refused, with an error naming the paths that qualify. Ten new configs use it.
  • Per-layer sensitivity ranking on every workload. Previously language-modeling only, it now runs from the shared pipeline for all seven workloads, alongside a slower estimator that re-runs the workload's own evaluator once per candidate layer as the reference the fast one is judged against. Ranking against a teacher distribution is a new metric, added because a benchmark's own sampling noise exceeds the effect of upgrading a single layer and cannot order layers at all — the layer that is ranked and the metric that is plotted are now separate choices.
  • The upgrade curve, with measured endpoints. Task accuracy is re-measured as layers are upgraded in ranked order, starting from the deployed quantized floor and ending at an independently evaluated float ceiling — both measured, neither assumed. Random-order curves are drawn alongside, since the question is whether the ranking beats chance.
  • Hardware cost pricing. A precision assignment is priced in DDR bytes or cycles against the compiler's core descriptor, prefill and decode separately, reporting the largest ranked prefix a given budget affords and naming what it could not price rather than omitting it.
  • AWQ weight-grid fold. Reallocates a weight's block grid toward the channels its activation is largest on. SmoothQuant improves an activation grid; at W4A16 there is no activation grid left to improve, so AWQ is the fold that applies. The two are mutually exclusive.
  • Smaller levers: vision-tower fold strength (every text-side builder is handed the language submodel, so the existing control could not reach the image tower), a per-head key fold target expressed as an exponent or as the key range it should reach, a clip floor bounding how far the weight solve may shrink a grid it is re-fitting, and sequential weight solving promoted to the shared pipeline.
  • Visibility: a per-layer quantization report — captured shapes, datapath, per-operand grid, parameter and arithmetic counts — makes a layer left at full precision by a config typo visible. Evaluation gains per-document result rows for paired comparison and a count of how often a multiple-choice parser fell back to a random pick, without changing the reported score.
  • Gates and tests: a new vision-language W4A16 regression gate plus twelve new gate assertions, so a fold that silently no-ops now turns a gate red. Distinct test names go 1,526 → 1,925.
  • New model configs: Qwen3-8B (dense), Qwen3-VL-4B-Instruct, new Qwen3-VL-2B-Instruct W4A16 recipes at two block sizes, and a ViT-B/16 residual-add variant. Each ships a base config plus one pipeline config per op point — W8A8, W4A8, W4A16, and W4A16 with an int4 KV cache. Config count goes 117 → 134, with 50 pre-existing configs re-tuned.
  • Two behaviour changes land with it — the removed global quantile flag and the widened asymmetric grid. Both are in the Migration Guide.

CCL

Performance
  • Gated DeltaNet, from first kernels to a quantized decode path. 26.07 landed the first Gated DeltaNet and Mixture-of-Experts kernels; 26.08 assembled the hybrid decoder — three DeltaNet layers plus one gated-attention layer as a single A-A-A-B instance — and 26.09 quantizes the decode projections. Both now route through the weight-only quantized matmul: integer codes (int8, or eight 4-bit codes packed per 32-bit word along K) with one symmetric scale per output channel, activations unquantized, each B tile dequantized in LRM. Weight reads fall −57.3% (q8 272,224,832 → 116,132,416 bytes) for −55.4% of all DDR bytes per token, and the in_proj product no longer visits DDR at all. Production 1-step decode — what autoregressive serving runs — improves −27.7%/−27.3% on q8, −44.6%/−44.3% on q16 and −39.8%/−39.5% on q32.
  • Gated DeltaNet prefill no longer needs a layout pass. The depthwise conv writes q/k/v head-major itself, deleting the DDR scratch, the ~48 per-head memCpys that rebuilt that order, the intermediate transpose and the 1 MiB L2 conv_out buffer.
  • WaveFormer attention, int8-first. int8 P·V with a deferred normalize, a 64-entry int8 exp table replacing the iterative expFast, K and V fed to flash out of L2 instead of round-tripping the DDR caches, and RoPE rotating the int8 branch in place. On QC-N, q8/m8, single core, 0.5 GHz, 4 MB L2, 4 kB LRM, 8 GB/s the attention op ends at 423,566 cycles — the dominant levers being the exp LUT (−23.0%, iterative execution cycles 137,088 → 0) and L2-staged K/V (−18.5%).
  • INT4 weights on INT8 MACs. 26.08's DoubleINT8 W4A16 matmul runs at ~1.7× the FP16 path at prefill M=64, q16/m16, 4 cores — a kernel-level matmul result. 26.09 fixes the activation side: a flat [1,1,1,K] activation has one channel, so the channel-major iterator ran the conversion on 32 of 1024 cores. Deriving the iterator from the shape puts all 1024 to work: the matMulNBits suite falls 26,087,951 → 25,450,179 cycles (−2.44%) on q32/m8, 69 of 153 instances improving and none regressing beyond noise.
  • *Partial overloads for strided 5×5 and 7×7 convolutions, and topK batches its per-element stores rather than issuing one at a time.
  • A no-RAU path for align-corners bilinear upscale where each output axis is 2N or 2N − 1: each PE owns one input pixel, assembles its 3×3 neighbourhood and writes a packed 2×2 output block. The resize API and DDR staging are unchanged, and other geometries keep the existing dispatch.
  • Row-stationary bilinear upscale for large ratios, and the stride-2 depthwise input shuffle moves word pairs rather than single bytes — the pairs were never split, so it was always 8 halfword moves wearing the costume of 16 byte moves.
New Kernels
  • DeepSeek-V4 token-level cache compressors: CSA and HCA pool projected KV windows into single compressed entries — softmax-gated convex combination, RMSNorm, partial RoPE — with the normalize fused into the pool rather than materialized back to LRM.
  • MoE experts run on quantized weights. The expert-gather views a per-expert dequant scale row alongside the weight slabs and forwards it to both matmuls of the SwiGLU expert: routed experts int4, the shared expert int8, matching the W4A16 delivery. Float call sites are unchanged.
  • GEMM unifies onto matrixMulMatrix with a shared exact core, and matmul and flash attention gain FP16 outputs.
API and Shapes
  • Runtime shapes across the library. FP16 flash prefill accepts a query with fewer rows than the compile-time seqLength, so one build serves any prefill length up to that bound — the FP16 MAC is IEEE-754, where 0.0 × NaN = NaN, so reads past the live KV extent are bounded, not masked. qLinearMul, gateUpProjection and fusedGateUpProjection take a runtime row count, and matmul gains channel-batched operation with runtime tensor reshape — closing a quarter that began with 26.07's multicore flash-attention partitioning for MHA, GQA and MQA.
  • tensor::slice takes a real 2D sub-rectangle. A TensorView sourced its row pitch from its own extents, so a window could only span leading dimensions — whole channels, whole batches, leading rows — and NUM_COLS had to match the source. A view now sources its strides from the root it is a view into, so a window narrower than the row it steps over keeps its compile-time extents and stays usable where a RuntimeTensor is not.
  • tensor::reshape and tensor::slice return flow-tracked views that share the source tensor's tracking slot, so flows issued on a sub-view order against the parent and its siblings automatically. sliceAs is unchanged and not deprecated.

LLVM

  • MEU-shadow scheduling landed in 26.08 at 1.042× on YOLOX-tiny (12.69M → 12.18M cycles; QC-N, 8 MACs/PE, 1 core, 1.7 GHz), and lifted the convolutional fleet more broadly than the YOLO family that motivated it.
  • Dangling .LBB label in sideband loopback fixed. The back edge's target block, not the loop header, now supplies both the jump-setup operand and the loopback target; a header reachable only by fallthrough never defined the label the setup referenced, failing assembly with a stol: no conversion error.
  • Chain loss in load/store fusion for ALU ops fixed.
  • MEU setup hoisting restored in loops that contain no calls. Modelling the MEU config registers as call-clobbered fixed a real correctness bug, but routed that modelling through the register grouping, so the one register that legitimately varies per iteration blocked hoisting of the two that do not — even with no call in the loop. The call-clobber modelling is now decoupled from the grouping and fires only for calls that can actually reach MEU-triggering code.

ISS and Archsim Performance

  • A Perfetto timeline for a GPNPU run, from one flag. sdk source kernel.cpp --timeline run.json — or sdk graph compile model.onnx --run --timeline run.json --timeline-level standard — writes a Chrome-Trace file that opens in ui.perfetto.dev and reads cycle by cycle: what each memory engine moved, what the compute array was doing, where it stalled, and which instruction each slice belongs to. --timeline-level selects coarse, standard or full, and reduces in the ingestor rather than in the simulator. coarse is the default, and --timeline-max-bytes caps each core at 256 MB — truncation is recorded in trace_metadata, so a cut trace never reads as a complete run. Behind the flag a Python ingestor consumes the ISS's existing --eventStreamTarget stream over a private FIFO, so the simulator emits events and never a trace, and gains no new CLI options. The ingestor can also be driven directly for custom tooling — --fifo to attach to a running ISS, --replay to read a recorded stream. This is stage 1 and the first release of the tool to users, at schema_version 1.0.
  • The simulator is ~1.3–1.5× faster on ResNet-18 and Qwen as of 26.08, with no regressions across the suite and no change to numeric output. Three ISS variants also collapsed into one owning a vector of clusters.
  • Power profiling was wrong in three ways, and all three moved every reported value. The average_power divisor counted the cycle window one short — dividing by zero on a single-cycle window, and able to report an average larger than any value it averaged. Neither profiler accounted for the clock: both divided energy by cycle count, so a run at 500 MHz reported twice the power it drew. And the periodic profiler received its PE count where its reporting period belongs, so --powerProfileReportPeriod was ignored and predication credits were priced against the wrong array size.
  • A Python power-profiler ingestor now reads per-cycle events off the profiler stream, prices each activity against an energy map, and reports energy and average power by activity label and by profile region. The simulator serializes its run config onto the stream, so an energy map with no entry for the run's hardware configuration is an error rather than a silently substituted default.
  • simulatorThreads caps to the core count, and dcpclr is modelled.

Bug Fixes

  • 3×3 stride-2 convolution corrupted at VAP partition edges (26.08), and ONNX Clip with a single bound was silently converted to a passthrough, dropping the clamp.
  • Lane-packed dense weight buffers were under-sized in LRM, and a buffered flow's LRM allocation failure path raised NameError instead of reporting the allocation failure.
  • Saturated FX32 activations are clamped in the DoubleInt8 requantize for MatMulNBits int8 activations.
  • Attention qIn zero-fill is bounded by its own extent.
  • Concat: a buffered external write flow dropped the field's start offset, a negative concat axis was never normalized, a per-core L2 concat buffer was sized to one field's ROI and dropped every other field, and a batched flow's min-ROI carried the whole batch.
  • Two-pass FP16 flash masks Pass 1 against the runtime KV extent, and row-granularity packCols in numTiles() is corrected.
  • Past-edge ELS underflow in multicore tilewise and transpose paths fixed.
  • An expert-weight plane was selected with the wrong element type, rebuilt from the activation width instead of derived from the weight stack — under a packed stack that is off by exactly 4×, so one expert's slab landed on another's plane. The offset arithmetic now asserts that plane and stack share an element type and that the plane spans the stack.
  • Quantization scheduling: fan-out packing no longer loses its range, and fan-outs feeding norms are left unpacked.
  • External right-hand splitting defects resolved, and a rank > 4 constant input to a custom op is squeezed to 4-D at the frontend.

Breaking Changes & Migration Notes

Upgrading from 26.06 crosses three monthly cuts, so this is the full quarter's surface. Seven items require action; the remaining three are re-baselines.

  • Quantization library ships as quadric_pyquant with the SDK (26.08). Repoint imports and re-quantize any checkpoint calibrated with an older in-tree copy.
  • Requires transformers v5 (26.08) — transformers == 5.14.1 is an exact pin and pulls huggingface-hub 1.x with it. Everything else is a floor rather than a pin: torch >= 2.1, torchvision >= 0.16, Python 3.10 or later — though CI installs torch and torchvision at exact versions for the qq extras. Install the SDK's extras rather than pinning by hand.
  • Twenty-two kernel-library functions removed across the quarter — fourteen in 26.08, eight in 26.09 — plus a batch of verified-unused internal helpers in 26.07.
  • nn::qwen3PrefillGateUpProjection is gone, completing an arc that began with its 26.07 rename to nn::gateUpProjection and its 26.08 deprecation.
  • nn::qwen3Fp16Attention no longer carves its own KV cache — the caller owns it. The cache was taken from the per-invocation scratch, which cannot survive the call that allocated it. Existing callers compile and run unchanged, with no diagnostic, so audit call sites.
  • The patch-mesh nn::multiheadAttention SDK library APIs are deprecated, removal targeted for 26.11 — CGC's general attention supersedes them, and models compiled through CGC need no change.
  • runKernelsAsync is superseded by simulatorThreads (26.09), and copyBufferFromDevice, allocateAndCopyToDevice and unused DeviceManager entry points are deprecated. Move run configurations across; the helpers still compile and emit a diagnostic.
  • Quantization tails round ties toward positive infinity (26.08) — INT8 output on affected paths can differ by one LSB from 26.07.
  • ISS cycle counts shift versus 26.06, from 26.07's patch-mesh pause-tracking correction and 26.08's single-core-per-cluster coalescer change.
  • Reported power values change (26.09), and runKernelsAsync, copyBufferFromDevice and allocateAndCopyToDevice are deprecated.

Performance Summary

The charts below compare 26.09 with 26.08 for the monthly delta and with 26.06 for the quarter, matching the 26.06 page's own 26.06-vs-26.03 baseline. The per-benchmark figures beneath them come from merged release PRs across the quarter and are independent of the fleet run.

From release PR benchmarks across the quarter:

  • YOLO26 end to end (26.09): 2.14M cycles / 1.26 ms / 796 FPS (N) and 10.3M / 6.06 ms (L), QC-U 1.7 GHz, 1 core, 16 MB L2, 4 kB LRM, 16 MACs/PE, 128 GB/s. External writes 2.69 MB → 0.01 MB.
  • DETR-R101 end to end (26.08): 25.9M → 18.5M cycles (−28%), QC-U 8 MB L2 / 4 kB LRM / 16 MACs/PE / 1.56 GHz, single core. ≥80% box-agreement gate unchanged and clearing.
  • pi0.5 VLA (26.08): 14.17M → 12.16M cycles (−14.2%) at seq=968 on 8 cores.
  • Gated DeltaNet decode (26.09): DDR bytes per token −55.4%; production 1-step decode −27.7%/−27.3% (q8), −44.6%/−44.3% (q16), −39.8%/−39.5% (q32).
  • Whisper end to end (26.09): 40.8 → 32.2 ms on QC-U, 1 core, 16 MB L2, 128 GB/s; decoder PSNR 46.98 → 49.34 dB. QC-P 4-core 25.07 ms.
  • Sarashina 2.2-3B decode (26.09): 32,600,227 cycles / token, ~46 tokens/s on QC-U array 32, 128 GB/s; rel_rmse_affine 0.0448 against ONNX Runtime, first token matching.
  • WaveFormer attention (26.09): op at 423,566 cycles, QC-N q8/m8, single core, 0.5 GHz, 4 MB L2, 4 kB LRM, 8 GB/s. End to end 5,648,820 cycles, ~18% below the previous int8 P·V measurement.
  • SegFormer (26.08): ~1.6× at 512×512. HRNet-family resize (26.08): 34.15M → 21.55M cycles (−36.9%), synthetic calibration.
  • Simulator (26.08): ISS ~1.3–1.5× on ResNet-18 and Qwen, no change to numeric output.
  • W4A16 flat-activation tiling (26.09): matMulNBits suite 26,087,951 → 25,450,179 cycles (−2.44%), q32/m8.
  • MEU-shadow scheduling (26.08): 1.042× on YOLOX-tiny, 12.69M → 12.18M cycles; QC-N, 8 MACs/PE, 1 core, 1.7 GHz.

Several changes move baselines rather than deliver speedups: ISS cycle counts shift versus 26.06, reported power values change, and quantization tails round ties toward positive infinity. All are in the Migration Guide.

from IPython.display import Image, display
from whats_new_utils import compare_current_release_vs_last_release

month = compare_current_release_vs_last_release("26.09", "26.08")
display(Image(month[0]))

quarter = compare_current_release_vs_last_release("26.09", "26.06")
display(Image(quarter[0]))


Migration Guide

Upgrading from the previous major (26.06) crosses the 26.07, 26.08 and 26.09 cuts. Ordered by blast radius; the first four require action, the rest are re-baselines.

Quantization Library Ships as quadric_pyquant

What changed (26.08): the library is packaged with the SDK rather than living in the compiler tree, with its workload configs, READMEs and scripts in the wheel.

Before (≤26.07):

from tvm.contrib.epu.quadric_quant.layers import QConv2d, QMatMul

After (26.08 and later):

from quadric_pyquant.layers import QConv2d, QMatMul

Action required: repoint imports, and re-quantize any checkpoint calibrated with an older in-tree copy. ChiPy's converters are keyed by class object, so classes from a stale copy match nothing: no submodule is opaque, torch.export traces into every quantized layer, and the graph arrives as raw quantization arithmetic. Nothing fails at import — the symptom is an unconvertible graph. To check first, confirm your model's layer classes import from quadric_pyquant.layers.

Why: two copies meant the converter registry was built from one of them, so a model quantized with the other could not be consumed.

transformers v5 Required

What changed (26.08): transformers is pinned exactly at 5.14.1. Python 3.10 or later is a real floor, from python_requires. Everything else is a floor rather than a pin — torch >= 2.1, and torchvision >= 0.16 in the qq extra.

Action required: install the SDK's extras rather than pinning versions by hand — qq-core for the quantization engine on its own, qq to add torchvision and transformers. setup.cfg is the authoritative source for what each extra requires.

On torch 2.13: the packaging does not pin torch, but CI does — the one job that installs the qq extras installs torch==2.13.0 and torchvision==0.28.0 from the cu129 wheels, matching the venv the transformers 5.14.1 bump was validated in. torchcodec >= 0.13, < 0.14 in qq-speech is ABI-locked to that same torch minor. So 2.13 is the environment the qq path is built and tested in rather than a floor the ingestion path requires — that path is verified on torch 2.6.

Why: transformers v5 is required by the tracked vision-language and speech recipes, and it pulls huggingface-hub 1.x with it.

Kernel-Library Functions Removed Across the Quarter

What changed: twenty-two named functions were removed — fourteen in 26.08 (memcpy, trunc, vectorMag, loadTensors, argmaxImpl, argmax, vectorAngle, writeTensor among them) and eight in 26.09 (qwen3PrefillGateUpProjection, sin, cos, getCoreMask, blockUntilPaused, unloadAndCompareTensors, fullyConnectedTileBlock, and the ElsDependencyQueue struct) — plus a batch of verified-unused internal helpers in 26.07.

nn::qwen3PrefillGateUpProjection completes a three-release arc: renamed to nn::gateUpProjection in 26.07 behind a forwarding wrapper, deprecated with a diagnostic in 26.08, removed in 26.09.

Action required: callers of qwen3PrefillGateUpProjection move to fusedGateUpProjection, which additionally accepts a runtime row count as of this release. The remaining removals have no in-tree callers; a build error naming one indicates an out-of-tree kernel that must move to the supported entry point.

Why: deprecation cycles announced in 26.07 and 26.08 completing on schedule.

nn::qwen3Fp16Attention: the Caller Owns the KV Cache

What changed (26.09): qwen3Fp16Attention carved its K/V cache out of the ext-temp scratch it is handed. Scratch is reallocated per kernel invocation, so the cache could not outlive the call that allocated it — fatal for token-by-token decode, where token t+1 must attend over the keys token t wrote. The caller now supplies the cache via kv_cache_ptr.

Action required: allocate the K/V cache outside the attention call and pass it as kv_cache_ptr. This applies to every caller — qwen3Fp16Attention is autoregressive by construction (one token per call, enforced by static_assert), so there is no single-invocation case that is safe to leave alone. Prefill goes through the separate qwen3PrefillAttention entry point, and qwen3GatedFp16Attention takes the same argument.

Audit your call sites rather than relying on the build to find them: a caller that does not pass a cache still compiles and runs, with no compile error and no runtime diagnostic. The symptom is correct output at the first token and wrong output from the second onward.

Why: per-invocation scratch is the wrong lifetime for state that must persist across decode steps.

Patch-Mesh Multihead Attention Deprecated — Removal in 26.11

What changed (26.09): the five patch-mesh nn::multiheadAttention overloads — the NDArray-level engine and its four L2-level wrappers — are marked deprecated (2026-08-27). This affects the SDK kernel-library APIs only. Patch-mesh attention is now compiled automatically as cgc::generalMultiheadAttention, which supersedes the hand-written entry points, and the attention layer itself has moved from the kernel library into the compiler.

Action required: for callers of the SDK kernel library, stop calling the hand-written entry points and let the compiler emit the attention. Models compiled through CGC need no change. They still compile and emit a deprecation diagnostic. Removal is targeted for 26.11, leaving two releases of overlap.

Why: two implementations of the same attention, one of which the compiler can generate and keep in step with the mesh.

Quantization Tails Round Ties Toward +inf

What changed (26.08): quantization tails round ties toward positive infinity.

Action required: none functionally, but re-baseline stored outputs — INT8 output on affected paths can differ by one LSB from 26.07. Bit-exact comparisons against pre-26.08 goldens will flag.

Why: one rounding convention across the quantization paths, matching the kernel library's math::round.

ISS Cycle Counts Shift Versus 26.06

What changed: two independent corrections. 26.07's patch-mesh pause-tracking fix corrects previously over-counted cycles on pause-heavy kernels, and 26.08's single-core-per-cluster devices no longer register a coalescer channel. Multi-core clusters are unaffected by the second.

Action required: re-baseline cycle-count goldens taken on 26.06 or earlier. The counts are more accurate, not faster; no kernel changed.

Why: both were counting artefacts rather than behaviour.

Reported Power Values Change

What changed (26.09): three defects in the ISS power profiler, each of which moved every value it reported. The average_power divisor counted the cycle window one short; energy was divided by cycle count rather than by time, so power was only correct at 1 GHz; and the periodic profiler received the PE count in place of its reporting period.

Action required: re-baseline any stored power figure. A 26.09 power report is not comparable with one captured earlier — most visibly on sub-GHz targets, where a 500 MHz run previously reported twice the power it drew. Energy maps must carry an entry for the run's hardware configuration; a missing entry is now an error rather than a silently substituted default.

Why: an average larger than every value averaged, and a clock-independent divisor, are both wrong in a way no downstream consumer could detect.

Simulator and Host-Backend Entry Points Deprecated

What changed (26.09): runKernelsAsync is superseded by simulatorThreads, which the ISS host bridge now caps to the run's total core count. copyBufferFromDevice and allocateAndCopyToDevice are deprecated, along with unused DeviceManager entry points in host_backend.hpp.

Before (≤26.08):

run_config.runKernelsAsync = True

After (26.09):

run_config.simulatorThreads = <threads>   # capped to ipCoreCount

Action required: move run configurations onto simulatorThreads; values above the core count are capped rather than rejected. The host-backend helpers still compile and emit a diagnostic — plan a move before a future removal, which has not yet been dated.

Why: one parameter expressing how much host parallelism the simulation gets, instead of a boolean plus an implicit thread count.

Sign in to your account

Don't have an account? 
By signing in, you are agreeing to our Terms of Use and Privacy Policy.
Quadric // One architecture. Every algorithm.

Develop.

Simulate.

Profile.

Collaborate.