NOTE: The Jupyter Notebook below is included in the Chimera SDK and can be run interactively by running the following CLI command:
$ quadric sdk notebook
From the Jupyter Notebook window in your browser, select the notebook named /quadric/sdk-cli/examples/models/whisper/whisper_tutorial.ipynb.
Whisper-Tiny Speech Recognition on Chimera GPNPU
This notebook runs OpenAI's Whisper-Tiny on the Quadric Chimera GPNPU, from a raw waveform to decoded token ids.
The pre- and postprocessing that normally runs on the host — the log-mel feature extractor and the greedy argmax — are CCL custom ops stitched to the ONNX graphs with ChiPy, so the host only decodes the audio file.
Pipeline
| Stage | Runs on | Input | Output |
|---|---|---|---|
| FLAC decode | host CPU | LibriSpeech clip | 16 kHz fp32 waveform |
Log-mel (whisper::logMel) | GPNPU | [1, 480000] waveform | [1, 80, 3000] log-mel |
Conv front-end (whisper::convStem) | GPNPU | log-mel | [1, 1500, 384] features |
| Encoder, 4 layers | GPNPU | features | [1, 1500, 384] hidden states |
| Decoder prefill, 4 layers | GPNPU | hidden states + 4 prompt ids | [1, 4, 51865] logits |
Greedy argmax (whisper::greedyTokens) | GPNPU | logits | [1, 4] token ids |
1. Setup
Install the Python dependencies and download the Whisper-Tiny encoder and decoder ONNX models from the Quadric model S3 bucket.
%pip install -q -r whisper_requirements.txt
import glob
import urllib.request
from pathlib import Path
import matplotlib.pyplot as plt
import numpy as np
import onnxruntime as ort
from tvm.contrib.epu import chipy
from tvm.contrib.epu.chimera_job import core as chimera_core
from tvm.contrib.epu.chimera_job.hw_config import HWConfig
from whisper_helper.encoder_helper import replace_encoder_custom_ops
from whisper_helper.decoder_helper import replace_decoder_custom_ops
from whisper_helper.generate_tranges import generate_tranges
from whisper_helper.prepost import (
DECODER_COMPILE_TRANGES,
KERNEL_HPP,
N_WAVE,
STEM_D,
STEM_T2,
build_decoder_pipeline,
build_encoder_pipeline,
build_mel_constants,
build_stem_constants,
compile_context,
encoder_compile_tranges,
ensure_decoder_models,
numpy_float_mel,
report,
)
[33mWARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv[0m[33m
[0mNote: you may need to restart the kernel to use updated packages.
S3_BASE = "https://sdk-cli-models.s3.amazonaws.com"
S3_MODELS = {
"encoder_model.onnx": f"{S3_BASE}/whisper_encoder_model.onnx",
"decoder_prefill.onnx": f"{S3_BASE}/whisper_decoder_prefill.onnx",
}
def download_models(model_dir: Path, models: dict = S3_MODELS) -> None:
"""Download model files from S3 if not already present locally.
Parameters
----------
model_dir : Path
Local directory where model files are stored.
models : dict, optional
Mapping of ``{local_filename: s3_url}``. Defaults to ``S3_MODELS``.
"""
for local_name, url in models.items():
local_path = model_dir / local_name
if local_path.exists():
print(f" {local_name}: exists ({local_path.stat().st_size:,} bytes)")
continue
print(f" Downloading {local_name}...")
urllib.request.urlretrieve(url, str(local_path))
print(f" {local_name}: {local_path.stat().st_size:,} bytes")
MODEL_DIR = Path(".")
print("Checking / downloading model files...")
download_models(MODEL_DIR)
ENCODER_ONNX = "encoder_model.onnx"
ENCODER_TRANGES = "encoder_model.tranges"
ENCODER_GPNPU = "encoder_model_gpnpu.onnx"
DECODER_ONNX = "decoder_prefill.onnx"
DECODER_TRANGES = "decoder_prefill.tranges"
DECODER_MATCHED = "decoder_model_matched.onnx"
ENCODER_ONNX_PATH = MODEL_DIR / ENCODER_ONNX
DECODER_ONNX_PATH = MODEL_DIR / DECODER_ONNX
## One hardware configuration for both programs.
hw_config = HWConfig(
product="QC-P",
ocm_size="32MB",
macs_per_pe=16,
num_cores=4,
ext_rd_bw="128GBps",
ext_wr_bw="128GBps",
)
Checking / downloading model files...
Downloading encoder_model.onnx...
encoder_model.onnx: 8,331,061 bytes
Downloading decoder_prefill.onnx...
decoder_prefill.onnx: 95,466,860 bytes
2. Audio & Calibration
Whisper consumes 16 kHz mono audio; the GPNPU takes the padded waveform directly and produces the mel itself.
whisper_helper/data_prep.py loads the LibriSpeech "dummy" validation split and runs each sample through the floating-point encoder to produce activations for the calibration step.
from whisper_helper.data_prep import prepare_whisper_data
SAMPLE_INDEX = 0
NUM_CALIBRATION_SAMPLES = 73
data = prepare_whisper_data(
encoder_onnx=ENCODER_ONNX_PATH,
decoder_onnx=DECODER_ONNX_PATH,
num_calibration_samples=NUM_CALIBRATION_SAMPLES,
sample_index=SAMPLE_INDEX,
)
encoder_session = data.encoder_session
decoder_session = data.decoder_session
calibration_mels = data.calibration_mels
calibration_encoder_outputs = data.calibration_encoder_outputs
## The GPNPU program starts at the waveform: pad the clip to Whisper's 30 s.
waveform = np.zeros(N_WAVE, dtype=np.float32)
waveform[: min(len(data.audio), N_WAVE)] = data.audio[:N_WAVE]
print(f"waveform: {len(data.audio) / 16000:.1f}s of audio -> {waveform.shape}")
print(f"reference: {data.reference_text}")
preprocessor_config.json: 0.00B [00:00, ?B/s]
tokenizer_config.json: 0.00B [00:00, ?B/s]
vocab.json: 0.00B [00:00, ?B/s]
tokenizer.json: 0.00B [00:00, ?B/s]
merges.txt: 0.00B [00:00, ?B/s]
normalizer.json: 0.00B [00:00, ?B/s]
added_tokens.json: 0.00B [00:00, ?B/s]
special_tokens_map.json: 0.00B [00:00, ?B/s]
Primary: 1272-128104-0000.flac | ref: MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME HIS GOSPEL
Calibration: 73 mels, 73 encoder outputs
waveform: 5.9s of audio -> (480000,)
reference: MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME HIS GOSPEL
3. Tensor-Range Calibration
Tensor ranges (tranges) record per-tensor min/max statistics from the calibration data. The compiler uses these to assign fixed-point fractional bits.
Encoder ranges come from the mel spectrograms, decoder ranges from the encoder's floating-point outputs.
print(f"Generating encoder tranges ({len(calibration_mels)} samples)...")
generate_tranges(
str(ENCODER_ONNX_PATH),
ENCODER_TRANGES,
num_samples=len(calibration_mels),
mel_spectrograms=calibration_mels,
)
print(f"Generated {ENCODER_TRANGES}")
print(f"\nGenerating decoder tranges ({len(calibration_encoder_outputs)} samples)...")
generate_tranges(
str(DECODER_ONNX_PATH),
DECODER_TRANGES,
num_samples=len(calibration_encoder_outputs),
encoder_outputs_list=calibration_encoder_outputs,
)
print(f"Generated {DECODER_TRANGES}")
Generating encoder tranges (73 samples)...
Processing encoder_model.onnx...
Inputs: [('input_features', [1, 80, 3000])]
Using 73 real audio samples
Running calibration with 73 real audio samples...
Saved 224 tensor ranges to encoder_model.tranges
Generated encoder_model.tranges
Generating decoder tranges (73 samples)...
Processing decoder_prefill.onnx...
Inputs: [('input_ids', [1, 4]), ('encoder_hidden_states', [1, 1500, 384])]
Using 73 real encoder outputs
Running calibration with 73 random samples...
Saved 345 tensor ranges to decoder_prefill.tranges
Generated decoder_prefill.tranges
4. Custom Op Replacement
Both graphs get the same treatment: every MatMulNBits subgraph becomes a linalg::matMulNBits custom op and every attention core becomes nn::whisperAttention. The attention-adjacent activation edges are then declared float16 so the FP16 kernels exchange them directly.
4.1 Encoder
The encoder adds one more pass: its convolutional front-end (Conv1 → GELU → Conv2 → GELU → Transpose → PositionalEmbedding, 15 nodes) collapses into a single whisper::convStem node. The result is one graph whose input is still input_features — the mel arrives from logMel.
encoder_graph_info = replace_encoder_custom_ops(
encoder_onnx=ENCODER_ONNX,
encoder_tranges=ENCODER_TRANGES,
encoder_path=ENCODER_GPNPU,
conv_stem_consts=build_stem_constants(),
)
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_392__v_393__v_394
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_391
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_350__v_351__v_352
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_349
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_315__v_316__v_317
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_314
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_310__v_312__v_318
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_309
L0 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
L0 attn core
L0 outproj → linalg::matMulNBits<31, 31, 28>
L0 fc1 → linalg::matMulNBits<31, 31, 29>
L0 fc2 → linalg::matMulNBits<31, 31, 29>
Layer 0: all custom ops replaced
L1 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
L1 attn core
L1 outproj → linalg::matMulNBits<31, 31, 30>
L1 fc1 → linalg::matMulNBits<31, 31, 28>
L1 fc2 → linalg::matMulNBits<31, 31, 28>
Layer 1: all custom ops replaced
L2 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
L2 attn core
L2 outproj → linalg::matMulNBits<31, 31, 28>
L2 fc1 → linalg::matMulNBits<31, 31, 27>
L2 fc2 → linalg::matMulNBits<31, 31, 28>
Layer 2: all custom ops replaced
L3 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
L3 attn core
L3 outproj → linalg::matMulNBits<31, 31, 29>
L3 fc1 → linalg::matMulNBits<31, 31, 28>
L3 fc2 → linalg::matMulNBits<31, 31, 27>
Layer 3: all custom ops replaced
Enabled fp16 attention edges (8 edges)
Saved encoder: 130 nodes, input=input_features → encoder_model_gpnpu.onnx
4.2 Decoder
The decoder prefill processes the four Whisper bootstrap tokens (<|startoftranscript|>, <|en|>, <|transcribe|>, <|notimestamps|>) with cross-attention to the encoder's hidden states, so the replacement also covers the cross-attention variant of nn::whisperAttention.
ensure_decoder_models then derives the two variants the pipeline needs: a logits-only float model for CPU mode, and the matched model with int32 input_ids.
replace_decoder_custom_ops(
decoder_onnx=DECODER_ONNX,
decoder_tranges=DECODER_TRANGES,
decoder_matched_path=DECODER_MATCHED,
)
decoder_float, decoder_gpnpu = ensure_decoder_models()
print(f"CPU-mode model: {decoder_float.name}\nGPNPU model: {decoder_gpnpu.name}")
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_612__v_614__v_615
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_489__v_490__v_491
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_480__v_481__v_482
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_471__v_472__v_473
/decoder/layers.0/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.0/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.1/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.1/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.2/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.2/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.3/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.3/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
fused_qkv/_v_611 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.0/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.0/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.0/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.0/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.0/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 30>
fused_qkv/_v_488 → linalg::matMulNBits<31, 31, 27>
/decoder/layers.1/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
/decoder/layers.1/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
/decoder/layers.1/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 30>
/decoder/layers.1/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
/decoder/layers.1/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
fused_qkv/_v_479 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.2/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
/decoder/layers.2/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.2/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 30>
/decoder/layers.2/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.2/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
fused_qkv/_v_470 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.3/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
/decoder/layers.3/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.3/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
/decoder/layers.3/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
/decoder/layers.3/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
lm_head → linalg::matMulNBits<31, 31, 31>
L0 self-attn FP16 MHA core
L0 cross-attn FP16 MHA core
L1 self-attn FP16 MHA core
L1 cross-attn FP16 MHA core
L2 self-attn FP16 MHA core
L2 cross-attn FP16 MHA core
L3 self-attn FP16 MHA core
L3 cross-attn FP16 MHA core
CPU-mode model: _decoder_prefill_logits.onnx
GPNPU model: _decoder_matched_ids32.onnx
5. Pipeline Setup and Build
ChiPy stitches each graph to its custom ops as a single traced program: the encoder takes the log-mel op on the input side, the decoder takes the greedy argmax on the output side.
Calling a traced function directly runs it in CPU mode, where each custom op falls back to its Python reference. That checks the wiring in seconds, before spending minutes on a compile.
5.1 Encoder
ref_mel, fb_float = numpy_float_mel(data.audio, "reflect")
mel_consts = build_mel_constants(fb_float)
ort_hidden = encoder_session.run(None, {"input_features": ref_mel[None].astype(np.float32)})[0]
encoder_pipeline = build_encoder_pipeline(mel_consts)
cpu_out = encoder_pipeline(
chipy.Tensor(shape=(1, N_WAVE), dtype="float32", values=waveform),
ENCODER_ONNX,
)
cpu_psnr = report("chipy CPU vs ORT hidden", ort_hidden, np.asarray(cpu_out.values))
## CPU mode runs the same graph under ONNX Runtime, so agreement is numerical
## noise, not quantization: a low PSNR here means the wiring is wrong.
CPU_PSNR_MIN_DB = 100.0
assert cpu_psnr >= CPU_PSNR_MIN_DB, f"CPU-mode wiring regression: {cpu_psnr:.2f} dB"
chipy CPU vs ORT hidden max|e|=0.00000 rms=0.000000 PSNR=630.91 dB
5.2 Decoder
CPU mode is checked against the floating-point hidden states; section 6 then feeds the decoder the encoder's device output.
decoder_input_ids = np.array([[50258, 50259, 50359, 50363]], dtype=np.int32)
decoder_ort = ort.InferenceSession(str(decoder_float), providers=["CPUExecutionProvider"])
def ort_tokens_for(hidden):
"""Greedy token ids from the float decoder, for a given hidden state."""
logits = decoder_ort.run(
None, {"input_ids": decoder_input_ids, "encoder_hidden_states": hidden}
)[0]
return logits[0].argmax(-1)
decoder_pipeline = build_decoder_pipeline()
cpu_out = decoder_pipeline(
chipy.Tensor(shape=(1, 4), dtype="int32", values=decoder_input_ids),
chipy.Tensor(shape=(1, STEM_T2, STEM_D), dtype="float32", values=ort_hidden),
str(decoder_float),
)
cpu_tokens = np.asarray(cpu_out.values).ravel().astype(np.int64)
host_tokens = ort_tokens_for(ort_hidden)
print(f"chipy CPU tokens: {cpu_tokens} match={np.array_equal(cpu_tokens, host_tokens)}")
assert np.array_equal(
cpu_tokens, host_tokens
), f"CPU-mode token mismatch: {cpu_tokens} != {host_tokens}"
chipy CPU tokens: [50259 50359 50363 2221] match=True
6. Compile, Run, and Validation
pipeline.compile() hands a whole program to the Chimera Graph Compiler: the custom-op kernels and the ONNX graph become one module under ccl_build/. .run() then executes it on the Instruction Set Simulator.
Both programs compile for the same target, so both calls pass the same hw_config.
| Check | Observed | Asserted gate |
|---|---|---|
| Encoder hidden states vs ONNX Runtime | ~54.5 dB PSNR | >= 20 dB and MSE <= 5 |
| Token ids vs ONNX Runtime, from device hidden states | exact match | exact |
| Token ids vs the all-host chain | exact match | exact |
The encoder gate is deliberately far below the observed value: it catches a broken pipeline, not a small numerical drift.
6.1 Encoder — waveform → hidden states
with compile_context(qlut=False):
encoder_compiled = encoder_pipeline.compile(
hw_config=hw_config,
module_name="whisper_e2e_chipy",
custom_op_header=str(KERNEL_HPP),
value_proto_ranges=encoder_compile_tranges(),
waveform=chipy.Tensor(shape=(1, N_WAVE), dtype="float32"),
encoder=ENCODER_GPNPU,
)
encoder_out = encoder_compiled.run(waveform=waveform)
iss_hidden = (
np.asarray(next(iter(encoder_out.values())).tensor).reshape(ort_hidden.shape).astype(np.float32)
)
psnr = report("ISS vs ORT hidden", ort_hidden, iss_hidden)
mse = float(np.mean((ort_hidden - iss_hidden) ** 2))
ENCODER_PSNR_MIN_DB = 20.0
assert (
psnr >= ENCODER_PSNR_MIN_DB
), f"Encoder PSNR regression: {psnr:.2f} dB < {ENCODER_PSNR_MIN_DB} dB"
## MSE catches blow-ups that PSNR can mask when the signal range is wider than
## expected (the encoder output's range grows with calibration drift, which
## can keep PSNR steady while errors climb in absolute terms).
ENCODER_MSE_MAX = 5 # ~5x headroom over current MSE; PSNR 20 dB equivalent
assert mse <= ENCODER_MSE_MAX, f"Encoder MSE regression: {mse:.4f} > {ENCODER_MSE_MAX}"
/quadric/sdk-cli/examples/models/whisper/whisper_helper/prepost.py:57: RuntimeWarning: invalid value encountered in cast
wave[: len(audio)] = audio[:N_WAVE]
2026-09-22 02:44 - INFO - epu - codegen - START==================================build_relay
2026-09-22 02:44 - INFO - epu - codegen - START===============================optimize_relay
2026-09-22 02:44 - INFO - epu - codegen - START====================quantize_to_cpu_runnable_fx
2026-09-22 02:44 - INFO - epu - fx -
Source name Op Output 0 Range Output 0 Frac Bits
------------------------------------------ ----------------------------- -------------------------- --------------------
contrib.epu.quadric_custom_op (-32.0, 31.99999998509884) 26
CustomOp/whisper::convStem0 contrib.epu.quadric_custom_op (-8.0, 7.99999999627471) 28
/layers.0/self_attn_layer_norm/Add_1 nn.layer_norm [-16.5387f, 11.5336f] 26
CustomOp/linalg::matMulNBits<31, 31, 28>2 contrib.epu.quadric_custom_op [-4.96675f, 2.94481f] 28
/layers.0/Add add [-2.08293f, 3.24672f] 27
/layers.0/final_layer_norm/Add_1 nn.layer_norm [-14.6758f, 21.1295f] 26
CustomOp/linalg::matMulNBits<31, 31, 29>3 contrib.epu.quadric_custom_op [-10.8754f, 12.0852f] 27
/layers.0/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 12.0852f] 27
CustomOp/linalg::matMulNBits<31, 31, 29>4 contrib.epu.quadric_custom_op [-4.11117f, 4.20129f] 28
/layers.0/Add_1 add [-4.7491f, 3.81342f] 27
/layers.1/self_attn_layer_norm/Add_1 nn.layer_norm [-10.1508f, 14.0432f] 27
CustomOp/linalg::matMulNBits<31, 31, 30>7 contrib.epu.quadric_custom_op [-1.57465f, 2.00129f] 29
/layers.1/Add add [-4.0395f, 3.76171f] 27
/layers.1/final_layer_norm/Add_1 nn.layer_norm [-13.3625f, 16.9489f] 26
CustomOp/linalg::matMulNBits<31, 31, 28>8 contrib.epu.quadric_custom_op [-10.0635f, 13.8182f] 27
/layers.1/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 13.8182f] 27
CustomOp/linalg::matMulNBits<31, 31, 28>9 contrib.epu.quadric_custom_op [-6.80559f, 7.79773f] 28
/layers.1/Add_1 add [-10.0096f, 9.14031f] 27
/layers.2/self_attn_layer_norm/Add_1 nn.layer_norm [-21.9773f, 22.2916f] 26
CustomOp/linalg::matMulNBits<31, 31, 28>12 contrib.epu.quadric_custom_op [-4.5468f, 3.53895f] 28
/layers.2/Add add [-9.32271f, 9.04946f] 27
/layers.2/final_layer_norm/Add_1 nn.layer_norm [-27.1201f, 24.2017f] 26
CustomOp/linalg::matMulNBits<31, 31, 27>13 contrib.epu.quadric_custom_op [-40.5755f, 128.95f] 23
/layers.2/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 128.95f] 23
CustomOp/linalg::matMulNBits<31, 31, 28>14 contrib.epu.quadric_custom_op [-27.8214f, 248.118f] 23
/layers.2/Add_1 add [-30.3549f, 253.457f] 22
/layers.3/self_attn_layer_norm/Add_1 nn.layer_norm [-21.1582f, 33.7291f] 25
CustomOp/linalg::matMulNBits<31, 31, 29>17 contrib.epu.quadric_custom_op [-5.37736f, 11.7757f] 27
/layers.3/Add add [-29.5622f, 264.989f] 22
/layers.3/final_layer_norm/Add_1 nn.layer_norm [-94.0882f, 120.342f] 24
CustomOp/linalg::matMulNBits<31, 31, 28>18 contrib.epu.quadric_custom_op [-119.89f, 329.025f] 22
/layers.3/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 329.025f] 22
CustomOp/linalg::matMulNBits<31, 31, 27>19 contrib.epu.quadric_custom_op [-121.728f, 52.5513f] 24
/layers.3/Add_1 add [-131.772f, 276.199f] 22
/layer_norm/Add_1 nn.layer_norm [-17.5736f, 19.0347f] -
2026-09-22 02:44 - INFO - epu - codegen - START====================build_cpu_runnable_fx_relay
2026-09-22 02:44 - INFO - epu - codegen - START=======================quantize_to_chimera_fx
2026-09-22 02:44 - INFO - epu - codegen - START=================================relay_to_tir
2026-09-22 02:44 - INFO - epu - codegen - START===========================relay_to_epu_relay
2026-09-22 02:44 - INFO - epu - codegen - START==============================adapt_and_order
2026-09-22 02:44 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 02:44 - INFO - epu - codegen - START=============================plan_lrm_virtual
2026-09-22 02:44 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 02:45 - INFO - epu - codegen - START===============================lrm_alloc_loop
2026-09-22 02:45 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 02:45 - INFO - epu - codegen - START================================lrm_splitting
2026-09-22 02:46 - INFO - epu - codegen - START==============================ext_split_relay
2026-09-22 02:46 - INFO - epu - codegen - START====================================build_tir
2026-09-22 02:48 - INFO - epu - iss_testing - Found tranges for input: <tvm.contrib.epu.interval.Interval object at 0x7d843e6ccf10>
2026-09-22 02:48 - INFO - epu - iss_testing - Started Running Graph on Chimera ISS...
FILM 46/46: 100%|███████████████████████████████████████████████████| 46/46 [33:46<00:00, 44.06s/it]
2026-09-22 03:27 - INFO - epu - iss_testing - Done 0:38:58.354943
ISS vs ORT hidden max|e|=10.82557 rms=0.065826 PSNR= 54.55 dB
6.2 Decoder — hidden states → token ids
The hidden states come from the encoder's ISS run in 6.1, so this is an end-to-end device check. The tokens are compared twice: against ORT on those same device hidden states, which isolates the decoder, and against the all-host chain from 5.2, which catches encoder drift that flips a token.
with compile_context(qlut=True):
decoder_compiled = decoder_pipeline.compile(
hw_config=hw_config,
module_name="whisper_dec_chipy",
custom_op_header=str(KERNEL_HPP),
value_proto_ranges=DECODER_COMPILE_TRANGES,
input_ids=chipy.Tensor(shape=(1, 4), dtype="int32"),
encoder_hidden=chipy.Tensor(shape=(1, STEM_T2, STEM_D), dtype="float32"),
decoder=str(decoder_gpnpu),
)
decoder_out = decoder_compiled.run(input_ids=decoder_input_ids, encoder_hidden=iss_hidden)
iss_tokens = np.asarray(next(iter(decoder_out.values())).tensor).ravel().astype(np.int64)
ort_tokens = ort_tokens_for(iss_hidden)
tokens_match = np.array_equal(iss_tokens, ort_tokens)
print(f" GPNPU tokens: {iss_tokens.tolist()}")
print(f" ORT tokens: {ort_tokens.tolist()}")
print(f" Token ids match ORT: {tokens_match}")
## Same hidden states on both sides: isolates the decoder's own error.
assert tokens_match, f"Decoder token mismatch: {iss_tokens} != {ort_tokens}"
print(f"\n Decoded: {data.processor.tokenizer.decode(iss_tokens)}")
## And against the all-host chain, so encoder drift that flips a token fails here.
assert np.array_equal(
iss_tokens, host_tokens
), f"Full-chain token mismatch: device {iss_tokens} != all-host {host_tokens}"
2026-09-22 03:27 - INFO - epu - codegen - START==================================build_relay
2026-09-22 03:27 - INFO - epu - codegen - START===============================optimize_relay
2026-09-22 03:27 - INFO - epu - codegen - START====================quantize_to_cpu_runnable_fx
2026-09-22 03:27 - INFO - epu - fx -
Source name Op Output 0 Range Output 0 Frac Bits
------------------------------------------------- ----------------------------- -------------------------- --------------------
/decoder/embed_tokens/Gather contrib.epu.embedding [-0.154297f, 0.20166f] 31
/decoder/Add add [-0.586182f, 0.470642f] 31
/decoder/layers.0/self_attn_layer_norm/Add_1 nn.layer_norm [-5.01829f, 4.52495f] 28
CustomOp/linalg::matMulNBits<31, 31, 28>8 contrib.epu.quadric_custom_op [-7.66943f, 8.27715f] 27
CustomOp/nn::whisperAttention<384, 6, 4, 14>33 contrib.epu.quadric_custom_op [-0.863014f, 1.08392f] 30
CustomOp/linalg::matMulNBits<31, 31, 31>9 contrib.epu.quadric_custom_op [-0.691889f, 0.544409f] 31
/decoder/layers.0/Add add [-0.680281f, 0.579606f] 30
/decoder/layers.0/encoder_attn_layer_norm/Add_1 nn.layer_norm [-8.14458f, 7.72507f] 27
CustomOp/linalg::matMulNBits<31, 31, 28>10 contrib.epu.quadric_custom_op [-9.41189f, 8.99864f] 27
CustomOp/linalg::matMulNBits<31, 31, 31>0 contrib.epu.quadric_custom_op [-6.84735f, 9.78763f] 27
CustomOp/linalg::matMulNBits<31, 31, 31>1 contrib.epu.quadric_custom_op [-3.05484f, 3.37927f] 29
CustomOp/nn::whisperAttention<384, 6, 1500, 14>34 contrib.epu.quadric_custom_op [-1.38653f, 0.946689f] 30
CustomOp/linalg::matMulNBits<31, 31, 31>11 contrib.epu.quadric_custom_op [-0.550637f, 0.728843f] 31
/decoder/layers.0/Add_1 add [-0.880807f, 1.2834f] 29
/decoder/layers.0/final_layer_norm/Add_1 nn.layer_norm [-8.31163f, 20.6317f] 26
CustomOp/linalg::matMulNBits<31, 31, 28>12 contrib.epu.quadric_custom_op [-7.3799f, 10.9955f] 27
/decoder/layers.0/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 10.9955f] 27
CustomOp/linalg::matMulNBits<31, 31, 30>13 contrib.epu.quadric_custom_op [-11.5507f, 1.70249f] 27
/decoder/layers.0/Add_2 add [-11.7435f, 1.55111f] 27
/decoder/layers.1/self_attn_layer_norm/Add_1 nn.layer_norm [-9.22594f, 7.0071f] 27
CustomOp/linalg::matMulNBits<31, 31, 27>14 contrib.epu.quadric_custom_op [-5.99958f, 9.36083f] 27
CustomOp/nn::whisperAttention<384, 6, 4, 14>35 contrib.epu.quadric_custom_op [-0.677723f, 0.539847f] 31
CustomOp/linalg::matMulNBits<31, 31, 31>15 contrib.epu.quadric_custom_op [-0.59268f, 1.41659f] 30
/decoder/layers.1/Add add [-11.9694f, 1.4109f] 27
/decoder/layers.1/encoder_attn_layer_norm/Add_1 nn.layer_norm [-31.863f, 18.7409f] 26
CustomOp/linalg::matMulNBits<31, 31, 27>16 contrib.epu.quadric_custom_op [-8.70827f, 11.8252f] 27
CustomOp/linalg::matMulNBits<31, 31, 31>2 contrib.epu.quadric_custom_op [-6.54716f, 5.91384f] 28
CustomOp/linalg::matMulNBits<31, 31, 31>3 contrib.epu.quadric_custom_op [-4.21574f, 3.62095f] 28
CustomOp/nn::whisperAttention<384, 6, 1500, 14>36 contrib.epu.quadric_custom_op [-1.40372f, 2.0292f] 29
CustomOp/linalg::matMulNBits<31, 31, 30>17 contrib.epu.quadric_custom_op [-1.2742f, 0.783488f] 30
/decoder/layers.1/Add_1 add [-11.7797f, 1.31481f] 27
/decoder/layers.1/final_layer_norm/Add_1 nn.layer_norm [-34.356f, 17.1086f] 25
CustomOp/linalg::matMulNBits<31, 31, 27>18 contrib.epu.quadric_custom_op [-18.9478f, 71.5836f] 24
/decoder/layers.1/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 71.5836f] 24
CustomOp/linalg::matMulNBits<31, 31, 29>19 contrib.epu.quadric_custom_op [-113.528f, 8.97997f] 24
/decoder/layers.1/Add_2 add [-115.08f, 10.1623f] 24
/decoder/layers.2/self_attn_layer_norm/Add_1 nn.layer_norm [-9.23459f, 15.3414f] 26
CustomOp/linalg::matMulNBits<31, 31, 28>20 contrib.epu.quadric_custom_op [-5.45088f, 5.13565f] 28
CustomOp/nn::whisperAttention<384, 6, 4, 14>37 contrib.epu.quadric_custom_op [-0.822134f, 0.739178f] 31
CustomOp/linalg::matMulNBits<31, 31, 29>21 contrib.epu.quadric_custom_op [-0.471779f, 3.29612f] 29
/decoder/layers.2/Add add [-115.007f, 9.87098f] 24
/decoder/layers.2/encoder_attn_layer_norm/Add_1 nn.layer_norm [-32.0424f, 18.1859f] 25
CustomOp/linalg::matMulNBits<31, 31, 28>22 contrib.epu.quadric_custom_op [-10.3365f, 8.57053f] 27
CustomOp/linalg::matMulNBits<31, 31, 31>4 contrib.epu.quadric_custom_op [-7.2206f, 6.83287f] 28
CustomOp/linalg::matMulNBits<31, 31, 31>5 contrib.epu.quadric_custom_op [-4.22371f, 4.76144f] 28
CustomOp/nn::whisperAttention<384, 6, 1500, 14>38 contrib.epu.quadric_custom_op [-1.2737f, 1.75648f] 30
CustomOp/linalg::matMulNBits<31, 31, 30>23 contrib.epu.quadric_custom_op [-1.6217f, 1.06462f] 30
/decoder/layers.2/Add_1 add [-115.65f, 10.1965f] 24
/decoder/layers.2/final_layer_norm/Add_1 nn.layer_norm [-27.5322f, 13.5401f] 26
CustomOp/linalg::matMulNBits<31, 31, 28>24 contrib.epu.quadric_custom_op [-7.36932f, 2.67375f] 28
/decoder/layers.2/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 2.66372f] 29
CustomOp/linalg::matMulNBits<31, 31, 29>25 contrib.epu.quadric_custom_op [-3.60221f, 1.07261f] 29
/decoder/layers.2/Add_2 add [-115.773f, 9.67873f] 24
/decoder/layers.3/self_attn_layer_norm/Add_1 nn.layer_norm [-15.5193f, 13.6514f] 27
CustomOp/linalg::matMulNBits<31, 31, 28>26 contrib.epu.quadric_custom_op [-7.85931f, 6.792f] 28
CustomOp/nn::whisperAttention<384, 6, 4, 14>39 contrib.epu.quadric_custom_op [-2.41125f, 1.44966f] 29
CustomOp/linalg::matMulNBits<31, 31, 27>27 contrib.epu.quadric_custom_op [-2.09421f, 3.09023f] 29
/decoder/layers.3/Add add [-115.747f, 9.95996f] 24
/decoder/layers.3/encoder_attn_layer_norm/Add_1 nn.layer_norm [-26.1427f, 9.04464f] 26
CustomOp/linalg::matMulNBits<31, 31, 28>28 contrib.epu.quadric_custom_op [-8.85646f, 9.43103f] 27
CustomOp/linalg::matMulNBits<31, 31, 31>6 contrib.epu.quadric_custom_op [-8.39829f, 7.86805f] 27
CustomOp/linalg::matMulNBits<31, 31, 31>7 contrib.epu.quadric_custom_op [-5.39744f, 6.1728f] 28
CustomOp/nn::whisperAttention<384, 6, 1500, 14>40 contrib.epu.quadric_custom_op [-2.70359f, 3.22433f] 29
CustomOp/linalg::matMulNBits<31, 31, 27>29 contrib.epu.quadric_custom_op [-4.77284f, 1.9294f] 28
/decoder/layers.3/Add_1 add [-115.479f, 10.4301f] 24
/decoder/layers.3/final_layer_norm/Add_1 nn.layer_norm [-46.1546f, 11.7791f] 25
CustomOp/linalg::matMulNBits<31, 31, 28>30 contrib.epu.quadric_custom_op [-17.636f, 17.31f] 26
/decoder/layers.3/activation_fn/Mul_1 contrib.epu.gelu [-0.169971f, 17.31f] 26
CustomOp/linalg::matMulNBits<31, 31, 29>31 contrib.epu.quadric_custom_op [-8.89668f, 11.8703f] 27
/decoder/layers.3/Add_2 add [-110.711f, 8.84814f] 24
/decoder/layer_norm/Add_1 nn.layer_norm [-238.585f, 80.7449f] 23
CustomOp/linalg::matMulNBits<31, 31, 31>32 contrib.epu.quadric_custom_op (-64.0, 63.99999997019768) 25
contrib.epu.quadric_custom_op (-2147483648, 2147483647) 0
2026-09-22 03:27 - INFO - epu - codegen - START====================build_cpu_runnable_fx_relay
2026-09-22 03:27 - INFO - epu - codegen - START=======================quantize_to_chimera_fx
2026-09-22 03:27 - INFO - epu - codegen - START=================================relay_to_tir
2026-09-22 03:27 - INFO - epu - codegen - START===========================relay_to_epu_relay
2026-09-22 03:27 - INFO - epu - codegen - START==============================adapt_and_order
2026-09-22 03:27 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 03:27 - INFO - epu - codegen - START=============================plan_lrm_virtual
2026-09-22 03:27 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 03:27 - INFO - epu - codegen - START===============================lrm_alloc_loop
2026-09-22 03:28 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 03:28 - INFO - epu - codegen - START================================lrm_splitting
2026-09-22 03:29 - INFO - epu - codegen - START==============================ext_split_relay
2026-09-22 03:29 - INFO - epu - codegen - START====================================build_tir
2026-09-22 03:31 - INFO - epu - iss_testing - No tranges found for input, use default float range: <tvm.contrib.epu.interval.Interval object at 0x7d83171d97e0>
2026-09-22 03:31 - INFO - epu - iss_testing - Found tranges for input: <tvm.contrib.epu.interval.Interval object at 0x7d83171d8d00>
2026-09-22 03:31 - INFO - epu - iss_testing - Started Running Graph on Chimera ISS...
FILM 73/73: 100%|███████████████████████████████████████████████████| 73/73 [02:03<00:00, 1.70s/it]
2026-09-22 03:33 - INFO - epu - iss_testing - Done 0:02:04.145357
GPNPU tokens: [50259, 50359, 50363, 2221]
ORT tokens: [50259, 50359, 50363, 2221]
Token ids match ORT: True
Decoded: <|en|><|transcribe|><|notimestamps|> Mr
7. GPNPU Performance
The ISS writes one profile_core*.json per core into each module's build directory; these read core 0. The chart is the SDK's own breakdown as an image, which the docs catalog uses as this notebook's card.
Hardware config QC-P, 32 MB OCM, 16 MACs/PE, 128 GBps, 1.7 GHz, 4 cores.
| Stage | Cycles | ms @ 1.7 GHz |
|---|---|---|
| Encoder — waveform → hidden states | 43.96 M | 25.86 |
whisper::logMel | 7.40 M | 4.35 |
whisper::convStem | 7.35 M | 4.33 |
| Decoder — hidden states → token ids | 4.28 M | 2.52 |
| End to end | 48.24 M | 28.38 |
for module, stage in (
("whisper_e2e_chipy", "Encoder: waveform -> hidden states"),
("whisper_dec_chipy", "Decoder: hidden states -> token ids"),
):
profile = sorted(glob.glob(f"ccl_build/{module}/**/profile_core0.json", recursive=True))[0]
counts = chimera_core._plot_profile_results(profile, clock_freq=hw_config.clock_freq_ghz * 1e9)
cycles = {k: v / 1e6 for k, v in counts.items() if k not in ("total", "ExtBytes")}
figure, axis = plt.subplots(figsize=(8, 3.2))
axis.barh(list(cycles), list(cycles.values()), color="#3d6098")
axis.invert_yaxis()
axis.set_xlabel("cycles (millions)")
axis.set_title(f"{stage} — {counts['total'] / 1e6:,.2f}M cycles")
for row, value in enumerate(cycles.values()):
axis.text(value, row, f" {value:,.1f}M", va="center", fontsize=9)
figure.tight_layout()
plt.show()
[SDK-CLI] : TotalCycles: 43,925,705
[SDK-CLI] : Executions/second: 39
compute : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 19.679M
data_array : ▇▇▇▇▇▇▇▇ 3.478M
mac : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 6.154M
data_ocm : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 11.834M
data_external: ▇▇▇▇▇ 2.302M

[SDK-CLI] : TotalCycles: 4,297,756
[SDK-CLI] : Executions/second: 396
compute : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 1.907M
data_array : ▇▇▇▇▇▇▇▇▇▇▇▇ 487.555K
mac : ▇▇▇▇▇▇▇▇▇ 380.016K
data_ocm : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 1.145M
data_external: ▇▇▇▇▇▇▇▇▇ 372.422K

Summary
| Model | Whisper-Tiny (39 M parameters, 4 encoder + 4 decoder layers, embed_dim 384) |
| On the GPNPU | log-mel, conv front-end, encoder, decoder prefill, greedy argmax |
| On the host | FLAC decode only |
| Programs | 2 — waveform → hidden states, (ids, hidden) → tokens; stitched with ChiPy |
| Target | QC-P, 32 MB OCM, 16 MACs/PE, 128 GBps, 4 cores |
| Quantization | INT4 weights via MatMulNBits, FP16 attention |
| Custom Ops | linalg::matMulNBits, nn::whisperAttention, whisper::convStem, whisper::logMel, whisper::greedyTokens — all in whisper_stubs.hpp |
| Accuracy (observed) | ~54.5 dB PSNR on the encoder hidden states; token ids exactly matching ONNX Runtime |
Citation
@article{radford2022whisper,
title = {Robust Speech Recognition via Large-Scale Weak Supervision},
author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg
and McLeavey, Christine and Sutskever, Ilya},
journal = {arXiv preprint arXiv:2212.04356},
year = {2022}
}
