Unlock this feature

This feature isn’t part of your plan yet

Contact sales to get upgraded to the full DevStudio experience.

Unlock this feature

This feature isn't part of your plan yet.

Model Demo: Whisper Tiny (Waveform to Tokens)


NOTE: The Jupyter Notebook below is included in the Chimera SDK and can be run interactively by running the following CLI command:

$ quadric sdk notebook

From the Jupyter Notebook window in your browser, select the notebook named /quadric/sdk-cli/examples/models/whisper/whisper_tutorial.ipynb.


Whisper-Tiny Speech Recognition on Chimera GPNPU

This notebook runs OpenAI's Whisper-Tiny on the Quadric Chimera GPNPU, from a raw waveform to decoded token ids.

The pre- and postprocessing that normally runs on the host — the log-mel feature extractor and the greedy argmax — are CCL custom ops stitched to the ONNX graphs with ChiPy, so the host only decodes the audio file.

Pipeline

Whisper-Tiny on the Chimera GPNPU: waveform through log-mel, conv stem, encoder, decoder prefill and greedy argmax, as two ChiPy programs

StageRuns onInputOutput
FLAC decodehost CPULibriSpeech clip16 kHz fp32 waveform
Log-mel (whisper::logMel)GPNPU[1, 480000] waveform[1, 80, 3000] log-mel
Conv front-end (whisper::convStem)GPNPUlog-mel[1, 1500, 384] features
Encoder, 4 layersGPNPUfeatures[1, 1500, 384] hidden states
Decoder prefill, 4 layersGPNPUhidden states + 4 prompt ids[1, 4, 51865] logits
Greedy argmax (whisper::greedyTokens)GPNPUlogits[1, 4] token ids

1. Setup

Install the Python dependencies and download the Whisper-Tiny encoder and decoder ONNX models from the Quadric model S3 bucket.

%pip install -q -r whisper_requirements.txt

import glob
import urllib.request
from pathlib import Path

import matplotlib.pyplot as plt
import numpy as np
import onnxruntime as ort

from tvm.contrib.epu import chipy
from tvm.contrib.epu.chimera_job import core as chimera_core
from tvm.contrib.epu.chimera_job.hw_config import HWConfig

from whisper_helper.encoder_helper import replace_encoder_custom_ops
from whisper_helper.decoder_helper import replace_decoder_custom_ops
from whisper_helper.generate_tranges import generate_tranges
from whisper_helper.prepost import (
    DECODER_COMPILE_TRANGES,
    KERNEL_HPP,
    N_WAVE,
    STEM_D,
    STEM_T2,
    build_decoder_pipeline,
    build_encoder_pipeline,
    build_mel_constants,
    build_stem_constants,
    compile_context,
    encoder_compile_tranges,
    ensure_decoder_models,
    numpy_float_mel,
    report,
)
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv
Note: you may need to restart the kernel to use updated packages.
S3_BASE = "https://sdk-cli-models.s3.amazonaws.com"
S3_MODELS = {
    "encoder_model.onnx": f"{S3_BASE}/whisper_encoder_model.onnx",
    "decoder_prefill.onnx": f"{S3_BASE}/whisper_decoder_prefill.onnx",
}


def download_models(model_dir: Path, models: dict = S3_MODELS) -> None:
    """Download model files from S3 if not already present locally.

    Parameters
    ----------
    model_dir : Path
        Local directory where model files are stored.
    models : dict, optional
        Mapping of ``{local_filename: s3_url}``.  Defaults to ``S3_MODELS``.
    """
    for local_name, url in models.items():
        local_path = model_dir / local_name
        if local_path.exists():
            print(f"  {local_name}: exists ({local_path.stat().st_size:,} bytes)")
            continue
        print(f"  Downloading {local_name}...")
        urllib.request.urlretrieve(url, str(local_path))
        print(f"  {local_name}: {local_path.stat().st_size:,} bytes")


MODEL_DIR = Path(".")
print("Checking / downloading model files...")
download_models(MODEL_DIR)

ENCODER_ONNX = "encoder_model.onnx"
ENCODER_TRANGES = "encoder_model.tranges"
ENCODER_GPNPU = "encoder_model_gpnpu.onnx"
DECODER_ONNX = "decoder_prefill.onnx"
DECODER_TRANGES = "decoder_prefill.tranges"
DECODER_MATCHED = "decoder_model_matched.onnx"
ENCODER_ONNX_PATH = MODEL_DIR / ENCODER_ONNX
DECODER_ONNX_PATH = MODEL_DIR / DECODER_ONNX

## One hardware configuration for both programs.
hw_config = HWConfig(
    product="QC-P",
    ocm_size="32MB",
    macs_per_pe=16,
    num_cores=4,
    ext_rd_bw="128GBps",
    ext_wr_bw="128GBps",
)
Checking / downloading model files...
  Downloading encoder_model.onnx...
  encoder_model.onnx: 8,331,061 bytes
  Downloading decoder_prefill.onnx...
  decoder_prefill.onnx: 95,466,860 bytes

2. Audio & Calibration

Whisper consumes 16 kHz mono audio; the GPNPU takes the padded waveform directly and produces the mel itself.

whisper_helper/data_prep.py loads the LibriSpeech "dummy" validation split and runs each sample through the floating-point encoder to produce activations for the calibration step.

from whisper_helper.data_prep import prepare_whisper_data

SAMPLE_INDEX = 0
NUM_CALIBRATION_SAMPLES = 73

data = prepare_whisper_data(
    encoder_onnx=ENCODER_ONNX_PATH,
    decoder_onnx=DECODER_ONNX_PATH,
    num_calibration_samples=NUM_CALIBRATION_SAMPLES,
    sample_index=SAMPLE_INDEX,
)

encoder_session = data.encoder_session
decoder_session = data.decoder_session
calibration_mels = data.calibration_mels
calibration_encoder_outputs = data.calibration_encoder_outputs

## The GPNPU program starts at the waveform: pad the clip to Whisper's 30 s.
waveform = np.zeros(N_WAVE, dtype=np.float32)
waveform[: min(len(data.audio), N_WAVE)] = data.audio[:N_WAVE]
print(f"waveform: {len(data.audio) / 16000:.1f}s of audio -> {waveform.shape}")
print(f"reference: {data.reference_text}")
preprocessor_config.json: 0.00B [00:00, ?B/s]



tokenizer_config.json: 0.00B [00:00, ?B/s]



vocab.json: 0.00B [00:00, ?B/s]



tokenizer.json: 0.00B [00:00, ?B/s]



merges.txt: 0.00B [00:00, ?B/s]



normalizer.json: 0.00B [00:00, ?B/s]



added_tokens.json: 0.00B [00:00, ?B/s]



special_tokens_map.json: 0.00B [00:00, ?B/s]


Primary: 1272-128104-0000.flac  |  ref: MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME HIS GOSPEL
Calibration: 73 mels, 73 encoder outputs
waveform: 5.9s of audio -> (480000,)
reference: MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME HIS GOSPEL

3. Tensor-Range Calibration

Tensor ranges (tranges) record per-tensor min/max statistics from the calibration data. The compiler uses these to assign fixed-point fractional bits.

Encoder ranges come from the mel spectrograms, decoder ranges from the encoder's floating-point outputs.

print(f"Generating encoder tranges ({len(calibration_mels)} samples)...")
generate_tranges(
    str(ENCODER_ONNX_PATH),
    ENCODER_TRANGES,
    num_samples=len(calibration_mels),
    mel_spectrograms=calibration_mels,
)
print(f"Generated {ENCODER_TRANGES}")

print(f"\nGenerating decoder tranges ({len(calibration_encoder_outputs)} samples)...")
generate_tranges(
    str(DECODER_ONNX_PATH),
    DECODER_TRANGES,
    num_samples=len(calibration_encoder_outputs),
    encoder_outputs_list=calibration_encoder_outputs,
)
print(f"Generated {DECODER_TRANGES}")
Generating encoder tranges (73 samples)...
Processing encoder_model.onnx...
  Inputs: [('input_features', [1, 80, 3000])]
  Using 73 real audio samples
  Running calibration with 73 real audio samples...
  Saved 224 tensor ranges to encoder_model.tranges
Generated encoder_model.tranges

Generating decoder tranges (73 samples)...
Processing decoder_prefill.onnx...
  Inputs: [('input_ids', [1, 4]), ('encoder_hidden_states', [1, 1500, 384])]
  Using 73 real encoder outputs
  Running calibration with 73 random samples...
  Saved 345 tensor ranges to decoder_prefill.tranges
Generated decoder_prefill.tranges

4. Custom Op Replacement

Both graphs get the same treatment: every MatMulNBits subgraph becomes a linalg::matMulNBits custom op and every attention core becomes nn::whisperAttention. The attention-adjacent activation edges are then declared float16 so the FP16 kernels exchange them directly.

4.1 Encoder

The encoder adds one more pass: its convolutional front-end (Conv1 → GELU → Conv2 → GELU → Transpose → PositionalEmbedding, 15 nodes) collapses into a single whisper::convStem node. The result is one graph whose input is still input_features — the mel arrives from logMel.

encoder_graph_info = replace_encoder_custom_ops(
    encoder_onnx=ENCODER_ONNX,
    encoder_tranges=ENCODER_TRANGES,
    encoder_path=ENCODER_GPNPU,
    conv_stem_consts=build_stem_constants(),
)
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_392__v_393__v_394
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_391
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_350__v_351__v_352
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_349
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_315__v_316__v_317
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_314
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_310__v_312__v_318
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_309


    L0 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
    L0 attn core
    L0 outproj → linalg::matMulNBits<31, 31, 28>
    L0 fc1 → linalg::matMulNBits<31, 31, 29>
    L0 fc2 → linalg::matMulNBits<31, 31, 29>
  Layer 0: all custom ops replaced
    L1 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
    L1 attn core
    L1 outproj → linalg::matMulNBits<31, 31, 30>
    L1 fc1 → linalg::matMulNBits<31, 31, 28>
    L1 fc2 → linalg::matMulNBits<31, 31, 28>
  Layer 1: all custom ops replaced
    L2 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
    L2 attn core
    L2 outproj → linalg::matMulNBits<31, 31, 28>
    L2 fc1 → linalg::matMulNBits<31, 31, 27>
    L2 fc2 → linalg::matMulNBits<31, 31, 28>
  Layer 2: all custom ops replaced
    L3 QKV proj (+bias) → linalg::matMulNBits<31, 31, 28>
    L3 attn core
    L3 outproj → linalg::matMulNBits<31, 31, 29>
    L3 fc1 → linalg::matMulNBits<31, 31, 28>
    L3 fc2 → linalg::matMulNBits<31, 31, 27>
  Layer 3: all custom ops replaced
Enabled fp16 attention edges (8 edges)
Saved encoder: 130 nodes, input=input_features → encoder_model_gpnpu.onnx

4.2 Decoder

The decoder prefill processes the four Whisper bootstrap tokens (<|startoftranscript|>, <|en|>, <|transcribe|>, <|notimestamps|>) with cross-attention to the encoder's hidden states, so the replacement also covers the cross-attention variant of nn::whisperAttention.

ensure_decoder_models then derives the two variants the pipeline needs: a logits-only float model for CPU mode, and the matched model with int32 input_ids.

replace_decoder_custom_ops(
    decoder_onnx=DECODER_ONNX,
    decoder_tranges=DECODER_TRANGES,
    decoder_matched_path=DECODER_MATCHED,
)

decoder_float, decoder_gpnpu = ensure_decoder_models()
print(f"CPU-mode model: {decoder_float.name}\nGPNPU model:    {decoder_gpnpu.name}")
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_612__v_614__v_615
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_489__v_490__v_491
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_480__v_481__v_482
WARNING:root:ONNX node does not have a name, deriving temp name from its outputs. Node was named to: _v_471__v_472__v_473


    /decoder/layers.0/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.0/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.1/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.1/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.2/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.2/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.3/encoder_attn/k_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.3/encoder_attn/v_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    fused_qkv/_v_611 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.0/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.0/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.0/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.0/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.0/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 30>
    fused_qkv/_v_488 → linalg::matMulNBits<31, 31, 27>
    /decoder/layers.1/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 31>
    /decoder/layers.1/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
    /decoder/layers.1/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 30>
    /decoder/layers.1/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
    /decoder/layers.1/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
    fused_qkv/_v_479 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.2/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
    /decoder/layers.2/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.2/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 30>
    /decoder/layers.2/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.2/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
    fused_qkv/_v_470 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.3/self_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
    /decoder/layers.3/encoder_attn/q_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.3/encoder_attn/out_proj/MatMul_Q4 → linalg::matMulNBits<31, 31, 27>
    /decoder/layers.3/fc1/MatMul_Q4 → linalg::matMulNBits<31, 31, 28>
    /decoder/layers.3/fc2/MatMul_Q4 → linalg::matMulNBits<31, 31, 29>
    lm_head → linalg::matMulNBits<31, 31, 31>
    L0 self-attn FP16 MHA core
    L0 cross-attn FP16 MHA core
    L1 self-attn FP16 MHA core
    L1 cross-attn FP16 MHA core
    L2 self-attn FP16 MHA core
    L2 cross-attn FP16 MHA core
    L3 self-attn FP16 MHA core
    L3 cross-attn FP16 MHA core
CPU-mode model: _decoder_prefill_logits.onnx
GPNPU model:    _decoder_matched_ids32.onnx

5. Pipeline Setup and Build

ChiPy stitches each graph to its custom ops as a single traced program: the encoder takes the log-mel op on the input side, the decoder takes the greedy argmax on the output side.

Calling a traced function directly runs it in CPU mode, where each custom op falls back to its Python reference. That checks the wiring in seconds, before spending minutes on a compile.

5.1 Encoder

ref_mel, fb_float = numpy_float_mel(data.audio, "reflect")
mel_consts = build_mel_constants(fb_float)

ort_hidden = encoder_session.run(None, {"input_features": ref_mel[None].astype(np.float32)})[0]

encoder_pipeline = build_encoder_pipeline(mel_consts)

cpu_out = encoder_pipeline(
    chipy.Tensor(shape=(1, N_WAVE), dtype="float32", values=waveform),
    ENCODER_ONNX,
)
cpu_psnr = report("chipy CPU vs ORT hidden", ort_hidden, np.asarray(cpu_out.values))

## CPU mode runs the same graph under ONNX Runtime, so agreement is numerical
## noise, not quantization: a low PSNR here means the wiring is wrong.
CPU_PSNR_MIN_DB = 100.0
assert cpu_psnr >= CPU_PSNR_MIN_DB, f"CPU-mode wiring regression: {cpu_psnr:.2f} dB"
  chipy CPU vs ORT hidden              max|e|=0.00000  rms=0.000000  PSNR=630.91 dB

5.2 Decoder

CPU mode is checked against the floating-point hidden states; section 6 then feeds the decoder the encoder's device output.

decoder_input_ids = np.array([[50258, 50259, 50359, 50363]], dtype=np.int32)
decoder_ort = ort.InferenceSession(str(decoder_float), providers=["CPUExecutionProvider"])


def ort_tokens_for(hidden):
    """Greedy token ids from the float decoder, for a given hidden state."""
    logits = decoder_ort.run(
        None, {"input_ids": decoder_input_ids, "encoder_hidden_states": hidden}
    )[0]
    return logits[0].argmax(-1)


decoder_pipeline = build_decoder_pipeline()

cpu_out = decoder_pipeline(
    chipy.Tensor(shape=(1, 4), dtype="int32", values=decoder_input_ids),
    chipy.Tensor(shape=(1, STEM_T2, STEM_D), dtype="float32", values=ort_hidden),
    str(decoder_float),
)
cpu_tokens = np.asarray(cpu_out.values).ravel().astype(np.int64)
host_tokens = ort_tokens_for(ort_hidden)
print(f"chipy CPU tokens: {cpu_tokens}  match={np.array_equal(cpu_tokens, host_tokens)}")
assert np.array_equal(
    cpu_tokens, host_tokens
), f"CPU-mode token mismatch: {cpu_tokens} != {host_tokens}"
chipy CPU tokens: [50259 50359 50363  2221]  match=True

6. Compile, Run, and Validation

pipeline.compile() hands a whole program to the Chimera Graph Compiler: the custom-op kernels and the ONNX graph become one module under ccl_build/. .run() then executes it on the Instruction Set Simulator.

Both programs compile for the same target, so both calls pass the same hw_config.

CheckObservedAsserted gate
Encoder hidden states vs ONNX Runtime~54.5 dB PSNR>= 20 dB and MSE <= 5
Token ids vs ONNX Runtime, from device hidden statesexact matchexact
Token ids vs the all-host chainexact matchexact

The encoder gate is deliberately far below the observed value: it catches a broken pipeline, not a small numerical drift.

6.1 Encoder — waveform → hidden states

with compile_context(qlut=False):
    encoder_compiled = encoder_pipeline.compile(
        hw_config=hw_config,
        module_name="whisper_e2e_chipy",
        custom_op_header=str(KERNEL_HPP),
        value_proto_ranges=encoder_compile_tranges(),
        waveform=chipy.Tensor(shape=(1, N_WAVE), dtype="float32"),
        encoder=ENCODER_GPNPU,
    )

encoder_out = encoder_compiled.run(waveform=waveform)
iss_hidden = (
    np.asarray(next(iter(encoder_out.values())).tensor).reshape(ort_hidden.shape).astype(np.float32)
)

psnr = report("ISS vs ORT hidden", ort_hidden, iss_hidden)
mse = float(np.mean((ort_hidden - iss_hidden) ** 2))

ENCODER_PSNR_MIN_DB = 20.0
assert (
    psnr >= ENCODER_PSNR_MIN_DB
), f"Encoder PSNR regression: {psnr:.2f} dB < {ENCODER_PSNR_MIN_DB} dB"

## MSE catches blow-ups that PSNR can mask when the signal range is wider than
## expected (the encoder output's range grows with calibration drift, which
## can keep PSNR steady while errors climb in absolute terms).
ENCODER_MSE_MAX = 5  # ~5x headroom over current MSE; PSNR 20 dB equivalent
assert mse <= ENCODER_MSE_MAX, f"Encoder MSE regression: {mse:.4f} > {ENCODER_MSE_MAX}"
/quadric/sdk-cli/examples/models/whisper/whisper_helper/prepost.py:57: RuntimeWarning: invalid value encountered in cast
  wave[: len(audio)] = audio[:N_WAVE]
2026-09-22 02:44 - INFO - epu - codegen - START==================================build_relay
2026-09-22 02:44 - INFO - epu - codegen - START===============================optimize_relay
2026-09-22 02:44 - INFO - epu - codegen - START====================quantize_to_cpu_runnable_fx
2026-09-22 02:44 - INFO - epu - fx - 

Source name                                 Op                             Output 0 Range              Output 0 Frac Bits
------------------------------------------  -----------------------------  --------------------------  --------------------
                                            contrib.epu.quadric_custom_op  (-32.0, 31.99999998509884)  26
CustomOp/whisper::convStem0                 contrib.epu.quadric_custom_op  (-8.0, 7.99999999627471)    28
/layers.0/self_attn_layer_norm/Add_1        nn.layer_norm                  [-16.5387f, 11.5336f]       26
CustomOp/linalg::matMulNBits<31, 31, 28>2   contrib.epu.quadric_custom_op  [-4.96675f, 2.94481f]       28
/layers.0/Add                               add                            [-2.08293f, 3.24672f]       27
/layers.0/final_layer_norm/Add_1            nn.layer_norm                  [-14.6758f, 21.1295f]       26
CustomOp/linalg::matMulNBits<31, 31, 29>3   contrib.epu.quadric_custom_op  [-10.8754f, 12.0852f]       27
/layers.0/activation_fn/Mul_1               contrib.epu.gelu               [-0.169971f, 12.0852f]      27
CustomOp/linalg::matMulNBits<31, 31, 29>4   contrib.epu.quadric_custom_op  [-4.11117f, 4.20129f]       28
/layers.0/Add_1                             add                            [-4.7491f, 3.81342f]        27
/layers.1/self_attn_layer_norm/Add_1        nn.layer_norm                  [-10.1508f, 14.0432f]       27
CustomOp/linalg::matMulNBits<31, 31, 30>7   contrib.epu.quadric_custom_op  [-1.57465f, 2.00129f]       29
/layers.1/Add                               add                            [-4.0395f, 3.76171f]        27
/layers.1/final_layer_norm/Add_1            nn.layer_norm                  [-13.3625f, 16.9489f]       26
CustomOp/linalg::matMulNBits<31, 31, 28>8   contrib.epu.quadric_custom_op  [-10.0635f, 13.8182f]       27
/layers.1/activation_fn/Mul_1               contrib.epu.gelu               [-0.169971f, 13.8182f]      27
CustomOp/linalg::matMulNBits<31, 31, 28>9   contrib.epu.quadric_custom_op  [-6.80559f, 7.79773f]       28
/layers.1/Add_1                             add                            [-10.0096f, 9.14031f]       27
/layers.2/self_attn_layer_norm/Add_1        nn.layer_norm                  [-21.9773f, 22.2916f]       26
CustomOp/linalg::matMulNBits<31, 31, 28>12  contrib.epu.quadric_custom_op  [-4.5468f, 3.53895f]        28
/layers.2/Add                               add                            [-9.32271f, 9.04946f]       27
/layers.2/final_layer_norm/Add_1            nn.layer_norm                  [-27.1201f, 24.2017f]       26
CustomOp/linalg::matMulNBits<31, 31, 27>13  contrib.epu.quadric_custom_op  [-40.5755f, 128.95f]        23
/layers.2/activation_fn/Mul_1               contrib.epu.gelu               [-0.169971f, 128.95f]       23
CustomOp/linalg::matMulNBits<31, 31, 28>14  contrib.epu.quadric_custom_op  [-27.8214f, 248.118f]       23
/layers.2/Add_1                             add                            [-30.3549f, 253.457f]       22
/layers.3/self_attn_layer_norm/Add_1        nn.layer_norm                  [-21.1582f, 33.7291f]       25
CustomOp/linalg::matMulNBits<31, 31, 29>17  contrib.epu.quadric_custom_op  [-5.37736f, 11.7757f]       27
/layers.3/Add                               add                            [-29.5622f, 264.989f]       22
/layers.3/final_layer_norm/Add_1            nn.layer_norm                  [-94.0882f, 120.342f]       24
CustomOp/linalg::matMulNBits<31, 31, 28>18  contrib.epu.quadric_custom_op  [-119.89f, 329.025f]        22
/layers.3/activation_fn/Mul_1               contrib.epu.gelu               [-0.169971f, 329.025f]      22
CustomOp/linalg::matMulNBits<31, 31, 27>19  contrib.epu.quadric_custom_op  [-121.728f, 52.5513f]       24
/layers.3/Add_1                             add                            [-131.772f, 276.199f]       22
/layer_norm/Add_1                           nn.layer_norm                  [-17.5736f, 19.0347f]       -

2026-09-22 02:44 - INFO - epu - codegen - START====================build_cpu_runnable_fx_relay
2026-09-22 02:44 - INFO - epu - codegen - START=======================quantize_to_chimera_fx
2026-09-22 02:44 - INFO - epu - codegen - START=================================relay_to_tir
2026-09-22 02:44 - INFO - epu - codegen - START===========================relay_to_epu_relay
2026-09-22 02:44 - INFO - epu - codegen - START==============================adapt_and_order
2026-09-22 02:44 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 02:44 - INFO - epu - codegen - START=============================plan_lrm_virtual
2026-09-22 02:44 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 02:45 - INFO - epu - codegen - START===============================lrm_alloc_loop
2026-09-22 02:45 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 02:45 - INFO - epu - codegen - START================================lrm_splitting
2026-09-22 02:46 - INFO - epu - codegen - START==============================ext_split_relay
2026-09-22 02:46 - INFO - epu - codegen - START====================================build_tir
2026-09-22 02:48 - INFO - epu - iss_testing - Found tranges for input: <tvm.contrib.epu.interval.Interval object at 0x7d843e6ccf10>
2026-09-22 02:48 - INFO - epu - iss_testing - Started Running Graph on Chimera ISS...
FILM 46/46: 100%|███████████████████████████████████████████████████| 46/46 [33:46<00:00, 44.06s/it]
2026-09-22 03:27 - INFO - epu - iss_testing - Done 0:38:58.354943


  ISS vs ORT hidden                    max|e|=10.82557  rms=0.065826  PSNR= 54.55 dB

6.2 Decoder — hidden states → token ids

The hidden states come from the encoder's ISS run in 6.1, so this is an end-to-end device check. The tokens are compared twice: against ORT on those same device hidden states, which isolates the decoder, and against the all-host chain from 5.2, which catches encoder drift that flips a token.

with compile_context(qlut=True):
    decoder_compiled = decoder_pipeline.compile(
        hw_config=hw_config,
        module_name="whisper_dec_chipy",
        custom_op_header=str(KERNEL_HPP),
        value_proto_ranges=DECODER_COMPILE_TRANGES,
        input_ids=chipy.Tensor(shape=(1, 4), dtype="int32"),
        encoder_hidden=chipy.Tensor(shape=(1, STEM_T2, STEM_D), dtype="float32"),
        decoder=str(decoder_gpnpu),
    )

decoder_out = decoder_compiled.run(input_ids=decoder_input_ids, encoder_hidden=iss_hidden)
iss_tokens = np.asarray(next(iter(decoder_out.values())).tensor).ravel().astype(np.int64)
ort_tokens = ort_tokens_for(iss_hidden)

tokens_match = np.array_equal(iss_tokens, ort_tokens)
print(f"  GPNPU tokens: {iss_tokens.tolist()}")
print(f"  ORT   tokens: {ort_tokens.tolist()}")
print(f"  Token ids match ORT: {tokens_match}")

## Same hidden states on both sides: isolates the decoder's own error.
assert tokens_match, f"Decoder token mismatch: {iss_tokens} != {ort_tokens}"

print(f"\n  Decoded: {data.processor.tokenizer.decode(iss_tokens)}")

## And against the all-host chain, so encoder drift that flips a token fails here.
assert np.array_equal(
    iss_tokens, host_tokens
), f"Full-chain token mismatch: device {iss_tokens} != all-host {host_tokens}"
2026-09-22 03:27 - INFO - epu - codegen - START==================================build_relay
2026-09-22 03:27 - INFO - epu - codegen - START===============================optimize_relay
2026-09-22 03:27 - INFO - epu - codegen - START====================quantize_to_cpu_runnable_fx
2026-09-22 03:27 - INFO - epu - fx - 

Source name                                        Op                             Output 0 Range                Output 0 Frac Bits
-------------------------------------------------  -----------------------------  --------------------------  --------------------
/decoder/embed_tokens/Gather                       contrib.epu.embedding          [-0.154297f, 0.20166f]                        31
/decoder/Add                                       add                            [-0.586182f, 0.470642f]                       31
/decoder/layers.0/self_attn_layer_norm/Add_1       nn.layer_norm                  [-5.01829f, 4.52495f]                         28
CustomOp/linalg::matMulNBits<31, 31, 28>8          contrib.epu.quadric_custom_op  [-7.66943f, 8.27715f]                         27
CustomOp/nn::whisperAttention<384, 6, 4, 14>33     contrib.epu.quadric_custom_op  [-0.863014f, 1.08392f]                        30
CustomOp/linalg::matMulNBits<31, 31, 31>9          contrib.epu.quadric_custom_op  [-0.691889f, 0.544409f]                       31
/decoder/layers.0/Add                              add                            [-0.680281f, 0.579606f]                       30
/decoder/layers.0/encoder_attn_layer_norm/Add_1    nn.layer_norm                  [-8.14458f, 7.72507f]                         27
CustomOp/linalg::matMulNBits<31, 31, 28>10         contrib.epu.quadric_custom_op  [-9.41189f, 8.99864f]                         27
CustomOp/linalg::matMulNBits<31, 31, 31>0          contrib.epu.quadric_custom_op  [-6.84735f, 9.78763f]                         27
CustomOp/linalg::matMulNBits<31, 31, 31>1          contrib.epu.quadric_custom_op  [-3.05484f, 3.37927f]                         29
CustomOp/nn::whisperAttention<384, 6, 1500, 14>34  contrib.epu.quadric_custom_op  [-1.38653f, 0.946689f]                        30
CustomOp/linalg::matMulNBits<31, 31, 31>11         contrib.epu.quadric_custom_op  [-0.550637f, 0.728843f]                       31
/decoder/layers.0/Add_1                            add                            [-0.880807f, 1.2834f]                         29
/decoder/layers.0/final_layer_norm/Add_1           nn.layer_norm                  [-8.31163f, 20.6317f]                         26
CustomOp/linalg::matMulNBits<31, 31, 28>12         contrib.epu.quadric_custom_op  [-7.3799f, 10.9955f]                          27
/decoder/layers.0/activation_fn/Mul_1              contrib.epu.gelu               [-0.169971f, 10.9955f]                        27
CustomOp/linalg::matMulNBits<31, 31, 30>13         contrib.epu.quadric_custom_op  [-11.5507f, 1.70249f]                         27
/decoder/layers.0/Add_2                            add                            [-11.7435f, 1.55111f]                         27
/decoder/layers.1/self_attn_layer_norm/Add_1       nn.layer_norm                  [-9.22594f, 7.0071f]                          27
CustomOp/linalg::matMulNBits<31, 31, 27>14         contrib.epu.quadric_custom_op  [-5.99958f, 9.36083f]                         27
CustomOp/nn::whisperAttention<384, 6, 4, 14>35     contrib.epu.quadric_custom_op  [-0.677723f, 0.539847f]                       31
CustomOp/linalg::matMulNBits<31, 31, 31>15         contrib.epu.quadric_custom_op  [-0.59268f, 1.41659f]                         30
/decoder/layers.1/Add                              add                            [-11.9694f, 1.4109f]                          27
/decoder/layers.1/encoder_attn_layer_norm/Add_1    nn.layer_norm                  [-31.863f, 18.7409f]                          26
CustomOp/linalg::matMulNBits<31, 31, 27>16         contrib.epu.quadric_custom_op  [-8.70827f, 11.8252f]                         27
CustomOp/linalg::matMulNBits<31, 31, 31>2          contrib.epu.quadric_custom_op  [-6.54716f, 5.91384f]                         28
CustomOp/linalg::matMulNBits<31, 31, 31>3          contrib.epu.quadric_custom_op  [-4.21574f, 3.62095f]                         28
CustomOp/nn::whisperAttention<384, 6, 1500, 14>36  contrib.epu.quadric_custom_op  [-1.40372f, 2.0292f]                          29
CustomOp/linalg::matMulNBits<31, 31, 30>17         contrib.epu.quadric_custom_op  [-1.2742f, 0.783488f]                         30
/decoder/layers.1/Add_1                            add                            [-11.7797f, 1.31481f]                         27
/decoder/layers.1/final_layer_norm/Add_1           nn.layer_norm                  [-34.356f, 17.1086f]                          25
CustomOp/linalg::matMulNBits<31, 31, 27>18         contrib.epu.quadric_custom_op  [-18.9478f, 71.5836f]                         24
/decoder/layers.1/activation_fn/Mul_1              contrib.epu.gelu               [-0.169971f, 71.5836f]                        24
CustomOp/linalg::matMulNBits<31, 31, 29>19         contrib.epu.quadric_custom_op  [-113.528f, 8.97997f]                         24
/decoder/layers.1/Add_2                            add                            [-115.08f, 10.1623f]                          24
/decoder/layers.2/self_attn_layer_norm/Add_1       nn.layer_norm                  [-9.23459f, 15.3414f]                         26
CustomOp/linalg::matMulNBits<31, 31, 28>20         contrib.epu.quadric_custom_op  [-5.45088f, 5.13565f]                         28
CustomOp/nn::whisperAttention<384, 6, 4, 14>37     contrib.epu.quadric_custom_op  [-0.822134f, 0.739178f]                       31
CustomOp/linalg::matMulNBits<31, 31, 29>21         contrib.epu.quadric_custom_op  [-0.471779f, 3.29612f]                        29
/decoder/layers.2/Add                              add                            [-115.007f, 9.87098f]                         24
/decoder/layers.2/encoder_attn_layer_norm/Add_1    nn.layer_norm                  [-32.0424f, 18.1859f]                         25
CustomOp/linalg::matMulNBits<31, 31, 28>22         contrib.epu.quadric_custom_op  [-10.3365f, 8.57053f]                         27
CustomOp/linalg::matMulNBits<31, 31, 31>4          contrib.epu.quadric_custom_op  [-7.2206f, 6.83287f]                          28
CustomOp/linalg::matMulNBits<31, 31, 31>5          contrib.epu.quadric_custom_op  [-4.22371f, 4.76144f]                         28
CustomOp/nn::whisperAttention<384, 6, 1500, 14>38  contrib.epu.quadric_custom_op  [-1.2737f, 1.75648f]                          30
CustomOp/linalg::matMulNBits<31, 31, 30>23         contrib.epu.quadric_custom_op  [-1.6217f, 1.06462f]                          30
/decoder/layers.2/Add_1                            add                            [-115.65f, 10.1965f]                          24
/decoder/layers.2/final_layer_norm/Add_1           nn.layer_norm                  [-27.5322f, 13.5401f]                         26
CustomOp/linalg::matMulNBits<31, 31, 28>24         contrib.epu.quadric_custom_op  [-7.36932f, 2.67375f]                         28
/decoder/layers.2/activation_fn/Mul_1              contrib.epu.gelu               [-0.169971f, 2.66372f]                        29
CustomOp/linalg::matMulNBits<31, 31, 29>25         contrib.epu.quadric_custom_op  [-3.60221f, 1.07261f]                         29
/decoder/layers.2/Add_2                            add                            [-115.773f, 9.67873f]                         24
/decoder/layers.3/self_attn_layer_norm/Add_1       nn.layer_norm                  [-15.5193f, 13.6514f]                         27
CustomOp/linalg::matMulNBits<31, 31, 28>26         contrib.epu.quadric_custom_op  [-7.85931f, 6.792f]                           28
CustomOp/nn::whisperAttention<384, 6, 4, 14>39     contrib.epu.quadric_custom_op  [-2.41125f, 1.44966f]                         29
CustomOp/linalg::matMulNBits<31, 31, 27>27         contrib.epu.quadric_custom_op  [-2.09421f, 3.09023f]                         29
/decoder/layers.3/Add                              add                            [-115.747f, 9.95996f]                         24
/decoder/layers.3/encoder_attn_layer_norm/Add_1    nn.layer_norm                  [-26.1427f, 9.04464f]                         26
CustomOp/linalg::matMulNBits<31, 31, 28>28         contrib.epu.quadric_custom_op  [-8.85646f, 9.43103f]                         27
CustomOp/linalg::matMulNBits<31, 31, 31>6          contrib.epu.quadric_custom_op  [-8.39829f, 7.86805f]                         27
CustomOp/linalg::matMulNBits<31, 31, 31>7          contrib.epu.quadric_custom_op  [-5.39744f, 6.1728f]                          28
CustomOp/nn::whisperAttention<384, 6, 1500, 14>40  contrib.epu.quadric_custom_op  [-2.70359f, 3.22433f]                         29
CustomOp/linalg::matMulNBits<31, 31, 27>29         contrib.epu.quadric_custom_op  [-4.77284f, 1.9294f]                          28
/decoder/layers.3/Add_1                            add                            [-115.479f, 10.4301f]                         24
/decoder/layers.3/final_layer_norm/Add_1           nn.layer_norm                  [-46.1546f, 11.7791f]                         25
CustomOp/linalg::matMulNBits<31, 31, 28>30         contrib.epu.quadric_custom_op  [-17.636f, 17.31f]                            26
/decoder/layers.3/activation_fn/Mul_1              contrib.epu.gelu               [-0.169971f, 17.31f]                          26
CustomOp/linalg::matMulNBits<31, 31, 29>31         contrib.epu.quadric_custom_op  [-8.89668f, 11.8703f]                         27
/decoder/layers.3/Add_2                            add                            [-110.711f, 8.84814f]                         24
/decoder/layer_norm/Add_1                          nn.layer_norm                  [-238.585f, 80.7449f]                         23
CustomOp/linalg::matMulNBits<31, 31, 31>32         contrib.epu.quadric_custom_op  (-64.0, 63.99999997019768)                    25
                                                   contrib.epu.quadric_custom_op  (-2147483648, 2147483647)                      0

2026-09-22 03:27 - INFO - epu - codegen - START====================build_cpu_runnable_fx_relay
2026-09-22 03:27 - INFO - epu - codegen - START=======================quantize_to_chimera_fx
2026-09-22 03:27 - INFO - epu - codegen - START=================================relay_to_tir
2026-09-22 03:27 - INFO - epu - codegen - START===========================relay_to_epu_relay
2026-09-22 03:27 - INFO - epu - codegen - START==============================adapt_and_order
2026-09-22 03:27 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 03:27 - INFO - epu - codegen - START=============================plan_lrm_virtual
2026-09-22 03:27 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 03:27 - INFO - epu - codegen - START===============================lrm_alloc_loop
2026-09-22 03:28 - INFO - epu - codegen - START==============================amend_ctrl_flow
2026-09-22 03:28 - INFO - epu - codegen - START================================lrm_splitting
2026-09-22 03:29 - INFO - epu - codegen - START==============================ext_split_relay
2026-09-22 03:29 - INFO - epu - codegen - START====================================build_tir
2026-09-22 03:31 - INFO - epu - iss_testing - No tranges found for input, use default float range: <tvm.contrib.epu.interval.Interval object at 0x7d83171d97e0>
2026-09-22 03:31 - INFO - epu - iss_testing - Found tranges for input: <tvm.contrib.epu.interval.Interval object at 0x7d83171d8d00>
2026-09-22 03:31 - INFO - epu - iss_testing - Started Running Graph on Chimera ISS...
FILM 73/73: 100%|███████████████████████████████████████████████████| 73/73 [02:03<00:00,  1.70s/it]
2026-09-22 03:33 - INFO - epu - iss_testing - Done 0:02:04.145357


  GPNPU tokens: [50259, 50359, 50363, 2221]
  ORT   tokens: [50259, 50359, 50363, 2221]
  Token ids match ORT: True

  Decoded: <|en|><|transcribe|><|notimestamps|> Mr

7. GPNPU Performance

The ISS writes one profile_core*.json per core into each module's build directory; these read core 0. The chart is the SDK's own breakdown as an image, which the docs catalog uses as this notebook's card.

Hardware config QC-P, 32 MB OCM, 16 MACs/PE, 128 GBps, 1.7 GHz, 4 cores.

StageCyclesms @ 1.7 GHz
Encoder — waveform → hidden states43.96 M25.86
whisper::logMel7.40 M4.35
whisper::convStem7.35 M4.33
Decoder — hidden states → token ids4.28 M2.52
End to end48.24 M28.38
for module, stage in (
    ("whisper_e2e_chipy", "Encoder: waveform -> hidden states"),
    ("whisper_dec_chipy", "Decoder: hidden states -> token ids"),
):
    profile = sorted(glob.glob(f"ccl_build/{module}/**/profile_core0.json", recursive=True))[0]
    counts = chimera_core._plot_profile_results(profile, clock_freq=hw_config.clock_freq_ghz * 1e9)
    cycles = {k: v / 1e6 for k, v in counts.items() if k not in ("total", "ExtBytes")}
    figure, axis = plt.subplots(figsize=(8, 3.2))
    axis.barh(list(cycles), list(cycles.values()), color="#3d6098")
    axis.invert_yaxis()
    axis.set_xlabel("cycles (millions)")
    axis.set_title(f"{stage} — {counts['total'] / 1e6:,.2f}M cycles")
    for row, value in enumerate(cycles.values()):
        axis.text(value, row, f" {value:,.1f}M", va="center", fontsize=9)
    figure.tight_layout()
    plt.show()
[SDK-CLI] : TotalCycles: 43,925,705
[SDK-CLI] : Executions/second: 39

compute      : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 19.679M
data_array   : ▇▇▇▇▇▇▇▇ 3.478M
mac          : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 6.154M
data_ocm     : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 11.834M
data_external: ▇▇▇▇▇ 2.302M

[SDK-CLI] : TotalCycles: 4,297,756
[SDK-CLI] : Executions/second: 396

compute      : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 1.907M
data_array   : ▇▇▇▇▇▇▇▇▇▇▇▇ 487.555K
mac          : ▇▇▇▇▇▇▇▇▇ 380.016K
data_ocm     : ▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇▇ 1.145M
data_external: ▇▇▇▇▇▇▇▇▇ 372.422K


Summary

ModelWhisper-Tiny (39 M parameters, 4 encoder + 4 decoder layers, embed_dim 384)
On the GPNPUlog-mel, conv front-end, encoder, decoder prefill, greedy argmax
On the hostFLAC decode only
Programs2 — waveform → hidden states, (ids, hidden) → tokens; stitched with ChiPy
TargetQC-P, 32 MB OCM, 16 MACs/PE, 128 GBps, 4 cores
QuantizationINT4 weights via MatMulNBits, FP16 attention
Custom Opslinalg::matMulNBits, nn::whisperAttention, whisper::convStem, whisper::logMel, whisper::greedyTokens — all in whisper_stubs.hpp
Accuracy (observed)~54.5 dB PSNR on the encoder hidden states; token ids exactly matching ONNX Runtime

Citation

@article{radford2022whisper,
  title   = {Robust Speech Recognition via Large-Scale Weak Supervision},
  author  = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg
             and McLeavey, Christine and Sutskever, Ilya},
  journal = {arXiv preprint arXiv:2212.04356},
  year    = {2022}
}

Sign in to your account

Don't have an account? 
By signing in, you are agreeing to our Terms of Use and Privacy Policy.
Quadric // One architecture. Every algorithm.

Develop.

Simulate.

Profile.

Collaborate.