Parakeet TDT v3 runtime benchmark · M1 Pro · 16 GB · macOS 14.7.3

One model.
Seven runtimes.

The model and recordings stay fixed. The runtime implementation changes. Compare the official and community Parakeet TDT v3 runtimes on one Mac, using direct native requests instead of Wasper.app.

Runtimes
7
Languages
9
Sequential passes
3

Short dictation

Performance map

7 complete runtimes of one model on one Mac. Lower word error rate is better. Higher processing speed is better. Focus or hover a point for exact values.

  • Wasper
  • MLX Community
  • Local MLX
  • Handy
  • NVIDIA
  • Istupakov
  • Fluid
Processing speed (x real time)20x real time means 1 hour of audio is processed in about 3 minutes.Word error rate (%)

Choose a workload

Compare quick recordings and long audio.

One model, one Mac, two recording lengths.Choose a workload to update the map and full comparison together.

Runtime results

Same model. Same Mac. Different runtimes.

Each row runs the same model through its own local runtime.Wasper uses its Metal server and Axiom's custom Metal kernels.

Test Mac
Showing M1 Pro · 16 GB · macOS 14.7.3 · Short recordings
Short recordings · complete, comparable results
Artifact
Wasper Metal INT8Wasper · Wasper 1.8.0
osa911/wasper-parakeet-tdt-0.6b-v3-onnx-int8 · v6INT8 encoder9.33% WER82.8×115.2 ms
MLX Community F32/BF16MLX Community · parakeet-mlx 0.5.2 / MLX 0.32.2
mlx-community/parakeet-tdt-0.6b-v3F32 weights, BF16 runtime8.36% WER28.8×342.5 ms
Local MLX INT8Wasper · parakeet-mlx 0.5.2 / MLX 0.32.2
Local group-64 INT8 derivativeBased on MLX CommunityINT8 weights, group 648.40% WER47.2×211.4 ms
Handy Q8Handy Computer · transcribe.cpp 0.2.3 / 63a44d9239
handy-computer/parakeet-tdt-0.6b-v3-gguf Q8_0Q8_08.63% WER52.3×190.0 ms
NVIDIA Q8NVIDIA · NeMo-Speech.cpp 0.1.0 / 4f96762
nvidia/parakeet-tdt-0.6b-v3 Q8_0 GGUFQ8_08.59% WER47.1×212.0 ms
Istupakov ONNX INT8Istupakov · onnx-asr 0.12.0 / ONNX Runtime 1.30.0
istupakov/parakeet-tdt-0.6b-v3-onnxINT89.82% WER33.9×293.5 ms
Fluid Core MLFluidInference · FluidAudio / 69e42dae8e
FluidInference/parakeet-tdt-0.6b-v3-coremlMixed precision8.59% WER12.8×775.1 ms

Method and sources

One model. One Mac. Direct runtime comparison.

The model, source recording, and Mac stay fixed. Only the local runtime implementation changes. This is a runtime comparison, not an app workflow comparison.

Selected run and reproducibility

Published September 2026 baseline. Wasper displayed version 1.8.0 but was a locally rebuilt development binary, not necessarily the public 1.8.0 release. The frozen snapshot does not include its exact binary hash.

Hardware, operating system and Wasper builds differ between runs. These results do not isolate the effect of macOS, and displayed app versions do not establish identical binaries.

FLEURS short recordings

Google FLEURS test set supplied 243 prompted recordings: 27 each in German, Greek, English, Spanish, French, Italian, Dutch, Polish, and Portuguese.

Comparison boundary

Each runtime received the same complete source recording directly on one Mac. Every row runs Parakeet TDT v3 through its local runtime implementation. The Wasper result measures the native Parakeet server, not Wasper.app.

Timing and memory scope

Response-only timing after one discarded warm-up request. The frozen public aggregate does not preserve matching short- and long-recording memory cohorts, so this comparison does not publish sortable physical-memory results.

How Local MLX INT8 was made Reproducible group-64 derivative of MLX Community’s public model.

This is a local derivative, not a Wasper download. Start from MLX Community’s FP32 model, then create and save a separate eight-bit, group-64 copy. The original model remains unchanged.

  1. Install parakeet-mlx==0.5.2 and mlx==0.32.2.
  2. Copy the source config.json and tokenizer files into a new model folder.
  3. Add "quantization": {"bits": 8, "group_size": 64} to the copied configuration.
model = parakeet_mlx.from_pretrained(source, dtype=mx.float32)
nn.quantize(model, bits=8, group_size=64)
mx.eval(model.parameters())
model.save_weights("local-int8/model.safetensors")

To match this benchmark, load the copied configuration before the saved weights, use BF16 activations, and process long audio in 120-second windows with 15 seconds of overlap.

Machine
MacBook Pro · Apple M1 Pro · 16 GiB unified memory · macOS 14.7.3 on AC power
Recordings
9 languages · 243 short recordings · 21 long recordings · 3 passes
Language
No language hint or external language-identification service

Long-audio policy

Every runtime received the complete source recording. MLX Community and Local MLX INT8 handle long recordings internally as 120-second windows with 15-second overlap; that is their runtime-owned behavior, not benchmark-created segmentation.

Runtime sources

  • Wasper Metal INT8 osa911/wasper-parakeet-tdt-0.6b-v3-onnx-int8 · v6 · Wasper 1.8.0
  • MLX Community F32/BF16 mlx-community/parakeet-tdt-0.6b-v3 · parakeet-mlx 0.5.2 / MLX 0.32.2
  • Local MLX INT8 Local group-64 INT8 derivative · based on mlx-community/parakeet-tdt-0.6b-v3 · parakeet-mlx 0.5.2 / MLX 0.32.2
  • Handy Q8 handy-computer/parakeet-tdt-0.6b-v3-gguf Q8_0 · transcribe.cpp 0.2.3 / 63a44d9239
  • NVIDIA Q8 nvidia/parakeet-tdt-0.6b-v3 Q8_0 GGUF · NeMo-Speech.cpp 0.1.0 / 4f96762
  • Istupakov ONNX INT8 istupakov/parakeet-tdt-0.6b-v3-onnx · onnx-asr 0.12.0 / ONNX Runtime 1.30.0
  • Fluid Core ML FluidInference/parakeet-tdt-0.6b-v3-coreml · FluidAudio / 69e42dae8e

Evidence SHA-2565b229c5b24408f7d439fb92924c2bfdda5ccb7419e0d134198980ebfdf8fa819

Download both configurations and build notes (JSON) Run the public benchmark on your Mac

Verify the runtime comparison.

Each row links to the model bundle and runtime version used here. The Local MLX INT8 row includes the group-64 recipe for the locally derived runtime.