Method and sources
One model. One Mac. Direct runtime comparison.
The model, source recording, and Mac stay fixed. Only the local runtime implementation changes. This is a runtime comparison, not an app workflow comparison.
Selected run and reproducibility
October 1, 2026 additional run. Runner 4000ea7 with shutdown timeout locally increased from 5 to 60 seconds. Installed Wasper displayed 1.8.0 and was built locally. The operator confirms the local change was README-only, with no runtime code changes. Native server SHA-256: bdd71e87e1bc032c0b056205a9c55c8e2ca964b76826cfa77525627f3bb9b076. Machine, RAM, macOS marketing version and AC power are reported by the operator; the structured evidence records arm64 and Darwin 27.0.0.
Hardware, operating system and Wasper builds differ between runs. These results do not isolate the effect of macOS, and displayed app versions do not establish identical binaries.
All 5,544 expected outcomes were saved: 5,430 successes and 114 long-recording errors. Every runtime completed 729/729 short requests. Missing detailed failure logs prevent attributing the long errors to a particular cause. Raw transcripts are not published.
FLEURS short recordings
Google FLEURS test set
supplied 243 prompted recordings: 27 each in German, Greek, English, Spanish, French,
Italian, Dutch, Polish, and Portuguese.
Long original recordings
21 complete talks, speeches, readings, and addresses. Each source stays with its original publisher.
Comparison boundary
Each runtime received the same complete source recording directly on one Mac. Every row runs Parakeet TDT v3 through its local runtime implementation. The Wasper result measures the native Parakeet server, not Wasper.app.
Timing and memory scope
Response-only timing after one discarded warm-up request. Physical footprint was sampled after each response. These samples are not peak memory or total system memory requirements; no sortable memory ranking is published.
How Local MLX INT8 was made Reproducible group-64 derivative of MLX Community’s public model.
This is a local derivative, not a Wasper download. Start from
MLX Community’s FP32 model,
then create and save a separate eight-bit, group-64 copy. The original model remains unchanged.
- Install
parakeet-mlx==0.5.2 and mlx==0.32.2. - Copy the source
config.json and tokenizer files into a new model folder. - Add
"quantization": {"bits": 8, "group_size": 64} to the copied configuration.
model = parakeet_mlx.from_pretrained(source, dtype=mx.float32)
nn.quantize(model, bits=8, group_size=64)
mx.eval(model.parameters())
model.save_weights("local-int8/model.safetensors")
To match this benchmark, load the copied configuration before the saved weights, use BF16
activations, and process long audio in 120-second windows with 15 seconds of overlap.
- Machine
- MacBook Air · Apple M1 · 8 GiB unified memory · macOS 27.0.1 on AC power
- Recordings
- 9 languages · 243 short recordings · 21 long recordings · 3 passes
- Language
- No language hint or external language-identification service
Long-audio policy
Every runtime received the complete source recording. MLX Community and Local MLX INT8 handle
long recordings internally as 120-second windows with 15-second overlap; that is their
runtime-owned behavior, not benchmark-created segmentation.
Runtime sources
- Wasper Metal INT8 osa911/wasper-parakeet-tdt-0.6b-v3-onnx-int8 · v6 · Wasper 1.8.0
- MLX Community F32/BF16 mlx-community/parakeet-tdt-0.6b-v3 · parakeet-mlx 0.5.2 / MLX 0.32.2
- Local MLX INT8 Local group-64 INT8 derivative · based on mlx-community/parakeet-tdt-0.6b-v3 · parakeet-mlx 0.5.2 / MLX 0.32.2
- Handy Q8 handy-computer/parakeet-tdt-0.6b-v3-gguf Q8_0 · transcribe.cpp 0.2.3 / 63a44d9239
- NVIDIA Q8 nvidia/parakeet-tdt-0.6b-v3 Q8_0 GGUF · NeMo-Speech.cpp 0.1.0 / 4f96762
- Istupakov ONNX INT8 istupakov/parakeet-tdt-0.6b-v3-onnx · onnx-asr 0.12.0 / ONNX Runtime 1.30.0
- Fluid Core ML FluidInference/parakeet-tdt-0.6b-v3-coreml · FluidAudio / 69e42dae8e