How Dictivo benchmarks local dictation on Mac
Listen to the input, read the output, and reproduce the published local-engine runs. Dictivo also uses a short calibration to help choose models for your Mac.
Short answer
The published runs use one Apple M4 Pro with 48 GB memory. A 25-second human-read English example includes the input, model outputs and run records. Separate five-second calibration results explain model selection. These are engine measurements, not full-app latency or a comparison with other dictation apps.
Hear a human voice. Inspect the local-engine result.
This vendor-run example uses 25.312 seconds of human-read English: five original VCTK recordings from one speaker, joined in order with 250 ms of silence between them. It is a reconstructed paragraph, not a continuous recording or spontaneous dictation. No speech was synthesized or cloned.
Measured 12 September 2026 with Dictivo 0.3.47, whisper.cpp 1.8.4, Balanced settings and automatic language detection. Machine: Apple M4 Pro, 48 GB, macOS 26.3.1 (a), Metal acceleration, running on battery. One warmup and three fresh CLI processes per model; operating-system file caches may be warm.
| Installed model | Three runs | Median | Normalized word errors |
|---|---|---|---|
| Small | 1.062 / 1.062 / 1.054 s | 1.062 s | 0 / 69 on this sample |
| Large v3 Turbo Q5 | 2.067 / 2.066 / 2.068 s | 2.067 s | 0 / 69 on this sample |
Both models produced the same words on all measured runs. Two reference commas were omitted. The word-error count ignores punctuation and capitalization; it does not mean an identical transcript or general 100% accuracy.
Actual output from both local models
Please call Stella. Ask her to bring these things with her from the store. Six spoons of fresh snow peas, five thick slabs of blue cheese and maybe a snack for her brother Bob. We also need a small plastic snake and a big toy frog for the kids. She can scoop these things into three red bags and we will go meet her Wednesday at the train station.
Compare with the reference text
Please call Stella. Ask her to bring these things with her from the store. Six spoons of fresh snow peas, five thick slabs of blue cheese, and maybe a snack for her brother Bob. We also need a small plastic snake and a big toy frog for the kids. She can scoop these things into three red bags, and we will go meet her Wednesday at the train station.
Read the reference · Small output · Turbo Q5 output
Timing includes process startup, model loading, transcription and writing the text file. It excludes microphone capture, app processing, insertion and proofreading, so it is not full-app dictation latency. This is one clean English sample, one speaker and one computer. Noise, accents, other languages, longer recordings and other hardware were not tested. Overlap with model training data is unknown. No competitor was tested.
Download input, every run, hashes and reproduction steps (2.5 MB ZIP) · Machine-readable results
Source: Junichi Yamagishi, Christophe Veaux and Kirsten MacDonald (2019), CSTR VCTK Corpus v0.92, University of Edinburgh. Recordings p225_001–005, resampled and joined as described above. CC BY 4.0. No endorsement by the speaker, authors or university is implied.
A separate check with the engine's network access denied
On the same machine, Small and Turbo Q5 each completed one additional run under a macOS sandbox profile that denied network access. Their output matched the unrestricted text files. A control connection succeeded outside the profile and failed with “Operation not permitted” inside it.
This checks the installed transcription CLI only. It is not a network capture or offline test of the complete desktop app, its microphone workflow, licensing, trial reports or other helper processes. The app can still use the network for those product operations. These functional checks are excluded from the timing table above.
Read the control and model results · Sandbox profile · Full-app network test procedure
What the benchmark measures
| Signal | How Dictivo uses it |
|---|---|
| Input | A bundled 5-second speech clip used for local calibration. |
| Metric | Real-time factor, or RTF. Lower is faster; below 1.0 means transcription finishes faster than the audio duration. |
| Hardware signal | CPU brand, system memory, and GPU names are used as the hardware fingerprint for cached results. |
| Output | Runnable Fast, Medium, and Quality local tiers, including model id, predicted or measured RTF, download state, and budget fit. |
Calibration steps
- Inspect the Mac hardware profile and create a fingerprint from CPU, memory, and GPU signals.
- Run the installed local model against Dictivo's bundled 5-second benchmark clip.
- Store the measured real-time factor against the current hardware fingerprint.
- Invalidate cached results if the hardware fingerprint changes.
- Map the measured profile to Fast, Medium, and Quality local tiers.
- Show Cloud Fast as a fallback when local performance or model download size is a poor fit.
Measured Whisper speed on Apple Silicon: Metal vs CPU
These are measured numbers, not predictions. Machine: Apple M4 Pro, 14-core CPU, 48 GB unified memory, macOS 26.3.1. Engine: the local Whisper engine bundled with Dictivo, using the 0.3.33 calibration update where Metal is benchmarked and active by default on Apple Silicon. The CPU column shows the same machine with GPU disabled, which is also how Dictivo versions before 0.3.33 ran.
Method: each cell is the median of 3 full runs after 1 warm-up, timed as complete wall-clock per dictation (process start, model load, and transcription of the bundled 5-second clip with default decode settings). These timings cover the calibration process; they do not include UI insertion or proofreading. Process startup and model loading add fixed costs, so multiplying a short-clip RTF by a longer recording duration is not a reliable latency measurement. Real-time factor (RTF) = processing time divided by audio duration; lower is faster.
| Model | Tier on this Mac | Metal RTF | CPU-only RTF | Metal speedup |
|---|---|---|---|---|
| Tiny | Free tier | 0.11 | 0.11 | 1.0x |
| Small | Fast | 0.11 | 0.21 | 1.9x |
| Large v3 Turbo Q5 | Medium | 0.21 | 0.61 | 2.9x |
| Large v3 | Quality | 0.41 | 0.81 | 2.0x |
- Tiny shows no GPU gain because process start and model load dominate its runtime.
- On this 5-second clip, Large v3 measured RTF 0.41 on Metal and 0.81 with the GPU disabled. These calibration results do not establish full-app responsiveness.
- One machine is published so far. Results vary with thermal state and background load; numbers for other Macs are added only after they are measured with this exact method.
Have a different Mac? Run Settings -> Engine -> Re-run setup in Dictivo and email the measured tier numbers to support@dictivo.app. Measured machines are added to this table with their macOS version and date.
Current local model tier logic
| Hardware capacity | Fast tier | Medium tier | Quality tier | Practical meaning |
|---|---|---|---|---|
| High local capacity | Small | Large v3 Turbo Q5 | Large v3 | Use larger local models when responsiveness and memory headroom are both acceptable. |
| Strong CPU profile | Base | Small | Large v3 Turbo Q5 | Keep everyday dictation responsive while still offering a higher-quality local option. |
| Constrained CPU profile | Tiny | Base | Small | Prefer small local models and use Cloud Fast when speed matters more than local-only processing. |
Model size and prediction ratios
| Model id | Display name | Approximate size | Prediction ratio | Role |
|---|---|---|---|---|
| tiny | Tiny | 75 MB | 0.2x | Starter model for constrained hardware. |
| base | Base | 142 MB | 0.4x | Quick feasibility checks and lightweight dictation. |
| small | Small | 469 MB | 0.7x | Default local model for resource-aware testing. |
| medium-q5_0 | Medium Q5 | 540 MB | 1.1x | CPU-friendly higher-accuracy local option. |
| large-v3-turbo-q5_0 | Large v3 Turbo Q5 | 600 MB | 1.5x | High-end balance of local speed and quality. |
| large-v3-turbo | Large v3 Turbo | 1.6 GB | 2.0x | Fast high-quality transcription on stronger hardware. |
| large-v3 | Large v3 | 3.1 GB | 2.5x | Highest-quality local transcription tier. |
Real-time factor is more useful than a generic benchmark score
A generic CPU score does not tell a dictation user whether a sentence will appear quickly enough after pressing the hotkey. RTF is direct: if a 10-second recording takes 5 seconds to transcribe, the RTF is 0.5. If it takes 20 seconds, the RTF is 2.0.
This is why Dictivo treats RTF as the operational metric for Local mode. It connects model choice to the actual dictation experience instead of to an abstract hardware ranking.
- Lower RTF is better for interactive dictation.
- Larger models can improve accuracy but increase download size, memory pressure, and processing time.
- The best local model is the largest model that still feels responsive on the user's Mac.
What this method proves, and what it does not prove
The calibration helps estimate which Local model tiers may fit a Mac. Model predictions are not measurements of every model or recording; the published tables identify the runs that were actually measured. Results do not establish a best Mac for every app, audio input or language.
Dictivo publishes hardware-specific numbers only for machines that were actually measured with the documented method. The table above covers an Apple M4 Pro; other Macs are added as they are measured, never predicted.
- Valid claim: Dictivo can calibrate local model fit on a specific Mac.
- Valid claim: Dictivo separates Local mode from optional Cloud Fast.
- Not claimed here: numbers for Mac models that have not been measured with this method yet.
How to use this when comparing dictation apps
When a dictation app says it runs locally, ask how it decides which local model is usable on the current machine. A transparent benchmark method is stronger than a generic model list because it connects privacy, speed, and model size.
Test a short message and a longer paragraph on your own Mac. Record the model, engine, audio duration, time until text is ready, and any corrections. Use the offline dictation guide to check the local and cloud options separately.
- Use the offline dictation guide for local-vs-cloud product comparisons.
- Use this benchmark method page for Dictivo's local model fit logic.
- Use the Mac model guide for a user-facing recommendation by Mac family and memory.
Benchmark questions
01 What is a good RTF for local dictation?
For interactive dictation, lower RTF is better. An RTF below 1.0 means transcription completes faster than the audio duration, but Dictivo may still recommend a smaller model when responsiveness matters more than maximum accuracy.
02 Does Dictivo publish M-series benchmark tables?
Yes, for measured machines only. The first table on this page covers an Apple M4 Pro (14-core, 48 GB) with Metal and CPU-only numbers per model. Other Macs are added once they are measured with the same method, not predicted.
03 How fast is Whisper Large v3 on an Apple M4 Pro?
On the published 5-second calibration clip, Large v3 measured RTF 0.41 and Large v3 Turbo Q5 measured RTF 0.21 with Metal. These results include process startup and model loading; a one-minute recording was not measured in this table.
04 Does Dictivo use the GPU on Apple Silicon?
Yes. Since the 0.3.33 engine update, calibration benchmarks both CPU and Metal and picks the faster path; on Apple Silicon, Metal is typically 2-3x faster end-to-end. Settings -> Engine shows which engine is active, and the app falls back to CPU automatically if the GPU path fails.
05 Why benchmark on the Mac instead of assuming a model?
Mac family, memory, background load, and local model size can change the real dictation experience. A local calibration result is more useful than assuming the same model is right for every Mac.
06 Does the benchmark audio leave the Mac?
No. Dictivo's local benchmark path runs against a bundled calibration clip on the device. Optional Cloud Fast is a separate mode for selected recordings.