Product docs

Transcription Parameters by Engine

Map tunable settings across Whisper file, Whisper realtime, realtime models, and advanced ASR—with scene presets and plan gates.

Transcription Parameters by Engine

Advanced

Advanced transcription tuning screenshot

Screenshot

What this page is for

Advanced settings differ for Whisper file, Whisper realtime, realtime models, and advanced ASR (ML Engine tab) routes. Mixing them up causes “docs mention a knob I cannot see.”

This is the hub page—read Advanced Parameter Transcription for tuning workflow first.

Plans (summary)

CapabilityFreeStandardPro
Advanced file parametersNoYesYes
Realtime presetsYesYes
Realtime custom paramsNoYes
Advanced ASR modelsYes

Whisper file transcription

When: offline batch workflows (files, links, watch folders) with a Whisper model.

Typical groups (visible with Standard+ advanced settings):

GroupWhy tune
Scene presetsOne-click VAD/decoding bias (table below)
VAD / no-speech thresholdsFewer hallucinations, cleaner cuts
DecodingGreedy vs beam, best-of, context
Segments & offsetsPartial files, length limits
Parallel & separationPro features such as diarization (separate toggles)

Whisper file scene presets

PresetBest forWhat moves
GeneralMost mediaBalanced VAD/decoding
DialogueInterviews, two-speakerShorter segments
SpeechMonologueLonger segments, more context
MeetingMulti-speakerMeeting-oriented VAD
CourseLecturesLong segments, large context
NoisyHigh background noiseStronger VAD/denoise bias, beam
MediaMusic + dialogueConservative no-speech, light denoise
CustomTeam templateSaved manual baseline

Whisper realtime

When: mic/app/global realtime with Whisper as the engine.

Not the same table as file mode. Common areas:

  • Step/window sizing (latency vs stability)
  • Segmentation mode (when lines appear)
  • Silero / energy VAD thresholds
  • Thread count

Stabilize RTF before chasing perfect wording. GPU can help but is not mandatory everywhere.

Realtime models

When: realtime workflows with models from the Realtime models group (e.g. Zipformer).

Focus on VAD and chunking, not Whisper beam search:

  • VAD scene presets
  • Min speech / silence durations
  • Min / max segment length
  • Padding, merge gaps
  • Threads

Do not apply Whisper file “no-speech / beam” advice here.

Advanced ASR (ML Engine tab)

When: Pro file jobs with a model from ML Engine.

  • Languages and limits vary by model—see Advanced ASR models.
  • Advanced fields live on a separate form—do not merge with Whisper VAD tables.
  • Offline file focus today; realtime-model knobs do not apply.

If the UI disagrees with this page

Ship cadence can add or hide controls. Trust the app UI and send feedback. Engineering uses whisper-form.factory visibility and CLI mappers as source of truth (marketing-content-audit.md in the repo).

FAQ

Can Standard customize realtime params?
Presets yes; custom realtime configuration requires Pro.

Do file presets affect link jobs?
Whisper file parameters apply to supported offline batch entry points.

Where is GPU?
GPU is an engine/runtime choice—see GPU Transcription.

Next steps

Whisper-Powered Live Transcription: Capture Speech from Mic, Apps & Media Files in Real Time

Contact us

Email
Copyright © 2026. Made by AudioNote, All rights reserved.