Product docs

Advanced ASR Models

Pro-only speech models under Settings → Models → ML Engine—languages, use cases, and how they compare to Whisper and realtime models.

Advanced ASR Models

Settings

Transcription settings overview screenshot

Screenshot

What this page is for

Advanced ASR models are Pro-only offline file routes in the ML Engine group in the model picker (e.g. Qwen3-ASR, Cohere, Parakeet). They sit beside Whisper and realtime models, but they target offline files—not low-latency live capture.

This page covers when to use each model.

Before you start

  • Mode: focused on offline file transcription today, not realtime-model mic latency.
  • Plan: Free and Standard cannot select these models; Pro (including Lifetime) gets the full ML Engine list.
  • Download: first use requires downloading the model in Settings; sizes are often larger than Tiny/Base Whisper.

Model overview

ModelGood forLanguagesNotes
Qwen3-ASRMeetings, courses, Chinese-heavy mediaChinese, English, Japanese, Korean, and moreMultilingual with punctuation
Cohere Transcribe (multilingual)International interviews, mixed sourcesMany widely used languagesHigh-quality offline multilingual
ParakeetEnglish/European bulk workEnglish plus European languagesRelatively compact for English-first libraries

See the in-app model card for the authoritative language list.

Whisper vs realtime vs advanced ASR

NeedLean toward
Live subtitles, low latencyRealtime models (or GPU Whisper realtime)
GGML ecosystem, links, watch foldersWhisper
Pro offline files, specific multilingual qualityAdvanced ASR (ML Engine tab)

See Models & Engines Overview and Concept.

Parameters

Advanced ASR scenes use their own advanced fields—not Whisper VAD tables. See parameters by engine.

FAQ

Does transcription need the internet?
Inference is local; downloads and licensing may use the network.

Can I run advanced ASR and GPU Whisper together?
Pick one model per job; GPU packages primarily serve Whisper routes.

After upgrading to Pro?
Past transcripts stay as-is; new jobs can pick advanced ASR models under ML Engine.

Next steps

Whisper-Powered Live Transcription: Capture Speech from Mic, Apps & Media Files in Real Time

Contact us

Email
Copyright © 2026. Made by AudioNote, All rights reserved.