Leaderboard · September 2026

Appen MedTerm-90 Benchmark

Seven commercial speech-to-text engines, scored on 90 hours of real physician dictation. A term counts only when the written form is one a clinician would accept: formatting differences count, misspellings do not.

Rankings

Medical terminology recall

30,017 clinician-validated entities on identical 1,413 files. A term counts only in a written form a clinician would accept.

90.5 h
of real physician dictation, 70% of it captured over telephone lines
30,017
clinical entities checked against every transcript, across 14 specialties
16.6 pp
recall gap between first place and last, on identical audio
190
misrenderings rated potentially fatal in the clinical risk review
#ModelProviderRecall95% CIVs medianSpeed
1 ElevenLabs Scribe v2 ElevenLabs 83.35% 82.60–84.09 +2.70 20.1×
2 OpenAI gpt-transcribe OpenAI 81.79% 80.69–82.77 +1.14 27.2×
3 Microsoft MAI-Transcribe 2 Microsoft 81.31% 80.53–82.09 +0.66 94.2×
4 Google Gemini 3.5 Transcribe Google 80.65% 79.48–81.77 ±0.00 44.2×
5 Meta Muse Transcribe Meta 77.76% 76.81–78.68 −2.89 6.5×
6 Amazon Transcribe Medical Amazon 76.70% 75.87–77.52 −3.95 8.8×
7 xAI Grok STT 1.0 xAI 66.74% 65.71–67.77 −13.91 102.4×

ElevenLabs Scribe v2 leads at 83.35%. OpenAI gpt-transcribe (81.79%) and Microsoft MAI-Transcribe 2 (81.31%) are a statistical tie for second with different characters: gpt-transcribe spells medicine cleanly, while MAI-2 pairs near-top recall with 94× real-time speed. Google's Gemini 3.5 debuts fourth at 80.65%, within the statistical margin of third, takes two categories outright, and needed the fewest corrections in the field. Muse and Amazon fill the middle, and Grok pairs the field's highest speed with its lowest recall.

Analysis

Recall by clinical category

◆ = best in field · strict scoring. Every model has a weak spot, and no two share the same one.

CategoryScribe v2gpt-transcribeMAI-2Gemini 3.5MuseAmazon Med.Grok 1.0
Diagnosis83.582.282.283.7 ◆79.778.769.7
Medication77.976.578.076.470.283.0 ◆52.5
Vital sign75.078.6 ◆75.476.068.451.154.8
Lab value65.975.769.978.7 ◆72.066.352.8
Procedure85.8 ◆82.680.583.479.274.666.6
Symptom90.0 ◆88.487.085.383.184.780.0
Anatomy84.1 ◆80.781.675.378.173.966.9

Scribe v2 leads three of seven categories (symptom 90.0%, procedure 85.7%, anatomy 84.0%). Gemini 3.5 takes diagnosis (83.7%) and lab values (78.7%). Amazon owns medications at 83.0% while dropping to 51.1% on vital signs, and gpt-transcribe takes vital signs (78.6%). Grok misses roughly one medication in two. Lab values are the weakest category field-wide, because the judge does not forgive unit and decimal garbles.

Analysis

Recall by clinical specialty

Strict micro recall (%) · specialties with ≥20 files · ◆ = best per specialty.

SpecialtyScribe v2gpt-transcribeMAI-2Gemini 3.5MuseAmazon Med.Grok 1.0
Medical-legal / IME n=50785.885.884.286.0 ◆81.380.672.4
Neurology n=19188.9 ◆88.185.487.784.376.767.1
Surgery n=18986.587.7 ◆83.883.281.979.167.9
Orthopedics n=17865.8 ◆52.062.450.452.158.743.0
Psychiatry n=9981.482.8 ◆80.981.179.980.668.8
Pulmonology n=9182.485.0 ◆81.882.277.977.468.8
Post-acute care n=5188.989.2 ◆88.387.585.682.875.2
Cardiology n=2785.287.6 ◆86.286.084.379.572.3
Urology n=2387.589.5 ◆85.286.985.681.670.5

The ranking moves with the specialty. gpt-transcribe takes psychiatry (82.8%), and Gemini 3.5 edges the medical-legal lead (86.0%). Scribe v2 leads neurology (88.9%) and orthopedics (65.8%). Orthopedics remains the hardest major specialty for every model, with a 23-point spread down to Grok's 43.0%.

Analysis

Speed vs. accuracy

X = audio seconds processed per API second (log scale) · Y = strict recall · top-right is best.

Grok is fastest at 102× and least accurate. MAI-Transcribe 2 holds the speed-accuracy frontier: 94× real time at 81.28%. Gemini is second fastest at 44.7×. Muse is slowest (6.5×, 28.9 s median), and Amazon's 8.8× includes its asynchronous upload, queue and poll pipeline, which is the latency production users actually experience. Per-file latency (p50 / p95): Scribe v2: 10.8s / 21.4s · gpt-transcribe: 6.6s / 18.3s · MAI-2: 2.1s / 4.4s · Gemini 3.5: 4.0s / 12.2s · Muse: 28.9s / 86.5s · Amazon Med.: 21.4s / 53.4s · Grok 1.0: 1.9s / 4.3s.

Insights

Where every model struggles

Clinical safety risk

On 28.0% of files no model reached 85% recall, and 192 files defeated all seven. The safety chart counts files where a model missed over half the medications or vital signs: Scribe v2 (20.4%) and MAI-2 (20.7%) fail that way least, and Grok does on more than half its files.

Clinical risk

When a miss is not just a miss

A separate risk review classified all 6,425 near-miss renderings by clinical consequence. It does not change any score. It shows what the recall numbers are made of.

Recall percentages treat every error the same. Clinically they are not the same. A dropped decimal turns Eliquis 2.5 mg into a tenfold anticoagulant overdose. A one-syllable slip turns Toradol, an NSAID, into tramadol, an opioid. These misrenderings look plausible on the page, they survive casual review, and several were written by five different models on the same file. Our risk judge classified every near-miss pair: 191 were rated potentially fatal and 1,027 serious. Every engine in the field produced some.

Each pair was rated on a five-level scale. Fatal: acting on the transcript could directly cause life-threatening harm, such as a wrong drug class, an order-of-magnitude dose, or a tripled frequency. Serious: significant clinical harm or a major treatment error is plausible, such as a similar but wrong drug or a lost critical order. Moderate: could mislead care but would likely be caught in context. Minor: a cosmetic deviation a reader recovers instantly. None: clinically equivalent. Across all 6,425 pairs: 191 fatal, 1,027 serious, 1,566 moderate, 2,027 minor, 1,614 none.

ModelFatal-risk renderingsSerious-risk renderings
Microsoft MAI-Transcribe 219181
ElevenLabs Scribe v226216
Meta Muse Transcribe28218
OpenAI gpt-transcribe31155
Google Gemini 3.5 Transcribe42174
Amazon Transcribe Medical43235
xAI Grok STT 1.067310

The drug-name garble gallery

516 of the fatal and serious renderings are outright drug-name garbles: one medication becoming another, or becoming a non-word. 444 of those were made by a single model; no other engine made the same mistake on the same audio. A sample of those single-model errors:

Reference (scribe note)What the model wroteOnly model to write itRisk
Cardene dripcodeine dripMAI-2FATAL
Mucinex 600 mg one twice daily for seven daysamoxicillin 600 milligrams 1 twice daily for 7 daysScribe v2FATAL
diclofenac 75 mg to take twice a day with foodxanax 75 milligrams to take twice a day with foodgpt-transcribeFATAL
Toradol 30 mgtramadol 30 mgAmazonFATAL
tramadoltoradolMuseFATAL
Humaloghema logGemini 3.5FATAL
hydrochlorothiazidehydrochlorideGrokFATAL
Vantin 200 mg twice a daydivantin 200 twice a dayMAI-2FATAL
clindamycin antibiotic solutionampicillin antibiotic solutionScribe v2FATAL
Ceftin 250 mg twice daily for 10 daysseptra 250 milligrams twice daily for 10 daysgpt-transcribeFATAL
QVAR 80 mcg two puffs twice a dayquavar 82 puffs twice a day andAmazonFATAL
doxycycline 100 mg two tablets for 10 daysoxycodone 100 milligrams 2 tablets for 10 daysMuseFATAL
calcium chloride 20 mEq q. daypotassium chloride 20 meq dailyGemini 3.5FATAL
FlorineffluorineGrokFATAL
0.5% Marcaine with epinephrinemarkane with epinephrine theMAI-2SERIOUS
Zoloft 50 mg one and a half tablets everydayzolpidem 50 milligrams 1 and a half tablets every dayScribe v2FATAL
Sinemet 25/100 mg one tablet dailysenna 25 100 1 tablet dailygpt-transcribeFATAL
lacosamide 100 mg b.i.d.glucosamide 100 mg b.i.d.AmazonFATAL

Sound-alike drug pairs are a known hazard in human transcription, and each model has its own: a cardiac drip becomes an opioid, an expectorant becomes an antibiotic, an NSAID becomes Xanax. No other engine made the same mistake on the same audio, which makes these model-specific confusions rather than audio problems. Each swap is a different pharmacology delivered to the chart in confident prose.

Methodology

How MedTerm-90 is built and scored

Real data, clinician-validated ground truth, and a verified scoring pipeline checked for vendor neutrality. Every layer is applied identically to every model.

The corpus: real data, zero contamination

1,413 scored files (of 1,424 collected): real physician dictations with paired scribe-authored SOAP notes. All of it is dictation audio: a single physician speaking close to the microphone, no crosstalk, no far-field capture. That makes the conditions favorable, and the remaining errors harder to excuse. 90.5 hours, median 3 min 05 s, 14 specialties led by medical-legal/IME (36%), neurology, surgery, orthopedics and psychiatry. The audio is dominated by telephony: 1,003 files (70%) are GSM 6.10-compressed WAV from telephone dictation lines, which carry codec-limited, telephony-band speech. The rest splits between 235 MP3 recordings (44.1 kHz mono, 320 kbps) and 186 uncompressed PCM WAV files (16 kHz). Estimated SNR by the frame-energy percentile method: PCM median 24 dB, MP3 33 dB, GSM 44 dB. The high GSM figure reflects codec noise-squelching, not better speech. Every recording is real data from real clinical environments. Nothing is synthetic, scripted or studio-read. The corpus is part of Appen's licensed dataset library, and only audio never sold to or acquired by any of the companies tested was used, so no model could have seen this data in training.

How each API was called

Every model ran through its standard production interface with vendor-recommended defaults. Runs were checkpointed per file, failures logged and retried, and latency measured as the full API round-trip. Expand a model for the exact request and response shape:

ElevenLabs Scribe v2

Request (per file):

POST https://api.elevenlabs.io/v1/speech-to-text
xi-api-key: ****
multipart/form-data:
  model_id = "scribe_v2"       language_code = "en"
  diarize = "true"             timestamps_granularity = "word"
  tag_audio_events = "false"   file = <raw audio bytes>

Response shape:

{"text": "... blood pressure one twenty-six over sixty-two ...", "words": [...], "language_code": "eng"}
OpenAI gpt-transcribe

Request (per file):

POST https://api.openai.com/v1/audio/transcriptions
Authorization: Bearer ****
multipart/form-data:
  model = "gpt-transcribe"
  file  = <raw audio bytes>      (language auto-detected)

Response shape:

{"text": "... medication like metformin, omeprazole, ... ...", "languages": ["en"], "usage": {...}}
Microsoft MAI-Transcribe 2

Request (per file):

POST {AZURE_ENDPOINT}/speechtotext/transcriptions:transcribe?api-version=2025-10-15
Ocp-Apim-Subscription-Key: ****
multipart/form-data:
  definition = {"enhancedMode": {"enabled": true, "model": "mai-transcribe-2",
                "transcribeStyle": "verbatim"}, "locales": ["en-US"]}
  audio = <raw audio bytes>

Response shape:

{"combinedPhrases": [{"text": "... on medication metformin, omeprazole, Invokana, Unglyza, glipizide ..."}]}
Google Gemini 3.5 Transcribe

Request (per file):

google-genai SDK -> models.generate_content
model = "gemini-3.5-transcribe"
contents = [<audio bytes inline>, "Transcribe this medical dictation verbatim."]
(one persistently empty response file excluded from the common subset for all
 models; see limitations)

Response shape:

{"candidates": [{"content": {"parts": [{"text": "... dictating a monthly follow-up ..."}]}}]}
Meta Muse Transcribe

Request (per file):

POST https://api.meta.ai/v1/asr/transcribe?sessionId=<uuid>
Authorization: Bearer ****
multipart/form-data:
  request = {"mode": "DIARIZATION", "model": "muse-voice-transcribe-1.0",
             "audioEncoding": "WAV"}
  audio = <16 kHz mono PCM WAV bytes>   (files > 580 s chunked)

Response shape:

{"text": "... blood pressure 126 over 62 respirations 18 pulse 74 ...", "segments": [...]}
Amazon Transcribe Medical

Request (per file):

boto3 (us-east-1): upload to S3, then
StartMedicalTranscriptionJob(
  LanguageCode="en-US", Specialty="PRIMARYCARE", Type="DICTATION",
  MediaFormat="mp3"|"wav", Media={"MediaFileUri": "s3://<bucket>/<key>"})
poll GetMedicalTranscriptionJob every 5 s -> fetch result JSON from S3

Response shape:

{"results": {"transcripts": [{"transcript": "... on medication metformin, omeprazole, Invokana. Onglyza ..."}]}}
xAI Grok STT 1.0

Request (per file):

POST https://api.x.ai/v1/stt
Authorization: Bearer ****
multipart/form-data:
  format = "true"    language = "en"
  file = <raw audio bytes>    (NO keyterm parameters; GSM WAV pre-converted to PCM)

Response shape:

{"text": "... blood pressure 126 over 62, respirations 18, pulse 74. ...", "words": [...]}

Scoring methodology

Reference entities were extracted from the scribe notes at temperature 0 and reviewed by a cohort of medical experts, who also spot-checked the final outputs and scores of every model. To score a transcript, entity and transcript first pass through the same clinical normalization: casing and punctuation are unified and spoken numbers canonicalized, so that “126/62”, “126 over 62” and “one twenty-six over sixty-two” are the same blood pressure. Matching then runs exact, fuzzy (indel ≥ 80), and bidirectional abbreviation, in that order.

Character similarity finds near-misses, but it cannot tell a formatting difference from a misspelling: “para vertebral” and “Unglyza” look equally close to their references. So every one of the 12,326 distinct near-miss pairs was reviewed the way a clinician would read it (Gemini 3.8 Flash, temperature 0, fixed policy). A rendering is credited when the written form is one clinicians actually use: spacing and hyphenation differences, recognized alternate spellings, standard abbreviations, equivalent number formats, minor inflections. This is what protects models from being punished for style — a model that writes “b.i.d.” where the note says “twice daily”, or hyphenates a compound term, loses nothing. What gets rejected is substance: misspellings and non-words (“Unglyza” for the drug “Onglyza”), different drugs or structures however similar, and garbles a clinician would flag. 21.2% of pairs were accepted; rejected credits count as misses. Canary checks passed: misspelled brand names rejected, spacing-only variants accepted.

Neutrality is verified rather than assumed. In decoy testing, identical clinical entities from unrelated files produced statistically indistinguishable false-match rates across models (7.6–8.4%), so no model gains an advantage from verbosity or output style. Every model faced identical files, entities, thresholds and judges, with no vocabulary hints, and every recall figure carries a file-level bootstrap 95% CI (2,000 resamples).

Worked example: one dictation, all seven models

Pulmonology follow-up dictated over a telephone line (GSM 6.10 WAV, 21 reference entities; internal id …163350). The scribe note lists the patient's medications, including “gabapentin 300 mg at bedtime … Incruse inhaler daily.” Each model's rendering of that passage, verbatim:

ModelRendering of the medication passage
Scribe v2“Gabapentin three hundred milligrams at bedtime, Qvar inhaler daily, Vent”
gpt-transcribe“gabapentin 300 milligrams at bedtime, Incruse inhaler daily, Ventolin PR”
MAI-2“gabapentin 300 milligrams at bedtime, Incruse inhaler daily, Ventolin, P”
Gemini 3.5“gabapentin 300 milligrams at bedtime Incruse Ellipta daily Ventolin PRN”
Muse“gabapentin 300 milligrams at bedtime incruise inhaler daily, Ventolin PR”
Amazon Med.“gabapentin 300 mg at bedtime, Incruse inhaler daily, Ventolin prn and pr”
Grok 1.0“gabapentin 300 milligrams at bedtime, in-cruz inhaler daily, ventolin PR”

Scoring of eight representative entities. ✓ credited · ✓f credited via accepted fuzzy match · ✗j near-miss rejected by the judge · ✗ missed:

Reference entityCategoryScribe v2gpt-transcribeMAI-2Gemini 3.5MuseAmazon Med.Grok 1.0
advanced COPDdiagnosis
gabapentin 300 mg at bedtimemedication
blood pressure 122/84vital sign
heart rate 84vital sign✗j✗j✗j✗j✗j
Incruse inhaler dailymedication✗j✗j
oxygen saturation is 98.4% on room airvital sign✗j✗j✗j✗j✗j
Nasopharynxanatomy✗j✗j
oropharynxanatomy

One inhaler, four fates: Amazon, MAI-2 and gpt-transcribe write “Incruse” and get credit. Grok writes “in-cruz” and Muse writes “incruise”, and the judge rejects both as non-words (✗j). Scribe v2 writes “Qvar”, a different inhaler entirely, and misses. Gemini writes “Incruse Ellipta daily”, which no longer matches the reference phrase and scores as a miss. “Oropharynx” defeats all seven on this recording. Full per-file results and raw API responses are available on request.

Limitations, read before quoting

About Appen

Appen builds expert data for the frontier of AI: the specialist, hard-to-reach data that generic web corpora cannot supply. The audio behind MedTerm-90 is a case in point. These are real dictations by real practicing physicians, captured in live clinical transcription workflows and licensed for AI development through Appen's data partnerships. Speech like this cannot be scripted, synthesized, or scraped. That access runs deep: licensed real-world archives across regulated domains, a global network of vetted specialists (clinicians among them) for expert annotation and validation, and coverage across 500+ locales and acoustic conditions. The same machinery powering this benchmark is available to model builders: expert domain corpora from Appen's data library, custom collection with specialist ground truth, and evaluation pipelines like this one run on your own models, your own data, your own metrics.