{"observation":{"id":"4962515f-4e04-46aa-808f-8dbe06374930","tool":"assemblyai","tool_name":"AssemblyAI","criterion":"output-quality","criterion_name":"Output quality","criterion_definition":"How accurately the returned transcript matches the human reference transcript, measured by WER.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)","scenario":"medical-anatomy-narration-with-dense-jargon","scenario_name":"Medical anatomy narration with dense jargon","group_tag":"speech-to-text-benchmark","scenario_description":"A long narrated excerpt from Henry Gray's Anatomy of the Human Body containing dense medical terminology and accented articulation. It was used to test lexical accuracy on domain-specific jargon and spelling of technical terms.","modality":"audio","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1","role":"input","filename":"medical_terms.mp3"}],"stresses":["domain-specific vocabulary recognition","medical term spelling accuracy","accented speech robustness","phoneme-to-grapheme precision","long-form audio handling"],"verdict":"worked","score":null,"score_total":null,"note":"Accuracy is strong on dense medical jargon: 3.78% WER with 71 substitutions, 14 deletions, and 18 insertions against a 2,728-word reference, and the run reports 100.0% jargon recall.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/4d74aea939f046b18bde5cd3b0764ce8.png?v=1","role":"input","alt":"4d74aea939f046b18bde5cd3b0764ce8.png"},{"url":"https://cdn.futuresmart.ai/public/aidemos/ff643967431a4804a67f8ce7f97a63b6.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/a93a20d74b4b4029bf96241aa5209ebb.png?v=1","role":"output","alt":null}],"run_id":"469de0c2-d727-4f8f-a60e-e3a5bf8e8588","study_title":"Transcribe Audio Accurately — Speech-to-Text Engine Benchmark","study_kind":"generation","research_task":"86baxegpu","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/071409ab3a0d4358bb238133eafcb60b.mp3?v=1","filename":"medical_terms.mp3","alt":"Medical anatomy narration with dense jargon","role":"input"}],"modality":"audio","stresses":["domain-specific vocabulary recognition","medical term spelling accuracy","accented speech robustness","phoneme-to-grapheme precision","long-form audio handling"]},"tool_page_slug":"assemblyai-speech-to-text","tool_url":"https://aidemos.com/tools/assemblyai-speech-to-text","permalink":"https://aidemos.com/evidence/4962515f-4e04-46aa-808f-8dbe06374930","api_url":"https://ai.aidemos.com/v1/observations/4962515f-4e04-46aa-808f-8dbe06374930"},"peers":[{"id":"00b55e91-870a-4430-8136-f7b81a548fba","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"worked","score":3.63,"score_total":null,"note":"Transcript quality is strong at 3.63% WER, with 75 substitutions, 12 deletions, and 12 insertions against 2728 reference words.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/902712fc2af843c2915508d758ab51c5.png?v=1","evidence_url":"https://aidemos.com/evidence/00b55e91-870a-4430-8136-f7b81a548fba"},{"id":"9e0d991b-d39f-4f0f-b7c0-ca643271064e","tool":"deepgram","tool_name":"Deepgram","verdict":"worked","score":5.43,"score_total":null,"note":"Keeps lexical accuracy high on dense medical narration: WER is 5.43% with 82 substitutions, 34 deletions, and 32 insertions against a 2728-word reference, and jargon recall is 100.0% (9/9 scored terms).","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1","evidence_url":"https://aidemos.com/evidence/9e0d991b-d39f-4f0f-b7c0-ca643271064e"},{"id":"d36c4860-c2e1-4ad4-8fdd-067e59f0cb8a","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"worked","score":3.01,"score_total":null,"note":"Transcript accuracy is strong on dense medical narration, with WER 3.01% and 51 substitutions, 8 deletions, and 23 insertions over a 2,728-word reference.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/799fe3740fed4fad962f4d443f4e44d2.png?v=1","evidence_url":"https://aidemos.com/evidence/d36c4860-c2e1-4ad4-8fdd-067e59f0cb8a"},{"id":"011f908c-afe9-4008-afe2-7f838349110d","tool":"gladia","tool_name":"Gladia","verdict":"worked","score":4.07,"score_total":null,"note":"On dense medical jargon, WER was 4.07% with 73 substitutions, 14 deletions, and 24 insertions against a 2,728-word reference; jargon recall was 100.0% (9/9 scored terms).","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1","evidence_url":"https://aidemos.com/evidence/011f908c-afe9-4008-afe2-7f838349110d"},{"id":"491d0a47-7871-458c-8beb-234875efd947","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"mixed","score":13.09,"score_total":null,"note":"On dense medical narration, the transcript reaches 13.09% WER with 228 substitutions, 49 deletions, and 80 insertions against 2728 reference words, and it recalls 100.0% of scored jargon terms.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/16302165d88c4986b79191a743eb0478.png?v=1","evidence_url":"https://aidemos.com/evidence/491d0a47-7871-458c-8beb-234875efd947"},{"id":"65a93d95-5cb0-41c2-96df-b2947129dba6","tool":"groqcloud","tool_name":"GroqCloud","verdict":"worked","score":3.15,"score_total":null,"note":"Delivers high lexical accuracy on dense medical narration: WER 3.15% with 58 substitutions, 13 deletions, and 15 insertions over 2728 reference words, returning 2730 hypothesis words.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/a466cf3e61c947b48d0be7aeef96433e.png?v=1","evidence_url":"https://aidemos.com/evidence/65a93d95-5cb0-41c2-96df-b2947129dba6"},{"id":"e4685302-d458-418f-bc6b-c0f6975013e6","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"worked","score":5.17,"score_total":null,"note":"Achieves low transcription error on dense jargon: WER 5.17% with 54 substitutions, 76 deletions, and 11 insertions against 2728 reference words; it returned 2663 words and missed only 'trabeculae'.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/01bceb9ce0d844619415fb9d1d128a10.png?v=1","evidence_url":"https://aidemos.com/evidence/e4685302-d458-418f-bc6b-c0f6975013e6"},{"id":"c4a05049-079d-45b2-bccd-9db5f8f83936","tool":"rev-ai","tool_name":"Rev AI","verdict":"worked","score":9.79,"score_total":null,"note":"On the dense medical narration, it held WER to 9.79% against a 2728-word reference, with 195 substitutions, 7 deletions and 65 insertions; jargon recall was 77.8% (7/9 scored terms), so it handled the terminology reasonably well but still missed specific terms such as cancellous and trabeculae.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/53a62dd8311a4f0c9ff89e699bc07877.mp3?v=1","evidence_url":"https://aidemos.com/evidence/c4a05049-079d-45b2-bccd-9db5f8f83936"},{"id":"42b6555a-0e00-49af-9f8f-91b173a74095","tool":"speechmatics","tool_name":"Speechmatics","verdict":"worked","score":3.01,"score_total":null,"note":"The transcript is highly accurate on dense medical narration, with 3.01% WER (49 substitutions, 17 deletions, 16 insertions over 2728 reference words) and 100.0% jargon recall.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/845fb0bf5bbd4d9e94e2645cdd415692.png?v=1","evidence_url":"https://aidemos.com/evidence/42b6555a-0e00-49af-9f8f-91b173a74095"}],"other_criteria":[{"id":"9c7fb0a9-ef0e-4b34-b866-e8084b7cf4fe","criterion":"export","criterion_name":"Export","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 5,437 timed tokens and JSON depth 5 across 5,443 objects, with payload depth 3/3.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/9c7fb0a9-ef0e-4b34-b866-e8084b7cf4fe"}],"appears_in":[{"page_type":"ranking","slug":"speech-to-text-apis","title":"Best AI Tools for Accurate Speech-to-Text on Hard Audio","url":"https://aidemos.com/best/speech-to-text-apis","binding":"run"}],"same_scenario":[{"id":"00b55e91-870a-4430-8136-f7b81a548fba","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"worked","score":3.63,"score_total":null,"note":"Transcript quality is strong at 3.63% WER, with 75 substitutions, 12 deletions, and 12 insertions against 2728 reference words."},{"id":"9e0d991b-d39f-4f0f-b7c0-ca643271064e","tool":"deepgram","tool_name":"Deepgram","verdict":"worked","score":5.43,"score_total":null,"note":"Keeps lexical accuracy high on dense medical narration: WER is 5.43% with 82 substitutions, 34 deletions, and 32 insertions against a 2728-word reference, and jargon recall is 100.0% (9/9 scored terms)."},{"id":"d36c4860-c2e1-4ad4-8fdd-067e59f0cb8a","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"worked","score":3.01,"score_total":null,"note":"Transcript accuracy is strong on dense medical narration, with WER 3.01% and 51 substitutions, 8 deletions, and 23 insertions over a 2,728-word reference."},{"id":"011f908c-afe9-4008-afe2-7f838349110d","tool":"gladia","tool_name":"Gladia","verdict":"worked","score":4.07,"score_total":null,"note":"On dense medical jargon, WER was 4.07% with 73 substitutions, 14 deletions, and 24 insertions against a 2,728-word reference; jargon recall was 100.0% (9/9 scored terms)."},{"id":"491d0a47-7871-458c-8beb-234875efd947","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"mixed","score":13.09,"score_total":null,"note":"On dense medical narration, the transcript reaches 13.09% WER with 228 substitutions, 49 deletions, and 80 insertions against 2728 reference words, and it recalls 100.0% of scored jargon terms."},{"id":"65a93d95-5cb0-41c2-96df-b2947129dba6","tool":"groqcloud","tool_name":"GroqCloud","verdict":"worked","score":3.15,"score_total":null,"note":"Delivers high lexical accuracy on dense medical narration: WER 3.15% with 58 substitutions, 13 deletions, and 15 insertions over 2728 reference words, returning 2730 hypothesis words."},{"id":"e4685302-d458-418f-bc6b-c0f6975013e6","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"worked","score":5.17,"score_total":null,"note":"Achieves low transcription error on dense jargon: WER 5.17% with 54 substitutions, 76 deletions, and 11 insertions against 2728 reference words; it returned 2663 words and missed only 'trabeculae'."},{"id":"c4a05049-079d-45b2-bccd-9db5f8f83936","tool":"rev-ai","tool_name":"Rev AI","verdict":"worked","score":9.79,"score_total":null,"note":"On the dense medical narration, it held WER to 9.79% against a 2728-word reference, with 195 substitutions, 7 deletions and 65 insertions; jargon recall was 77.8% (7/9 scored terms), so it handled the terminology reasonably well but still missed specific terms such as cancellous and trabeculae."},{"id":"42b6555a-0e00-49af-9f8f-91b173a74095","tool":"speechmatics","tool_name":"Speechmatics","verdict":"worked","score":3.01,"score_total":null,"note":"The transcript is highly accurate on dense medical narration, with 3.01% WER (49 substitutions, 17 deletions, 16 insertions over 2728 reference words) and 100.0% jargon recall."}]}