{"observation":{"id":"3905d745-03b6-4af4-a42d-3ce36307541a","tool":"speechmatics","tool_name":"Speechmatics","criterion":"output-quality","criterion_name":"Output quality","criterion_definition":"How accurately the returned transcript matches the human reference transcript, measured by WER.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)","scenario":"overlapping-meeting-speech-with-cross-talk","scenario_name":"Overlapping meeting speech with cross-talk","group_tag":"speech-to-text-benchmark","scenario_description":"A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.","modality":"audio","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1","role":"input","filename":"crosstalk.wav"}],"stresses":["overlapping speech","background noise robustness","multi-speaker separation","speaker diarization accuracy","long-form audio handling"],"verdict":"struggled","score":26.63,"score_total":null,"note":"The transcript is only moderately accurate on overlapping speech, with 26.63% WER (507 substitutions, 1430 deletions, 81 insertions over 7579 reference words) and diarization over-segmenting the meeting into 5 speaker labels.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/770b31cffe56485dac23dbf4206bdcdd.png?v=1","role":"input","alt":"770b31cffe56485dac23dbf4206bdcdd.png"},{"url":"https://cdn.futuresmart.ai/public/aidemos/a94c9cce49c24738b51df89ea7f25e86.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/3a8757927c674f7bade5b63bc785df9f.png?v=1","role":"output","alt":null}],"run_id":"469de0c2-d727-4f8f-a60e-e3a5bf8e8588","study_title":"Transcribe Audio Accurately — Speech-to-Text Engine Benchmark","study_kind":"generation","research_task":"86baxegpu","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1","filename":"crosstalk.wav","alt":"Overlapping meeting speech with cross-talk","role":"input"}],"modality":"audio","stresses":["overlapping speech","background noise robustness","multi-speaker separation","speaker diarization accuracy","long-form audio handling"]},"tool_page_slug":"speechmatics","tool_url":"https://aidemos.com/tools/speechmatics","permalink":"https://aidemos.com/evidence/3905d745-03b6-4af4-a42d-3ce36307541a","api_url":"https://ai.aidemos.com/v1/observations/3905d745-03b6-4af4-a42d-3ce36307541a"},"peers":[{"id":"67f8ec61-1974-472b-9d73-24058a75af43","tool":"assemblyai","tool_name":"AssemblyAI","verdict":"struggled","score":null,"score_total":null,"note":"Accuracy is weak on overlapping speech: 33.16% WER with 461 substitutions, 1,976 deletions, and 76 insertions against a 7,579-word reference, ranking 4th of 8 on this input.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/33522c2a9c0d4f29b49147f68f855f65.png?v=1","evidence_url":"https://aidemos.com/evidence/67f8ec61-1974-472b-9d73-24058a75af43"},{"id":"8b0f2b7f-f263-4002-88ec-ccda7fa02a4e","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"struggled","score":33.88,"score_total":null,"note":"Transcript quality is weak at 33.88% WER, with 363 substitutions, 2162 deletions, and 43 insertions against 7579 reference words.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/7d1886b880644b5ab0ccd11ee2011055.png?v=1","evidence_url":"https://aidemos.com/evidence/8b0f2b7f-f263-4002-88ec-ccda7fa02a4e"},{"id":"fc8aff3b-8b30-419d-87cd-aac42ea7db8d","tool":"deepgram","tool_name":"Deepgram","verdict":"struggled","score":36.27,"score_total":null,"note":"On overlapping meeting crosstalk, the transcript quality is weak: WER is 36.27% with 1096 substitutions, 1143 deletions, and 510 insertions against a 7579-word reference, returning 6946 words overall.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1","evidence_url":"https://aidemos.com/evidence/fc8aff3b-8b30-419d-87cd-aac42ea7db8d"},{"id":"667df61a-e337-48ae-b65c-6d910e2ccf83","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"struggled","score":26.67,"score_total":null,"note":"Transcript accuracy is weak on crosstalk, with WER 26.67% and 856 substitutions, 782 deletions, and 383 insertions over a 7,579-word reference.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/afa9e0579a6b4d3687bd1e7729b54546.png?v=1","evidence_url":"https://aidemos.com/evidence/667df61a-e337-48ae-b65c-6d910e2ccf83"},{"id":"d50a3776-7637-40c4-a330-9fab63cfab76","tool":"gladia","tool_name":"Gladia","verdict":"struggled","score":37.35,"score_total":null,"note":"On overlapping crosstalk, it still scored the run but WER was 37.35% with 529 substitutions, 2,213 deletions, and 89 insertions against a 7,579-word reference, leaving 5,455 words in the hypothesis.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1","evidence_url":"https://aidemos.com/evidence/d50a3776-7637-40c4-a330-9fab63cfab76"},{"id":"a04e7fb4-6535-45db-83e4-d36ad4ec0258","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"struggled","score":43.5,"score_total":null,"note":"On overlapping crosstalk, the transcript quality is poor at 43.50% WER, with 633 substitutions, 2614 deletions, and 50 insertions against 7579 reference words, yielding 5015 hypothesis words.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/7b330832f55d4a9ca7a5b1fbaf9f2dc2.png?v=1","evidence_url":"https://aidemos.com/evidence/a04e7fb4-6535-45db-83e4-d36ad4ec0258"},{"id":"c48a4462-c6e1-48ae-8339-455ac16704ed","tool":"groqcloud","tool_name":"GroqCloud","verdict":"mixed","score":null,"score_total":null,"note":"Produces no transcript on the oversized file, so WER is unmeasured here; the report explicitly treats accuracy for this input as untested rather than poor.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/a1e0fc3436944f299d84aaff8f30b14d.png?v=1","evidence_url":"https://aidemos.com/evidence/c48a4462-c6e1-48ae-8339-455ac16704ed"},{"id":"5e2cd2f8-b93d-4732-9b67-643919f9c29f","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"Transcript accuracy is untested because no transcript was produced; the run stopped at the size rejection before any WER could be measured.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/3093c5df674d41179041b5987186ba6d.png?v=1","evidence_url":"https://aidemos.com/evidence/5e2cd2f8-b93d-4732-9b67-643919f9c29f"},{"id":"74b7d615-f56e-4e63-a724-6cf62c0413b9","tool":"rev-ai","tool_name":"Rev AI","verdict":"struggled","score":28.33,"score_total":null,"note":"On the overlapping meeting audio, it reached WER 28.33% against a 7579-word reference, with 601 substitutions, 1423 deletions and 123 insertions; it returned 6279 words, so accuracy degraded substantially under crosstalk.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1","evidence_url":"https://aidemos.com/evidence/74b7d615-f56e-4e63-a724-6cf62c0413b9"}],"other_criteria":[{"id":"9909b52b-84f0-40bb-b464-d0465e3bd1c4","criterion":"export","criterion_name":"Export","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload with payload_depth 3/3, including word-level timing, confidence values, and speaker labels in the JSON response.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/9909b52b-84f0-40bb-b464-d0465e3bd1c4"}],"appears_in":[{"page_type":"ranking","slug":"speech-to-text-apis","title":"Best AI Tools for Accurate Speech-to-Text on Hard Audio","url":"https://aidemos.com/best/speech-to-text-apis","binding":"run"}],"same_scenario":[{"id":"67f8ec61-1974-472b-9d73-24058a75af43","tool":"assemblyai","tool_name":"AssemblyAI","verdict":"struggled","score":null,"score_total":null,"note":"Accuracy is weak on overlapping speech: 33.16% WER with 461 substitutions, 1,976 deletions, and 76 insertions against a 7,579-word reference, ranking 4th of 8 on this input."},{"id":"8b0f2b7f-f263-4002-88ec-ccda7fa02a4e","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"struggled","score":33.88,"score_total":null,"note":"Transcript quality is weak at 33.88% WER, with 363 substitutions, 2162 deletions, and 43 insertions against 7579 reference words."},{"id":"fc8aff3b-8b30-419d-87cd-aac42ea7db8d","tool":"deepgram","tool_name":"Deepgram","verdict":"struggled","score":36.27,"score_total":null,"note":"On overlapping meeting crosstalk, the transcript quality is weak: WER is 36.27% with 1096 substitutions, 1143 deletions, and 510 insertions against a 7579-word reference, returning 6946 words overall."},{"id":"667df61a-e337-48ae-b65c-6d910e2ccf83","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"struggled","score":26.67,"score_total":null,"note":"Transcript accuracy is weak on crosstalk, with WER 26.67% and 856 substitutions, 782 deletions, and 383 insertions over a 7,579-word reference."},{"id":"d50a3776-7637-40c4-a330-9fab63cfab76","tool":"gladia","tool_name":"Gladia","verdict":"struggled","score":37.35,"score_total":null,"note":"On overlapping crosstalk, it still scored the run but WER was 37.35% with 529 substitutions, 2,213 deletions, and 89 insertions against a 7,579-word reference, leaving 5,455 words in the hypothesis."},{"id":"a04e7fb4-6535-45db-83e4-d36ad4ec0258","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"struggled","score":43.5,"score_total":null,"note":"On overlapping crosstalk, the transcript quality is poor at 43.50% WER, with 633 substitutions, 2614 deletions, and 50 insertions against 7579 reference words, yielding 5015 hypothesis words."},{"id":"c48a4462-c6e1-48ae-8339-455ac16704ed","tool":"groqcloud","tool_name":"GroqCloud","verdict":"mixed","score":null,"score_total":null,"note":"Produces no transcript on the oversized file, so WER is unmeasured here; the report explicitly treats accuracy for this input as untested rather than poor."},{"id":"5e2cd2f8-b93d-4732-9b67-643919f9c29f","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"Transcript accuracy is untested because no transcript was produced; the run stopped at the size rejection before any WER could be measured."},{"id":"74b7d615-f56e-4e63-a724-6cf62c0413b9","tool":"rev-ai","tool_name":"Rev AI","verdict":"struggled","score":28.33,"score_total":null,"note":"On the overlapping meeting audio, it reached WER 28.33% against a 7579-word reference, with 601 substitutions, 1423 deletions and 123 insertions; it returned 6279 words, so accuracy degraded substantially under crosstalk."}]}