{"observation":{"id":"818d5ab9-cd17-4b1b-b244-887a4102298d","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","criterion":"export","criterion_name":"Export","criterion_definition":"How complete the returned transcript payload is, as reflected in the depth or richness of what the tool outputs.","criterion_evidence_type":"transformation","criterion_rank_role":"context","criterion_rank_role_reason":"Transcript payload richness is useful for comparison, but it does not measure transcription correctness against the reference. (3 of 3 judges)","scenario":"overlapping-meeting-speech-with-cross-talk","scenario_name":"Overlapping meeting speech with cross-talk","group_tag":"speech-to-text-benchmark","scenario_description":"A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.","modality":"audio","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1","role":"input","filename":"crosstalk.wav"}],"stresses":["overlapping speech","background noise robustness","multi-speaker separation","speaker diarization accuracy","long-form audio handling"],"verdict":"mixed","score":null,"score_total":null,"note":"Returns a mid-depth transcript payload: payload depth 2/3 with 5009 word-level timed tokens, confidence present, and no speaker labels.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/0e057a02d9d64d09b7aa20a2d5e4da1e.png?v=1","role":"output","alt":null},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-raw-response-2586aea6fce2.json","role":"output","alt":null}],"run_id":"469de0c2-d727-4f8f-a60e-e3a5bf8e8588","study_title":"Transcribe Audio Accurately — Speech-to-Text Engine Benchmark","study_kind":"generation","research_task":"86baxegpu","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/12990982a1614cc5b18a3a450e0e657f.wav?v=1","filename":"crosstalk.wav","alt":"Overlapping meeting speech with cross-talk","role":"input"}],"modality":"audio","stresses":["overlapping speech","background noise robustness","multi-speaker separation","speaker diarization accuracy","long-form audio handling"]},"tool_page_slug":"google-cloud-speech-to-text","tool_url":"https://aidemos.com/tools/google-cloud-speech-to-text","permalink":"https://aidemos.com/evidence/818d5ab9-cd17-4b1b-b244-887a4102298d","api_url":"https://ai.aidemos.com/v1/observations/818d5ab9-cd17-4b1b-b244-887a4102298d"},"peers":[{"id":"28eb1a61-9d3a-41d6-bc14-f2807721c53b","tool":"assemblyai","tool_name":"AssemblyAI","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 11,743 timed tokens and JSON depth 5 across 11,749 objects, with payload depth 3/3.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/33522c2a9c0d4f29b49147f68f855f65.png?v=1","evidence_url":"https://aidemos.com/evidence/28eb1a61-9d3a-41d6-bc14-f2807721c53b"},{"id":"4962545f-f4ad-47ca-8668-c44f159ada99","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"Returns a full developer payload with payload depth 3/3, 5912 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 19457 objects.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1","evidence_url":"https://aidemos.com/evidence/4962545f-f4ad-47ca-8668-c44f159ada99"},{"id":"466df8cc-f4d1-4e62-9f07-0d1fc6fcb815","tool":"deepgram","tool_name":"Deepgram","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload on hard crosstalk audio, with word-level timing, confidence, speaker labels, 13,555 timed tokens, 4 distinct speakers, payload depth 3/3, and JSON depth 11 across 14,065 objects.","artifact_count":4,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1","evidence_url":"https://aidemos.com/evidence/466df8cc-f4d1-4e62-9f07-0d1fc6fcb815"},{"id":"a9c3d354-c6ec-4468-ba77-6df4de2e7c41","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"The returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 14,506 timed tokens.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/e44c577a896a41ed81b79e2986cb72f3.png?v=1","evidence_url":"https://aidemos.com/evidence/a9c3d354-c6ec-4468-ba77-6df4de2e7c41"},{"id":"87244b43-3466-42eb-bce3-5c163985466f","tool":"gladia","tool_name":"Gladia","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 12,968 timed tokens, and 7-level JSON nesting across 12,976 objects.","artifact_count":4,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1","evidence_url":"https://aidemos.com/evidence/87244b43-3466-42eb-bce3-5c163985466f"},{"id":"db8677ed-d684-4483-8779-d921258ca5b2","tool":"groqcloud","tool_name":"GroqCloud","verdict":"failed","score":null,"score_total":null,"note":"Returns no transcript payload at all on the oversized upload; the only returned content is an error object, so export richness is effectively zero.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/a1e0fc3436944f299d84aaff8f30b14d.png?v=1","evidence_url":"https://aidemos.com/evidence/db8677ed-d684-4483-8779-d921258ca5b2"},{"id":"d8c37195-9af6-49c9-b834-cac2d58088dc","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"failed","score":null,"score_total":null,"note":"Returns no transcript payload at all on this input; the only output is a 413 error body, so there is nothing transcript-like to export.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/3093c5df674d41179041b5987186ba6d.png?v=1","evidence_url":"https://aidemos.com/evidence/d8c37195-9af6-49c9-b834-cac2d58088dc"},{"id":"3331a18c-72cf-4c07-9898-60b6ff8de924","tool":"rev-ai","tool_name":"Rev AI","verdict":"worked","score":null,"score_total":null,"note":"Returned a rich developer payload with word-level timing, confidence and speaker labels; the raw response walk reports 6249 timed tokens, payload depth 3/3, and JSON depth 5 across 14484 objects.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/24b7f1fb3d354789b95a764f0a5ef329.wav?v=1","evidence_url":"https://aidemos.com/evidence/3331a18c-72cf-4c07-9898-60b6ff8de924"},{"id":"9909b52b-84f0-40bb-b464-d0465e3bd1c4","tool":"speechmatics","tool_name":"Speechmatics","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload with payload_depth 3/3, including word-level timing, confidence values, and speaker labels in the JSON response.","artifact_count":4,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/770b31cffe56485dac23dbf4206bdcdd.png?v=1","evidence_url":"https://aidemos.com/evidence/9909b52b-84f0-40bb-b464-d0465e3bd1c4"}],"other_criteria":[{"id":"30c8040c-4369-44fa-a1f9-2c9315d276f5","criterion":"automation-level","criterion_name":"Automation level","rank_role":"context","verdict":"mixed","score":null,"score_total":null,"note":"The recorded workflow submits the audio as a POST to the v2 endpoint and reaches a scored result without operator input, but the trace explicitly says per-call timings were not instrumented, so no measured call count is claimed.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/30c8040c-4369-44fa-a1f9-2c9315d276f5"},{"id":"a04e7fb4-6535-45db-83e4-d36ad4ec0258","criterion":"output-quality","criterion_name":"Output quality","rank_role":"decisive","verdict":"struggled","score":43.5,"score_total":null,"note":"On overlapping crosstalk, the transcript quality is poor at 43.50% WER, with 633 substitutions, 2614 deletions, and 50 insertions against 7579 reference words, yielding 5015 hypothesis words.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/a04e7fb4-6535-45db-83e4-d36ad4ec0258"}],"appears_in":[{"page_type":"ranking","slug":"speech-to-text-apis","title":"Best AI Tools for Accurate Speech-to-Text on Hard Audio","url":"https://aidemos.com/best/speech-to-text-apis","binding":"run"}],"same_scenario":[{"id":"28eb1a61-9d3a-41d6-bc14-f2807721c53b","tool":"assemblyai","tool_name":"AssemblyAI","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 11,743 timed tokens and JSON depth 5 across 11,749 objects, with payload depth 3/3."},{"id":"4962545f-f4ad-47ca-8668-c44f159ada99","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"Returns a full developer payload with payload depth 3/3, 5912 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 19457 objects."},{"id":"466df8cc-f4d1-4e62-9f07-0d1fc6fcb815","tool":"deepgram","tool_name":"Deepgram","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload on hard crosstalk audio, with word-level timing, confidence, speaker labels, 13,555 timed tokens, 4 distinct speakers, payload depth 3/3, and JSON depth 11 across 14,065 objects."},{"id":"a9c3d354-c6ec-4468-ba77-6df4de2e7c41","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"The returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 14,506 timed tokens."},{"id":"87244b43-3466-42eb-bce3-5c163985466f","tool":"gladia","tool_name":"Gladia","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 12,968 timed tokens, and 7-level JSON nesting across 12,976 objects."},{"id":"db8677ed-d684-4483-8779-d921258ca5b2","tool":"groqcloud","tool_name":"GroqCloud","verdict":"failed","score":null,"score_total":null,"note":"Returns no transcript payload at all on the oversized upload; the only returned content is an error object, so export richness is effectively zero."},{"id":"d8c37195-9af6-49c9-b834-cac2d58088dc","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"failed","score":null,"score_total":null,"note":"Returns no transcript payload at all on this input; the only output is a 413 error body, so there is nothing transcript-like to export."},{"id":"3331a18c-72cf-4c07-9898-60b6ff8de924","tool":"rev-ai","tool_name":"Rev AI","verdict":"worked","score":null,"score_total":null,"note":"Returned a rich developer payload with word-level timing, confidence and speaker labels; the raw response walk reports 6249 timed tokens, payload depth 3/3, and JSON depth 5 across 14484 objects."},{"id":"9909b52b-84f0-40bb-b464-d0465e3bd1c4","tool":"speechmatics","tool_name":"Speechmatics","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload with payload_depth 3/3, including word-level timing, confidence values, and speaker labels in the JSON response."}]}