{"observation":{"id":"e55bde63-c293-4d27-b989-b971ead4bbf7","tool":"speechmatics","tool_name":"Speechmatics","criterion":"export","criterion_name":"Export","criterion_definition":"How complete the returned transcript payload is, as reflected in the depth or richness of what the tool outputs.","criterion_evidence_type":"transformation","criterion_rank_role":"context","criterion_rank_role_reason":"Transcript payload richness is useful for comparison, but it does not measure transcription correctness against the reference. (3 of 3 judges)","scenario":"bilingual-spanish-english-code-switching-speech","scenario_name":"Bilingual Spanish-English code-switching speech","group_tag":"speech-to-text-benchmark","scenario_description":"A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.","modality":"audio","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1","role":"input","filename":"mix_language.mp3"}],"stresses":["code-switching detection","multilingual language ID","mid-sentence language transitions","word preservation across language flips","hallucination resistance in bilingual speech"],"verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload with payload_depth 3/3, including word-level timing, confidence values, and speaker labels in the JSON response.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/45e5abbf1e184a9890ebb7fb4e05d09f.png?v=1","role":"input","alt":"45e5abbf1e184a9890ebb7fb4e05d09f.png"},{"url":"https://cdn.futuresmart.ai/public/aidemos/40f84735bf7846e6876389b73c81fc19.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/f06f7d3281604687a1d7b3f0081a6564.png?v=1","role":"output","alt":null},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-raw-response-3-e83898e84ebd.json","role":"output","alt":null}],"run_id":"469de0c2-d727-4f8f-a60e-e3a5bf8e8588","study_title":"Transcribe Audio Accurately — Speech-to-Text Engine Benchmark","study_kind":"generation","research_task":"86baxegpu","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1","filename":"mix_language.mp3","alt":"Bilingual Spanish-English code-switching speech","role":"input"}],"modality":"audio","stresses":["code-switching detection","multilingual language ID","mid-sentence language transitions","word preservation across language flips","hallucination resistance in bilingual speech"]},"tool_page_slug":"speechmatics","tool_url":"https://aidemos.com/tools/speechmatics","permalink":"https://aidemos.com/evidence/e55bde63-c293-4d27-b989-b971ead4bbf7","api_url":"https://ai.aidemos.com/v1/observations/e55bde63-c293-4d27-b989-b971ead4bbf7"},"peers":[{"id":"4bd8fb2e-e1e0-475d-a6c6-e879538edcaf","tool":"assemblyai","tool_name":"AssemblyAI","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 12,181 timed tokens and JSON depth 5 across 12,187 objects, with payload depth 3/3.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/e319813d52304d89a559933409e837ec.png?v=1","evidence_url":"https://aidemos.com/evidence/4bd8fb2e-e1e0-475d-a6c6-e879538edcaf"},{"id":"00f8839f-963c-4043-8b0a-9f374a3daec5","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"Returns a full developer payload with payload depth 3/3, 6311 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 21039 objects.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1","evidence_url":"https://aidemos.com/evidence/00f8839f-963c-4043-8b0a-9f374a3daec5"},{"id":"14209d19-de86-48f5-95cb-1458098b1164","tool":"deepgram","tool_name":"Deepgram","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload on code-switching audio, with word-level timing, confidence, speaker labels, 11,684 timed tokens, 3 distinct speakers, payload depth 3/3, and JSON depth 11 across 11,962 objects.","artifact_count":4,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1","evidence_url":"https://aidemos.com/evidence/14209d19-de86-48f5-95cb-1458098b1164"},{"id":"6303e8f9-cdc4-49e6-ab0b-8f38c22ff8a3","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"The returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 12,162 timed tokens.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/d7196ed7e0a44b68b455e434e0c6de83.png?v=1","evidence_url":"https://aidemos.com/evidence/6303e8f9-cdc4-49e6-ab0b-8f38c22ff8a3"},{"id":"ac711648-1b75-4389-a5c6-a72604ffeab6","tool":"gladia","tool_name":"Gladia","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 26,166 timed tokens, and 7-level JSON nesting across 26,174 objects.","artifact_count":4,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1","evidence_url":"https://aidemos.com/evidence/ac711648-1b75-4389-a5c6-a72604ffeab6"},{"id":"a4a565c6-afc7-4855-8d13-360a1f8aa34e","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"Returns a mid-depth transcript payload: payload depth 2/3 with 3688 word-level timed tokens, confidence present, and no speaker labels.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/8373e06aed4d43d4a27f7e9351f178b4.png?v=1","evidence_url":"https://aidemos.com/evidence/a4a565c6-afc7-4855-8d13-360a1f8aa34e"},{"id":"a8f4bff9-5f10-484a-81f7-54f78fa7e92b","tool":"groqcloud","tool_name":"GroqCloud","verdict":"worked","score":null,"score_total":null,"note":"Returns the same verbose JSON shape on code-switching speech, including task, language, duration, segments, word timestamps, confidence data, and no speaker labels; the payload depth is 2/3 with 877 timed tokens.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/08d9beea3d2349288f4ad5d0f3776231.png?v=1","evidence_url":"https://aidemos.com/evidence/a8f4bff9-5f10-484a-81f7-54f78fa7e92b"},{"id":"69e1f4f4-c76a-49d2-b48e-787e6b416be6","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"Returns 5703 word-level timed tokens, but no confidence values or speaker labels; the JSON is shallow at depth 3 with 5705 objects.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/7a3dbe25ec7541579bb3d1e651b20ffe.png?v=1","evidence_url":"https://aidemos.com/evidence/69e1f4f4-c76a-49d2-b48e-787e6b416be6"},{"id":"b8d04964-1003-4de2-b766-52060b761904","tool":"rev-ai","tool_name":"Rev AI","verdict":"worked","score":null,"score_total":null,"note":"Returned a full developer payload with word-level timing, confidence and speaker labels; the raw response walk reports 5855 timed tokens, payload depth 3/3, and JSON depth 5 across 13384 objects.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/52600606357c41998634fe876a6f214d.mp3?v=1","evidence_url":"https://aidemos.com/evidence/b8d04964-1003-4de2-b766-52060b761904"}],"other_criteria":[{"id":"9ac338da-02d8-4078-91d8-e0d6a5910839","criterion":"output-quality","criterion_name":"Output quality","rank_role":"decisive","verdict":"struggled","score":25.06,"score_total":null,"note":"The transcript struggles on code-switching speech, with 25.06% WER (560 substitutions, 981 deletions, 92 insertions over 6517 reference words) and only 12.5% Spanish token recall (10/80 types).","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/9ac338da-02d8-4078-91d8-e0d6a5910839"}],"appears_in":[{"page_type":"ranking","slug":"speech-to-text-apis","title":"Best AI Tools for Accurate Speech-to-Text on Hard Audio","url":"https://aidemos.com/best/speech-to-text-apis","binding":"run"}],"same_scenario":[{"id":"4bd8fb2e-e1e0-475d-a6c6-e879538edcaf","tool":"assemblyai","tool_name":"AssemblyAI","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 12,181 timed tokens and JSON depth 5 across 12,187 objects, with payload depth 3/3."},{"id":"00f8839f-963c-4043-8b0a-9f374a3daec5","tool":"aws-transcribe","tool_name":"AWS Transcribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"Returns a full developer payload with payload depth 3/3, 6311 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 21039 objects."},{"id":"14209d19-de86-48f5-95cb-1458098b1164","tool":"deepgram","tool_name":"Deepgram","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich developer payload on code-switching audio, with word-level timing, confidence, speaker labels, 11,684 timed tokens, 3 distinct speakers, payload depth 3/3, and JSON depth 11 across 11,962 objects."},{"id":"6303e8f9-cdc4-49e6-ab0b-8f38c22ff8a3","tool":"elevenlabs-scribe","tool_name":"ElevenLabs Scribe","verdict":"worked","score":3.0,"score_total":3.0,"note":"The returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 12,162 timed tokens."},{"id":"ac711648-1b75-4389-a5c6-a72604ffeab6","tool":"gladia","tool_name":"Gladia","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 26,166 timed tokens, and 7-level JSON nesting across 26,174 objects."},{"id":"a4a565c6-afc7-4855-8d13-360a1f8aa34e","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"Returns a mid-depth transcript payload: payload depth 2/3 with 3688 word-level timed tokens, confidence present, and no speaker labels."},{"id":"a8f4bff9-5f10-484a-81f7-54f78fa7e92b","tool":"groqcloud","tool_name":"GroqCloud","verdict":"worked","score":null,"score_total":null,"note":"Returns the same verbose JSON shape on code-switching speech, including task, language, duration, segments, word timestamps, confidence data, and no speaker labels; the payload depth is 2/3 with 877 timed tokens."},{"id":"69e1f4f4-c76a-49d2-b48e-787e6b416be6","tool":"openai-speech-to-text","tool_name":"OpenAI Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"Returns 5703 word-level timed tokens, but no confidence values or speaker labels; the JSON is shallow at depth 3 with 5705 objects."},{"id":"b8d04964-1003-4de2-b766-52060b761904","tool":"rev-ai","tool_name":"Rev AI","verdict":"worked","score":null,"score_total":null,"note":"Returned a full developer payload with word-level timing, confidence and speaker labels; the raw response walk reports 5855 timed tokens, payload depth 3/3, and JSON depth 5 across 13384 objects."}]}