{"observation":{"id":"d55e9863-9ab8-47c6-afac-8370461e2157","tool":"gladia","tool_name":"Gladia","criterion":"automation-level","criterion_name":"Automation level","criterion_definition":"How many API steps or calls the workflow requires, and whether it completes without operator input.","criterion_evidence_type":"capability","criterion_rank_role":"context","criterion_rank_role_reason":"How many API steps or whether operator input is needed affects convenience and workflow, but not whether the engine transcribes accurately once run. (3 of 3 judges)","scenario":"bilingual-spanish-english-code-switching-speech","scenario_name":"Bilingual Spanish-English code-switching speech","group_tag":"speech-to-text-benchmark","scenario_description":"A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.","modality":"audio","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1","role":"input","filename":"mix_language.mp3"}],"stresses":["code-switching detection","multilingual language ID","mid-sentence language transitions","word preservation across language flips","hallucination resistance in bilingual speech"],"verdict":"worked","score":null,"score_total":null,"note":"The workflow completes end to end without operator input, using the vendor's documented three-call sequence (upload, submit audio_url to /pre-recorded, then poll result_url until done); this run was multi-stage and did not instrument per-call timings.","evidence_state":"observed","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-gladia-bc37c24594b6.md","role":"context","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/f84aa813e76242f58daae74ae5190208.png?v=1","role":"context","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/fce72a78e0cc44989d44de66037f8d9f.png?v=1","role":"context","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/5e57637448d1453dbfcfce4f0799a344.png?v=1","role":"context","alt":null}],"run_id":"469de0c2-d727-4f8f-a60e-e3a5bf8e8588","study_title":"Transcribe Audio Accurately — Speech-to-Text Engine Benchmark","study_kind":"generation","research_task":"86baxegpu","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/8883983a83e848b2a3aef2f1c82c01db.mp3?v=1","filename":"mix_language.mp3","alt":"Bilingual Spanish-English code-switching speech","role":"input"}],"modality":"audio","stresses":["code-switching detection","multilingual language ID","mid-sentence language transitions","word preservation across language flips","hallucination resistance in bilingual speech"]},"tool_page_slug":"gladia","tool_url":"https://aidemos.com/tools/gladia","permalink":"https://aidemos.com/evidence/d55e9863-9ab8-47c6-afac-8370461e2157","api_url":"https://ai.aidemos.com/v1/observations/d55e9863-9ab8-47c6-afac-8370461e2157"},"peers":[{"id":"0ea7540d-eecb-456a-9896-1801df51307f","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"The recorded workflow submits the audio as a POST to the v2 endpoint and reaches a scored result without operator input, but the trace explicitly says per-call timings were not instrumented, so no measured call count is claimed.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/0563f1e21d4341d7ab72a176de39a054.png?v=1","evidence_url":"https://aidemos.com/evidence/0ea7540d-eecb-456a-9896-1801df51307f"},{"id":"40fbd740-b1ba-4d5a-a1fd-bfbc6990b8bc","tool":"groqcloud","tool_name":"GroqCloud","verdict":"worked","score":null,"score_total":null,"note":"Completes the bilingual run as a single POST with the response returned inline and no operator intervention; the run trace again says per-call timings were not instrumented.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/4c51807b815742a08ed7fcf03d8165c8.png?v=1","evidence_url":"https://aidemos.com/evidence/40fbd740-b1ba-4d5a-a1fd-bfbc6990b8bc"}],"other_criteria":[{"id":"ac711648-1b75-4389-a5c6-a72604ffeab6","criterion":"export","criterion_name":"Export","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Returns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 26,166 timed tokens, and 7-level JSON nesting across 26,174 objects.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/ac711648-1b75-4389-a5c6-a72604ffeab6"},{"id":"fcccec3d-9bf1-46cc-8e80-e79e347f3665","criterion":"output-quality","criterion_name":"Output quality","rank_role":"decisive","verdict":"failed","score":88.45,"score_total":null,"note":"On bilingual code-switching, the response duplicated across 2 channels, inflating output to 10,765 words against a 6,517-word reference and making the reported 88.45% WER an artefact of duplication rather than a usable accuracy score.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/fcccec3d-9bf1-46cc-8e80-e79e347f3665"}],"appears_in":[{"page_type":"ranking","slug":"speech-to-text-apis","title":"Best AI Tools for Accurate Speech-to-Text on Hard Audio","url":"https://aidemos.com/best/speech-to-text-apis","binding":"run"}],"same_scenario":[{"id":"0ea7540d-eecb-456a-9896-1801df51307f","tool":"google-cloud-speech-to-text","tool_name":"Google Cloud Speech-to-Text","verdict":"mixed","score":null,"score_total":null,"note":"The recorded workflow submits the audio as a POST to the v2 endpoint and reaches a scored result without operator input, but the trace explicitly says per-call timings were not instrumented, so no measured call count is claimed."},{"id":"40fbd740-b1ba-4d5a-a1fd-bfbc6990b8bc","tool":"groqcloud","tool_name":"GroqCloud","verdict":"worked","score":null,"score_total":null,"note":"Completes the bilingual run as a single POST with the response returned inline and no operator intervention; the run trace again says per-call timings were not instrumented."}]}