{"observation":{"id":"e34ea237-175b-4101-8914-9c620f5bc051","tool":"uberduck","tool_name":"Uberduck","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","criterion_definition":"How well the tool handles names, technical terms, numbers, acronyms, and other tricky pronunciations.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If a cloned voice cannot handle names, numbers, acronyms, and technical terms correctly, it fails at producing usable voiceover. (3 of 3 judges)","scenario":"low-quality-voice-sample","scenario_name":"Low-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A noisy voice recording with background noise, room ambience, and minor disturbances, used to test whether voice-cloning tools can preserve speaker identity when the source audio is imperfect.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","role":"input","filename":"speechify-input-low-quality.wav"}],"stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"],"verdict":"worked","score":null,"score_total":null,"note":"Standard-script English words stayed intelligible; the observed problem was delivery/prosody, not misread or garbled words.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-low-quality-input-recording-6fd885c280cb.mp4","role":"input","alt":"Research media low quality input recording.mp4"},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-uberduck-output-low-quality-a89e3da7ab64.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","filename":"speechify-input-low-quality.wav","alt":"Low-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"]},"tool_page_slug":"uberduck","tool_url":"https://aidemos.com/tools/uberduck","permalink":"https://aidemos.com/evidence/e34ea237-175b-4101-8914-9c620f5bc051","api_url":"https://ai.aidemos.com/v1/observations/e34ea237-175b-4101-8914-9c620f5bc051"},"peers":[{"id":"abac210b-03cb-41ef-b8d5-1709b9b95535","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Keeps pronunciation stable even when the source audio is noisy, with no major pronunciation breakdown reported in the generated narration.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-lowquality-input-8230726a9b39.wav","evidence_url":"https://aidemos.com/evidence/abac210b-03cb-41ef-b8d5-1709b9b95535"},{"id":"3d70358d-58fc-4add-8fdb-b9452f058e36","tool":"inworld","tool_name":"Inworld","verdict":"worked","score":null,"score_total":null,"note":"Across the low-quality pass, all words stayed intelligible and no misread or garbled words were noted; no dedicated names/numbers/acronyms stress test was run on this scenario.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-lowquality-input-437133b0167f.wav","evidence_url":"https://aidemos.com/evidence/3d70358d-58fc-4add-8fdb-b9452f058e36"},{"id":"5180604a-c600-4409-a0a8-7d1afeac41c8","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"The low-quality English generation mispronounced a small number of words, with the report calling out 2–3 misreads in this clip.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-output-low-quality-27c7c65786a5.mp3","evidence_url":"https://aidemos.com/evidence/5180604a-c600-4409-a0a8-7d1afeac41c8"},{"id":"69856bcc-f984-4afc-8701-290d4c80a13e","tool":"speechify","tool_name":"Speechify","verdict":"worked","score":null,"score_total":null,"note":"The noisy-output pass stayed pronunciation-clear, and the researcher flagged no specific mispronounced words.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-speechify-lowquality-input-86c01eae771e.wav","evidence_url":"https://aidemos.com/evidence/69856bcc-f984-4afc-8701-290d4c80a13e"},{"id":"2d1ef955-9a21-441b-94ee-c885f8bd300a","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"worked","score":null,"score_total":null,"note":"Gen keeps the standard script intelligible on the noisy sample; the robotic tone is a delivery issue, not a misread-word problem.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-lowquality-input-5c3da58b8b17.wav","evidence_url":"https://aidemos.com/evidence/2d1ef955-9a21-441b-94ee-c885f8bd300a"},{"id":"94823a2b-92cc-498d-8d17-b5cebea9516d","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The noisy-sample run did not show major pronunciation problems.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-output-low-quality-89f3e7da2000.wav","evidence_url":"https://aidemos.com/evidence/94823a2b-92cc-498d-8d17-b5cebea9516d"}],"other_criteria":[{"id":"aab42c38-9b8f-47e6-aea3-f731d0fffffb","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Quality stayed uniformly poor from start to finish rather than degrading later in the passage, so no long-form drift was observed.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/aab42c38-9b8f-47e6-aea3-f731d0fffffb"},{"id":"5959899b-bf1b-4519-b6c4-33074f3d30c1","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"The output sounded heavily robotic, with frequent unnatural pauses that made it immediately identifiable as AI-generated.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/5959899b-bf1b-4519-b6c4-33074f3d30c1"},{"id":"2fb082d1-14df-479f-8289-a6b6d624a2b4","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"The noisy-source clone barely resembled the original speaker and was described as the weakest voice-match result in the round.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/2fb082d1-14df-479f-8289-a6b6d624a2b4"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"abac210b-03cb-41ef-b8d5-1709b9b95535","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Keeps pronunciation stable even when the source audio is noisy, with no major pronunciation breakdown reported in the generated narration."},{"id":"3d70358d-58fc-4add-8fdb-b9452f058e36","tool":"inworld","tool_name":"Inworld","verdict":"worked","score":null,"score_total":null,"note":"Across the low-quality pass, all words stayed intelligible and no misread or garbled words were noted; no dedicated names/numbers/acronyms stress test was run on this scenario."},{"id":"5180604a-c600-4409-a0a8-7d1afeac41c8","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"The low-quality English generation mispronounced a small number of words, with the report calling out 2–3 misreads in this clip."},{"id":"69856bcc-f984-4afc-8701-290d4c80a13e","tool":"speechify","tool_name":"Speechify","verdict":"worked","score":null,"score_total":null,"note":"The noisy-output pass stayed pronunciation-clear, and the researcher flagged no specific mispronounced words."},{"id":"2d1ef955-9a21-441b-94ee-c885f8bd300a","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"worked","score":null,"score_total":null,"note":"Gen keeps the standard script intelligible on the noisy sample; the robotic tone is a delivery issue, not a misread-word problem."},{"id":"94823a2b-92cc-498d-8d17-b5cebea9516d","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The noisy-sample run did not show major pronunciation problems."}]}