{"observation":{"id":"fe74b58b-85ed-41a9-8e51-2fb32bc978c3","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","criterion_definition":"Whether voice quality, pacing, and pronunciation stay consistent over longer passages instead of degrading after a few sentences.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"For text-to-voiceover work, the voice must stay stable across longer scripts; degradation means the output is not reliable. (3 of 3 judges)","scenario":"low-quality-voice-sample","scenario_name":"Low-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A noisy voice recording with background noise, room ambience, and minor disturbances, used to test whether voice-cloning tools can preserve speaker identity when the source audio is imperfect.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","role":"input","filename":"speechify-input-low-quality.wav"}],"stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"],"verdict":"worked","score":null,"score_total":null,"note":"HD stays consistent through long-form narration on the noisy sample, with no abrupt changes in voice quality or pronunciation.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-lowquality-input-5c3da58b8b17.wav","role":"input","alt":"Research media topmediai lowquality input.wav"},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-output-low-quality-hd-390d1925e5b0.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","filename":"speechify-input-low-quality.wav","alt":"Low-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"]},"tool_page_slug":"topmediai-voice-cloning","tool_url":"https://aidemos.com/tools/topmediai-voice-cloning","permalink":"https://aidemos.com/evidence/fe74b58b-85ed-41a9-8e51-2fb32bc978c3","api_url":"https://ai.aidemos.com/v1/observations/fe74b58b-85ed-41a9-8e51-2fb32bc978c3"},"peers":[{"id":"0e1d6aba-d454-473a-b230-39892621a392","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Holds voice quality and pronunciation stable across extended narration, with no major degradation reported during long-form generation.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-lowquality-input-8230726a9b39.wav","evidence_url":"https://aidemos.com/evidence/0e1d6aba-d454-473a-b230-39892621a392"},{"id":"ac775bf1-ecc4-4d54-a9ed-623341ccd3fe","tool":"heygen","tool_name":"Heygen","verdict":"struggled","score":null,"score_total":null,"note":"In the longer-script generation, the voice stayed human-like but broke conversational flow, mispronounced certain words, and became inconsistent across longer passages.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-low-quality-variant-3-best-5e23070dc089.wav","evidence_url":"https://aidemos.com/evidence/ac775bf1-ecc4-4d54-a9ed-623341ccd3fe"},{"id":"a2ea2d42-3d95-437d-b7c1-c45c49784931","tool":"inworld","tool_name":"Inworld","verdict":"struggled","score":null,"score_total":null,"note":"The ~53-second low-quality output drifted further from the source voice by the end of the clip, making long-form consistency a clear weak point.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-lowquality-input-437133b0167f.wav","evidence_url":"https://aidemos.com/evidence/a2ea2d42-3d95-437d-b7c1-c45c49784931"},{"id":"be49be53-e125-48b2-b302-ca32d4a3710e","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"On the complete ~18-second output, pacing stayed steady and natural pauses held up through the clip.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-output-low-quality-27c7c65786a5.mp3","evidence_url":"https://aidemos.com/evidence/be49be53-e125-48b2-b302-ca32d4a3710e"},{"id":"aab42c38-9b8f-47e6-aea3-f731d0fffffb","tool":"uberduck","tool_name":"Uberduck","verdict":"worked","score":null,"score_total":null,"note":"Quality stayed uniformly poor from start to finish rather than degrading later in the passage, so no long-form drift was observed.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-low-quality-input-recording-6fd885c280cb.mp4","evidence_url":"https://aidemos.com/evidence/aab42c38-9b8f-47e6-aea3-f731d0fffffb"},{"id":"0568d899-df5c-41ca-afa5-263120606661","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The output stayed consistent throughout the generated script, with no major pronunciation issues, although the pacing was noticeably fast.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-output-low-quality-89f3e7da2000.wav","evidence_url":"https://aidemos.com/evidence/0568d899-df5c-41ca-afa5-263120606661"}],"other_criteria":[{"id":"6f2471dd-f005-4b7d-8798-ad028ed58810","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD is the most human-like output on the noisy sample, with better emotional tone, speech rhythm, and vocal realism than Gen or Gen+.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/6f2471dd-f005-4b7d-8798-ad028ed58810"},{"id":"b14dd1ff-8fa6-4073-bf38-0250dbcc735e","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Gen sounds noticeably robotic and lacks emotional depth and natural speech rhythm on the noisy sample.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/b14dd1ff-8fa6-4073-bf38-0250dbcc735e"},{"id":"ced3a32c-621b-48e4-ab5d-0259933c2238","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen+ is more natural than Gen, but the gender inconsistency keeps the overall cloning quality from feeling convincing.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/ced3a32c-621b-48e4-ab5d-0259933c2238"},{"id":"2d1ef955-9a21-441b-94ee-c885f8bd300a","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen keeps the standard script intelligible on the noisy sample; the robotic tone is a delivery issue, not a misread-word problem.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/2d1ef955-9a21-441b-94ee-c885f8bd300a"},{"id":"88aa7f94-5cdd-436e-9da3-d4a6c7d43f40","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen+ keeps the script intelligible on the noisy sample despite the gender-shift issue, and the report notes no dedicated pronunciation stress test for this scenario.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/88aa7f94-5cdd-436e-9da3-d4a6c7d43f40"},{"id":"e6e02fc5-bd69-4bad-a81f-c679c77f1540","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD delivers the best pronunciation result on the noisy sample, with no misread or garbled words noted during the listening pass.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/e6e02fc5-bd69-4bad-a81f-c679c77f1540"},{"id":"610203fe-2243-4679-a01c-d2a71a528601","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Gen+ loses speaker identity on the noisy sample by drifting toward a feminine vocal tone instead of the source male voice.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/610203fe-2243-4679-a01c-d2a71a528601"},{"id":"86b07cc2-a3a8-4612-85ce-cef2a6853f8a","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen keeps some resemblance to the original speaker on the noisy source sample, but it only partially preserves identity and still sounds robotic.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/86b07cc2-a3a8-4612-85ce-cef2a6853f8a"},{"id":"8452d11e-1149-4195-9eb0-5817433e2cc4","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD is the closest match on the noisy sample and preserves vocal identity best among the three variants.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/8452d11e-1149-4195-9eb0-5817433e2cc4"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"0e1d6aba-d454-473a-b230-39892621a392","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Holds voice quality and pronunciation stable across extended narration, with no major degradation reported during long-form generation."},{"id":"ac775bf1-ecc4-4d54-a9ed-623341ccd3fe","tool":"heygen","tool_name":"Heygen","verdict":"struggled","score":null,"score_total":null,"note":"In the longer-script generation, the voice stayed human-like but broke conversational flow, mispronounced certain words, and became inconsistent across longer passages."},{"id":"a2ea2d42-3d95-437d-b7c1-c45c49784931","tool":"inworld","tool_name":"Inworld","verdict":"struggled","score":null,"score_total":null,"note":"The ~53-second low-quality output drifted further from the source voice by the end of the clip, making long-form consistency a clear weak point."},{"id":"be49be53-e125-48b2-b302-ca32d4a3710e","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"On the complete ~18-second output, pacing stayed steady and natural pauses held up through the clip."},{"id":"aab42c38-9b8f-47e6-aea3-f731d0fffffb","tool":"uberduck","tool_name":"Uberduck","verdict":"worked","score":null,"score_total":null,"note":"Quality stayed uniformly poor from start to finish rather than degrading later in the passage, so no long-form drift was observed."},{"id":"0568d899-df5c-41ca-afa5-263120606661","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The output stayed consistent throughout the generated script, with no major pronunciation issues, although the pacing was noticeably fast."}]}