{"observation":{"id":"11bf336a-1887-480d-afb6-b231fbc3d21c","tool":"uberduck","tool_name":"Uberduck","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","criterion_definition":"Whether voice quality, pacing, and pronunciation stay consistent over longer passages instead of degrading after a few sentences.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"For text-to-voiceover work, the voice must stay stable across longer scripts; degradation means the output is not reliable. (3 of 3 judges)","scenario":"high-quality-voice-sample","scenario_name":"High-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","role":"input","filename":"heygen-input-high-quality.wav"}],"stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"],"verdict":"worked","score":null,"score_total":null,"note":"The output stayed poor throughout the passage rather than degrading over length, so no later-stage dropoff was observed.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-high-quality-input-recording-f326629d7802.mp4","role":"input","alt":"Research media high quality input recording.mp4"},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-uberduck-output-high-quality-e9f6b745f2cd.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","filename":"heygen-input-high-quality.wav","alt":"High-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"]},"tool_page_slug":"uberduck","tool_url":"https://aidemos.com/tools/uberduck","permalink":"https://aidemos.com/evidence/11bf336a-1887-480d-afb6-b231fbc3d21c","api_url":"https://ai.aidemos.com/v1/observations/11bf336a-1887-480d-afb6-b231fbc3d21c"},"peers":[{"id":"44e04da3-764d-4c4f-9b78-58fbdcfe7fd9","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Keeps voice stability across extended scripts and is described as good for long-form narration, with no reported degradation over longer passages.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-highquality-input-4dac94a57d04.wav","evidence_url":"https://aidemos.com/evidence/44e04da3-764d-4c4f-9b78-58fbdcfe7fd9"},{"id":"8ff9b798-baad-4160-8b98-7a90b05a1ca0","tool":"heygen","tool_name":"Heygen","verdict":"struggled","score":null,"score_total":null,"note":"The longer-script high-quality generation had flow interruptions, word mispronunciations, inconsistent delivery, and the report says multiple regenerations may be required for production-ready results.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-high-quality-variant-3-bes-3b44d7550666.wav","evidence_url":"https://aidemos.com/evidence/8ff9b798-baad-4160-8b98-7a90b05a1ca0"},{"id":"ff632063-0449-48b6-b605-f9c7176feabe","tool":"inworld","tool_name":"Inworld","verdict":"struggled","score":null,"score_total":null,"note":"In the ~1:04 high-quality output, matching accuracy degraded further as the clip progressed, so consistency weakened over the longer passage.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-highquality-input-5a7af892998c.wav","evidence_url":"https://aidemos.com/evidence/ff632063-0449-48b6-b605-f9c7176feabe"},{"id":"c62a8366-7e61-4481-89d2-713802e118f4","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"The complete ~19-second output held up fine at that length, with no visible degradation reported within the generated clip.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-output-high-quality-64a65a87ed76.mp3","evidence_url":"https://aidemos.com/evidence/c62a8366-7e61-4481-89d2-713802e118f4"},{"id":"eedd48b7-de42-4024-a225-dfe870fd6e87","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"worked","score":null,"score_total":null,"note":"Gen remains consistent through the generated narration on the clean sample, with no interruptions or instability observed.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-highquality-input-6c26a9323e7b.wav","evidence_url":"https://aidemos.com/evidence/eedd48b7-de42-4024-a225-dfe870fd6e87"},{"id":"2ae4e858-ea5a-4e0e-816a-197ac880c680","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"Consistency was maintained throughout the script, though the speech still ran faster than expected.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-output-high-quality-1583ecfff017.wav","evidence_url":"https://aidemos.com/evidence/2ae4e858-ea5a-4e0e-816a-197ac880c680"}],"other_criteria":[{"id":"36d63d33-707c-4265-95d6-5bbea891d148","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"The cleaner input still produced a robotic delivery with awkward pauses breaking up the speech.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/36d63d33-707c-4265-95d6-5bbea891d148"},{"id":"7d5bd84f-3369-49b9-9539-32d3849c8a66","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"English words remained intelligible, and the cleaner source did not change that ceiling on this run.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/7d5bd84f-3369-49b9-9539-32d3849c8a66"},{"id":"17d6e619-374f-40bd-af46-c8d70522df45","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Cleaner source audio did not materially improve identity matching; the clone still did not closely track the original voice.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/17d6e619-374f-40bd-af46-c8d70522df45"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"44e04da3-764d-4c4f-9b78-58fbdcfe7fd9","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Keeps voice stability across extended scripts and is described as good for long-form narration, with no reported degradation over longer passages."},{"id":"8ff9b798-baad-4160-8b98-7a90b05a1ca0","tool":"heygen","tool_name":"Heygen","verdict":"struggled","score":null,"score_total":null,"note":"The longer-script high-quality generation had flow interruptions, word mispronunciations, inconsistent delivery, and the report says multiple regenerations may be required for production-ready results."},{"id":"ff632063-0449-48b6-b605-f9c7176feabe","tool":"inworld","tool_name":"Inworld","verdict":"struggled","score":null,"score_total":null,"note":"In the ~1:04 high-quality output, matching accuracy degraded further as the clip progressed, so consistency weakened over the longer passage."},{"id":"c62a8366-7e61-4481-89d2-713802e118f4","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"The complete ~19-second output held up fine at that length, with no visible degradation reported within the generated clip."},{"id":"eedd48b7-de42-4024-a225-dfe870fd6e87","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"worked","score":null,"score_total":null,"note":"Gen remains consistent through the generated narration on the clean sample, with no interruptions or instability observed."},{"id":"2ae4e858-ea5a-4e0e-816a-197ac880c680","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"Consistency was maintained throughout the script, though the speech still ran faster than expected."}]}