{"observation":{"id":"0e63740f-6710-4e59-896e-d26271ddc32f","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","criterion_definition":"Whether voice quality, pacing, and pronunciation stay consistent over longer passages instead of degrading after a few sentences.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"For text-to-voiceover work, the voice must stay stable across longer scripts; degradation means the output is not reliable. (3 of 3 judges)","scenario":"high-quality-voice-sample","scenario_name":"High-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","role":"input","filename":"heygen-input-high-quality.wav"}],"stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"],"verdict":"worked","score":null,"score_total":null,"note":"Gen+ stays stable through long-form generation on the clean sample, with no voice breaks or instability observed.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-highquality-input-6c26a9323e7b.wav","role":"input","alt":"Research media topmediai highquality input.wav"},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-output-high-quality-genplus-5e1ea54f0a66.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","filename":"heygen-input-high-quality.wav","alt":"High-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"]},"tool_page_slug":"topmediai-voice-cloning","tool_url":"https://aidemos.com/tools/topmediai-voice-cloning","permalink":"https://aidemos.com/evidence/0e63740f-6710-4e59-896e-d26271ddc32f","api_url":"https://ai.aidemos.com/v1/observations/0e63740f-6710-4e59-896e-d26271ddc32f"},"peers":[{"id":"44e04da3-764d-4c4f-9b78-58fbdcfe7fd9","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Keeps voice stability across extended scripts and is described as good for long-form narration, with no reported degradation over longer passages.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-highquality-input-4dac94a57d04.wav","evidence_url":"https://aidemos.com/evidence/44e04da3-764d-4c4f-9b78-58fbdcfe7fd9"},{"id":"8ff9b798-baad-4160-8b98-7a90b05a1ca0","tool":"heygen","tool_name":"Heygen","verdict":"struggled","score":null,"score_total":null,"note":"The longer-script high-quality generation had flow interruptions, word mispronunciations, inconsistent delivery, and the report says multiple regenerations may be required for production-ready results.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-high-quality-variant-3-bes-3b44d7550666.wav","evidence_url":"https://aidemos.com/evidence/8ff9b798-baad-4160-8b98-7a90b05a1ca0"},{"id":"ff632063-0449-48b6-b605-f9c7176feabe","tool":"inworld","tool_name":"Inworld","verdict":"struggled","score":null,"score_total":null,"note":"In the ~1:04 high-quality output, matching accuracy degraded further as the clip progressed, so consistency weakened over the longer passage.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-highquality-input-5a7af892998c.wav","evidence_url":"https://aidemos.com/evidence/ff632063-0449-48b6-b605-f9c7176feabe"},{"id":"c62a8366-7e61-4481-89d2-713802e118f4","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"The complete ~19-second output held up fine at that length, with no visible degradation reported within the generated clip.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-output-high-quality-64a65a87ed76.mp3","evidence_url":"https://aidemos.com/evidence/c62a8366-7e61-4481-89d2-713802e118f4"},{"id":"11bf336a-1887-480d-afb6-b231fbc3d21c","tool":"uberduck","tool_name":"Uberduck","verdict":"worked","score":null,"score_total":null,"note":"The output stayed poor throughout the passage rather than degrading over length, so no later-stage dropoff was observed.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-high-quality-input-recording-f326629d7802.mp4","evidence_url":"https://aidemos.com/evidence/11bf336a-1887-480d-afb6-b231fbc3d21c"},{"id":"2ae4e858-ea5a-4e0e-816a-197ac880c680","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"Consistency was maintained throughout the script, though the speech still ran faster than expected.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-output-high-quality-1583ecfff017.wav","evidence_url":"https://aidemos.com/evidence/2ae4e858-ea5a-4e0e-816a-197ac880c680"}],"other_criteria":[{"id":"29181eec-a14d-4508-97c6-e0a5e0f9345c","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Gen still sounds robotic on the clean source sample, making it less convincing than HD.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/29181eec-a14d-4508-97c6-e0a5e0f9345c"},{"id":"13dae80b-2bff-4c46-bed3-b2429db82ad6","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD is the most human-like and realistic output on the clean sample, with better emotional delivery and conversational flow.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/13dae80b-2bff-4c46-bed3-b2429db82ad6"},{"id":"ff84cb60-0eed-46f9-9f33-90c410ef58f3","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen+ sounds more natural than Gen, but the weak identity preservation keeps the output from feeling fully convincing.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/ff84cb60-0eed-46f9-9f33-90c410ef58f3"},{"id":"d72cf668-6291-4489-b905-cbb5deacabb6","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD has no misread or garbled words on the clean sample and is the cleanest pronunciation result across all nine English generations tested.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/d72cf668-6291-4489-b905-cbb5deacabb6"},{"id":"b62a2ccb-2b14-4585-b4de-0b510c9887c2","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen keeps the script intelligible on the clean sample; the robotic character is a delivery issue, not misread words.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/b62a2ccb-2b14-4585-b4de-0b510c9887c2"},{"id":"91aabbbb-5272-4dcf-9cb8-bd369fbfa7a2","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen+ keeps the script intelligible on the clean sample despite the gender shift, so pronunciation remains intact.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/91aabbbb-5272-4dcf-9cb8-bd369fbfa7a2"},{"id":"7dd3ad3b-020a-4688-beaa-d24097791ff2","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD has the highest similarity to the clean source voice and preserves speaker identity best among the high-quality outputs.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/7dd3ad3b-020a-4688-beaa-d24097791ff2"},{"id":"078f1ee7-15c1-4a69-870b-ecfb34672ef0","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen keeps some similarity to the clean source voice, but speaker identity remains only moderately accurate and still sounds robotic.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/078f1ee7-15c1-4a69-870b-ecfb34672ef0"},{"id":"8167b0ac-f9dd-4ff4-aab8-d1bbda394f44","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Gen+ reduces similarity to the clean source voice by shifting toward a feminine tone, so speaker representation is inaccurate.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/8167b0ac-f9dd-4ff4-aab8-d1bbda394f44"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"44e04da3-764d-4c4f-9b78-58fbdcfe7fd9","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Keeps voice stability across extended scripts and is described as good for long-form narration, with no reported degradation over longer passages."},{"id":"8ff9b798-baad-4160-8b98-7a90b05a1ca0","tool":"heygen","tool_name":"Heygen","verdict":"struggled","score":null,"score_total":null,"note":"The longer-script high-quality generation had flow interruptions, word mispronunciations, inconsistent delivery, and the report says multiple regenerations may be required for production-ready results."},{"id":"ff632063-0449-48b6-b605-f9c7176feabe","tool":"inworld","tool_name":"Inworld","verdict":"struggled","score":null,"score_total":null,"note":"In the ~1:04 high-quality output, matching accuracy degraded further as the clip progressed, so consistency weakened over the longer passage."},{"id":"c62a8366-7e61-4481-89d2-713802e118f4","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"The complete ~19-second output held up fine at that length, with no visible degradation reported within the generated clip."},{"id":"11bf336a-1887-480d-afb6-b231fbc3d21c","tool":"uberduck","tool_name":"Uberduck","verdict":"worked","score":null,"score_total":null,"note":"The output stayed poor throughout the passage rather than degrading over length, so no later-stage dropoff was observed."},{"id":"2ae4e858-ea5a-4e0e-816a-197ac880c680","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"Consistency was maintained throughout the script, though the speech still ran faster than expected."}]}