{"observation":{"id":"86b07cc2-a3a8-4612-85ce-cef2a6853f8a","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","criterion_definition":"Whether the generated voice actually sounds like the original speaker, including tone, pitch, rhythm, and identity.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"This ranking is fundamentally about whether the generated speech still sounds like the target speaker, so identity match is core. (3 of 3 judges)","scenario":"low-quality-voice-sample","scenario_name":"Low-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A noisy voice recording with background noise, room ambience, and minor disturbances, used to test whether voice-cloning tools can preserve speaker identity when the source audio is imperfect.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","role":"input","filename":"speechify-input-low-quality.wav"}],"stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"],"verdict":"mixed","score":null,"score_total":null,"note":"Gen keeps some resemblance to the original speaker on the noisy source sample, but it only partially preserves identity and still sounds robotic.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-lowquality-input-5c3da58b8b17.wav","role":"input","alt":null},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-output-low-quality-gen-77e8de3857f1.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","filename":"speechify-input-low-quality.wav","alt":"Low-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"]},"tool_page_slug":"topmediai-voice-cloning","tool_url":"https://aidemos.com/tools/topmediai-voice-cloning","permalink":"https://aidemos.com/evidence/86b07cc2-a3a8-4612-85ce-cef2a6853f8a","api_url":"https://ai.aidemos.com/v1/observations/86b07cc2-a3a8-4612-85ce-cef2a6853f8a"},"peers":[{"id":"101630ef-8867-4f0a-adc2-b422996a45c8","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"Voice cloning stays strong even from a noisy source, with the report estimating about 95% similarity to the original speaker and preservation of most vocal characteristics.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-input-low-quality-d1a2d94e8e38.wav","evidence_url":"https://aidemos.com/evidence/101630ef-8867-4f0a-adc2-b422996a45c8"},{"id":"5d946255-1867-40d2-8dc5-17695ecf4038","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"mixed","score":50.0,"score_total":null,"note":"Clones only about half of the source-speaker identity in the noisy sample: the report rates the match at approximately 50% and says the output sounds heavily polished, which lowers resemblance instead of faithfully reproducing the original voice.","artifact_count":6,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-lowquality-input-8230726a9b39.wav","evidence_url":"https://aidemos.com/evidence/5d946255-1867-40d2-8dc5-17695ecf4038"},{"id":"533fcf8b-1d9b-4828-81a2-3cd52e3f90ec","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Preserves speaker identity well even from a noisy source: both low-quality English outputs were described as strong, close matches to the original voice.","artifact_count":4,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-fishaudio-lowquality-input-a57be1fc5dd8.wav","evidence_url":"https://aidemos.com/evidence/533fcf8b-1d9b-4828-81a2-3cd52e3f90ec"},{"id":"5c83e9e6-5abd-49fb-b1bf-94d0d6b128d0","tool":"heygen","tool_name":"Heygen","verdict":"mixed","score":null,"score_total":null,"note":"A mid-tier low-quality clone reached roughly 70–80% similarity to the original voice, making it acceptable for short-form use but still not fully faithful.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-low-quality-variant-2-90acc9fed5f4.wav","evidence_url":"https://aidemos.com/evidence/5c83e9e6-5abd-49fb-b1bf-94d0d6b128d0"},{"id":"90b33678-feed-449c-a618-4f55590cfe75","tool":"inworld","tool_name":"Inworld","verdict":"mixed","score":null,"score_total":null,"note":"On the ~53-second low-quality clone, identity match was strongest at the start and then faded gradually as the clip progressed, so the voice was only partially consistent rather than exact.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-lowquality-input-437133b0167f.wav","evidence_url":"https://aidemos.com/evidence/90b33678-feed-449c-a618-4f55590cfe75"},{"id":"72f96b15-3f48-4e62-9f91-db706af4f1cf","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"With a noisy source recording, the clone only partially preserved speaker identity: the generated voice was noticeably softer than the original and the report rates the match as Fair (~35–45%).","artifact_count":5,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-lowquality-input-6696703e9405.wav","evidence_url":"https://aidemos.com/evidence/72f96b15-3f48-4e62-9f91-db706af4f1cf"},{"id":"212d6091-fe7b-4e74-89be-5b2a0e3cf377","tool":"speechify","tool_name":"Speechify","verdict":"failed","score":null,"score_total":null,"note":"On the noisy source, Speechify produced a clone that diverged strongly from the speaker and even sounded female despite a male input.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-speechify-lowquality-input-86c01eae771e.wav","evidence_url":"https://aidemos.com/evidence/212d6091-fe7b-4e74-89be-5b2a0e3cf377"},{"id":"2fb082d1-14df-479f-8289-a6b6d624a2b4","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"The noisy-source clone barely resembled the original speaker and was described as the weakest voice-match result in the round.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-low-quality-input-recording-6fd885c280cb.mp4","evidence_url":"https://aidemos.com/evidence/2fb082d1-14df-479f-8289-a6b6d624a2b4"},{"id":"97e20d2d-bf92-4a26-b8be-1cbd2f3fd861","tool":"vocalai","tool_name":"VocalAI","verdict":"struggled","score":null,"score_total":null,"note":"The clone preserved only about 10–15% of the original speaker's identity, and the output sounded heavily polished and processed rather than speaker-faithful.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-lowquality-input-447ab8d8cc52.wav","evidence_url":"https://aidemos.com/evidence/97e20d2d-bf92-4a26-b8be-1cbd2f3fd861"}],"other_criteria":[{"id":"39db1c35-17db-4d0d-aeae-27f01f22c7ab","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen stays consistent through longer passages on the noisy sample, with no major voice breaks, glitches, or abrupt tonal shifts.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/39db1c35-17db-4d0d-aeae-27f01f22c7ab"},{"id":"fe74b58b-85ed-41a9-8e51-2fb32bc978c3","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD stays consistent through long-form narration on the noisy sample, with no abrupt changes in voice quality or pronunciation.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/fe74b58b-85ed-41a9-8e51-2fb32bc978c3"},{"id":"b787124b-be9f-4b59-b40a-89b769cac8fa","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen+ remains stable through long-form generation on the noisy sample, with no noticeable interruptions or voice degradation.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/b787124b-be9f-4b59-b40a-89b769cac8fa"},{"id":"6f2471dd-f005-4b7d-8798-ad028ed58810","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD is the most human-like output on the noisy sample, with better emotional tone, speech rhythm, and vocal realism than Gen or Gen+.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/6f2471dd-f005-4b7d-8798-ad028ed58810"},{"id":"ced3a32c-621b-48e4-ab5d-0259933c2238","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen+ is more natural than Gen, but the gender inconsistency keeps the overall cloning quality from feeling convincing.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/ced3a32c-621b-48e4-ab5d-0259933c2238"},{"id":"b14dd1ff-8fa6-4073-bf38-0250dbcc735e","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Gen sounds noticeably robotic and lacks emotional depth and natural speech rhythm on the noisy sample.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/b14dd1ff-8fa6-4073-bf38-0250dbcc735e"},{"id":"2d1ef955-9a21-441b-94ee-c885f8bd300a","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen keeps the standard script intelligible on the noisy sample; the robotic tone is a delivery issue, not a misread-word problem.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/2d1ef955-9a21-441b-94ee-c885f8bd300a"},{"id":"e6e02fc5-bd69-4bad-a81f-c679c77f1540","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD delivers the best pronunciation result on the noisy sample, with no misread or garbled words noted during the listening pass.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/e6e02fc5-bd69-4bad-a81f-c679c77f1540"},{"id":"88aa7f94-5cdd-436e-9da3-d4a6c7d43f40","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen+ keeps the script intelligible on the noisy sample despite the gender-shift issue, and the report notes no dedicated pronunciation stress test for this scenario.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/88aa7f94-5cdd-436e-9da3-d4a6c7d43f40"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"101630ef-8867-4f0a-adc2-b422996a45c8","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"Voice cloning stays strong even from a noisy source, with the report estimating about 95% similarity to the original speaker and preservation of most vocal characteristics."},{"id":"5d946255-1867-40d2-8dc5-17695ecf4038","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"mixed","score":50.0,"score_total":null,"note":"Clones only about half of the source-speaker identity in the noisy sample: the report rates the match at approximately 50% and says the output sounds heavily polished, which lowers resemblance instead of faithfully reproducing the original voice."},{"id":"533fcf8b-1d9b-4828-81a2-3cd52e3f90ec","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Preserves speaker identity well even from a noisy source: both low-quality English outputs were described as strong, close matches to the original voice."},{"id":"5c83e9e6-5abd-49fb-b1bf-94d0d6b128d0","tool":"heygen","tool_name":"Heygen","verdict":"mixed","score":null,"score_total":null,"note":"A mid-tier low-quality clone reached roughly 70–80% similarity to the original voice, making it acceptable for short-form use but still not fully faithful."},{"id":"90b33678-feed-449c-a618-4f55590cfe75","tool":"inworld","tool_name":"Inworld","verdict":"mixed","score":null,"score_total":null,"note":"On the ~53-second low-quality clone, identity match was strongest at the start and then faded gradually as the clip progressed, so the voice was only partially consistent rather than exact."},{"id":"72f96b15-3f48-4e62-9f91-db706af4f1cf","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"With a noisy source recording, the clone only partially preserved speaker identity: the generated voice was noticeably softer than the original and the report rates the match as Fair (~35–45%)."},{"id":"212d6091-fe7b-4e74-89be-5b2a0e3cf377","tool":"speechify","tool_name":"Speechify","verdict":"failed","score":null,"score_total":null,"note":"On the noisy source, Speechify produced a clone that diverged strongly from the speaker and even sounded female despite a male input."},{"id":"2fb082d1-14df-479f-8289-a6b6d624a2b4","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"The noisy-source clone barely resembled the original speaker and was described as the weakest voice-match result in the round."},{"id":"97e20d2d-bf92-4a26-b8be-1cbd2f3fd861","tool":"vocalai","tool_name":"VocalAI","verdict":"struggled","score":null,"score_total":null,"note":"The clone preserved only about 10–15% of the original speaker's identity, and the output sounded heavily polished and processed rather than speaker-faithful."}]}