{"observation":{"id":"212d6091-fe7b-4e74-89be-5b2a0e3cf377","tool":"speechify","tool_name":"Speechify","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","criterion_definition":"Whether the generated voice actually sounds like the original speaker, including tone, pitch, rhythm, and identity.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"This ranking is fundamentally about whether the generated speech still sounds like the target speaker, so identity match is core. (3 of 3 judges)","scenario":"low-quality-voice-sample","scenario_name":"Low-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A noisy voice recording with background noise, room ambience, and minor disturbances, used to test whether voice-cloning tools can preserve speaker identity when the source audio is imperfect.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","role":"input","filename":"speechify-input-low-quality.wav"}],"stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"],"verdict":"failed","score":null,"score_total":null,"note":"On the noisy source, Speechify produced a clone that diverged strongly from the speaker and even sounded female despite a male input.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-speechify-lowquality-input-86c01eae771e.wav","role":"input","alt":null},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-speechify-output-low-quality-53d2fb249a32.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/22549a3c02994d8b9fe38fff9bfda6a0.wav?v=1","filename":"speechify-input-low-quality.wav","alt":"Low-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Cloning accuracy from degraded audio","Noise and ambience robustness","Speaker identity preservation under poor recording conditions","Distinguishing enhancement from true cloning"]},"tool_page_slug":"speechify","tool_url":"https://aidemos.com/tools/speechify","permalink":"https://aidemos.com/evidence/212d6091-fe7b-4e74-89be-5b2a0e3cf377","api_url":"https://ai.aidemos.com/v1/observations/212d6091-fe7b-4e74-89be-5b2a0e3cf377"},"peers":[{"id":"101630ef-8867-4f0a-adc2-b422996a45c8","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"Voice cloning stays strong even from a noisy source, with the report estimating about 95% similarity to the original speaker and preservation of most vocal characteristics.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-input-low-quality-d1a2d94e8e38.wav","evidence_url":"https://aidemos.com/evidence/101630ef-8867-4f0a-adc2-b422996a45c8"},{"id":"5d946255-1867-40d2-8dc5-17695ecf4038","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"mixed","score":50.0,"score_total":null,"note":"Clones only about half of the source-speaker identity in the noisy sample: the report rates the match at approximately 50% and says the output sounds heavily polished, which lowers resemblance instead of faithfully reproducing the original voice.","artifact_count":6,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-lowquality-input-8230726a9b39.wav","evidence_url":"https://aidemos.com/evidence/5d946255-1867-40d2-8dc5-17695ecf4038"},{"id":"533fcf8b-1d9b-4828-81a2-3cd52e3f90ec","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Preserves speaker identity well even from a noisy source: both low-quality English outputs were described as strong, close matches to the original voice.","artifact_count":4,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-fishaudio-lowquality-input-a57be1fc5dd8.wav","evidence_url":"https://aidemos.com/evidence/533fcf8b-1d9b-4828-81a2-3cd52e3f90ec"},{"id":"5f56a175-1dc1-405b-bb24-ba0c2dae580b","tool":"heygen","tool_name":"Heygen","verdict":"worked","score":null,"score_total":null,"note":"The best low-quality clone reached approximately 95–99% similarity to the original voice and was described as the closest to the original speaker.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-low-quality-variant-3-best-5e23070dc089.wav","evidence_url":"https://aidemos.com/evidence/5f56a175-1dc1-405b-bb24-ba0c2dae580b"},{"id":"90b33678-feed-449c-a618-4f55590cfe75","tool":"inworld","tool_name":"Inworld","verdict":"mixed","score":null,"score_total":null,"note":"On the ~53-second low-quality clone, identity match was strongest at the start and then faded gradually as the clip progressed, so the voice was only partially consistent rather than exact.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-lowquality-input-437133b0167f.wav","evidence_url":"https://aidemos.com/evidence/90b33678-feed-449c-a618-4f55590cfe75"},{"id":"72f96b15-3f48-4e62-9f91-db706af4f1cf","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"With a noisy source recording, the clone only partially preserved speaker identity: the generated voice was noticeably softer than the original and the report rates the match as Fair (~35–45%).","artifact_count":5,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-lowquality-input-6696703e9405.wav","evidence_url":"https://aidemos.com/evidence/72f96b15-3f48-4e62-9f91-db706af4f1cf"},{"id":"8452d11e-1149-4195-9eb0-5817433e2cc4","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"worked","score":null,"score_total":null,"note":"HD is the closest match on the noisy sample and preserves vocal identity best among the three variants.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-lowquality-input-5c3da58b8b17.wav","evidence_url":"https://aidemos.com/evidence/8452d11e-1149-4195-9eb0-5817433e2cc4"},{"id":"2fb082d1-14df-479f-8289-a6b6d624a2b4","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"The noisy-source clone barely resembled the original speaker and was described as the weakest voice-match result in the round.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-low-quality-input-recording-6fd885c280cb.mp4","evidence_url":"https://aidemos.com/evidence/2fb082d1-14df-479f-8289-a6b6d624a2b4"},{"id":"97e20d2d-bf92-4a26-b8be-1cbd2f3fd861","tool":"vocalai","tool_name":"VocalAI","verdict":"struggled","score":null,"score_total":null,"note":"The clone preserved only about 10–15% of the original speaker's identity, and the output sounded heavily polished and processed rather than speaker-faithful.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-lowquality-input-447ab8d8cc52.wav","evidence_url":"https://aidemos.com/evidence/97e20d2d-bf92-4a26-b8be-1cbd2f3fd861"}],"other_criteria":[{"id":"df45fb93-e306-4fbe-a469-741589f4fd2b","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The noisy source still yielded a natural-sounding voice in isolation.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/df45fb93-e306-4fbe-a469-741589f4fd2b"},{"id":"69856bcc-f984-4afc-8701-290d4c80a13e","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The noisy-output pass stayed pronunciation-clear, and the researcher flagged no specific mispronounced words.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/69856bcc-f984-4afc-8701-290d4c80a13e"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"101630ef-8867-4f0a-adc2-b422996a45c8","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"Voice cloning stays strong even from a noisy source, with the report estimating about 95% similarity to the original speaker and preservation of most vocal characteristics."},{"id":"5d946255-1867-40d2-8dc5-17695ecf4038","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"mixed","score":50.0,"score_total":null,"note":"Clones only about half of the source-speaker identity in the noisy sample: the report rates the match at approximately 50% and says the output sounds heavily polished, which lowers resemblance instead of faithfully reproducing the original voice."},{"id":"533fcf8b-1d9b-4828-81a2-3cd52e3f90ec","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Preserves speaker identity well even from a noisy source: both low-quality English outputs were described as strong, close matches to the original voice."},{"id":"5f56a175-1dc1-405b-bb24-ba0c2dae580b","tool":"heygen","tool_name":"Heygen","verdict":"worked","score":null,"score_total":null,"note":"The best low-quality clone reached approximately 95–99% similarity to the original voice and was described as the closest to the original speaker."},{"id":"90b33678-feed-449c-a618-4f55590cfe75","tool":"inworld","tool_name":"Inworld","verdict":"mixed","score":null,"score_total":null,"note":"On the ~53-second low-quality clone, identity match was strongest at the start and then faded gradually as the clip progressed, so the voice was only partially consistent rather than exact."},{"id":"72f96b15-3f48-4e62-9f91-db706af4f1cf","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"With a noisy source recording, the clone only partially preserved speaker identity: the generated voice was noticeably softer than the original and the report rates the match as Fair (~35–45%)."},{"id":"8452d11e-1149-4195-9eb0-5817433e2cc4","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"worked","score":null,"score_total":null,"note":"HD is the closest match on the noisy sample and preserves vocal identity best among the three variants."},{"id":"2fb082d1-14df-479f-8289-a6b6d624a2b4","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"The noisy-source clone barely resembled the original speaker and was described as the weakest voice-match result in the round."},{"id":"97e20d2d-bf92-4a26-b8be-1cbd2f3fd861","tool":"vocalai","tool_name":"VocalAI","verdict":"struggled","score":null,"score_total":null,"note":"The clone preserved only about 10–15% of the original speaker's identity, and the output sounded heavily polished and processed rather than speaker-faithful."}]}