{"observation":{"id":"078f1ee7-15c1-4a69-870b-ecfb34672ef0","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","criterion_definition":"Whether the generated voice actually sounds like the original speaker, including tone, pitch, rhythm, and identity.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"This ranking is fundamentally about whether the generated speech still sounds like the target speaker, so identity match is core. (3 of 3 judges)","scenario":"high-quality-voice-sample","scenario_name":"High-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","role":"input","filename":"heygen-input-high-quality.wav"}],"stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"],"verdict":"mixed","score":null,"score_total":null,"note":"Gen keeps some similarity to the clean source voice, but speaker identity remains only moderately accurate and still sounds robotic.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-highquality-input-6c26a9323e7b.wav","role":"input","alt":null},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-output-high-quality-gen-7f8d16f66f7a.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","filename":"heygen-input-high-quality.wav","alt":"High-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"]},"tool_page_slug":"topmediai-voice-cloning","tool_url":"https://aidemos.com/tools/topmediai-voice-cloning","permalink":"https://aidemos.com/evidence/078f1ee7-15c1-4a69-870b-ecfb34672ef0","api_url":"https://ai.aidemos.com/v1/observations/078f1ee7-15c1-4a69-870b-ecfb34672ef0"},"peers":[{"id":"e536c0ef-9074-43e8-998e-b304828aeb04","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"The cloned voice closely matches the clean source speaker, with the report calling the match about 95% and saying tone and vocal characteristics were closely preserved.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-input-high-quality-1c369166e815.wav","evidence_url":"https://aidemos.com/evidence/e536c0ef-9074-43e8-998e-b304828aeb04"},{"id":"26a0a7d9-d7d9-4bf8-b763-04b96a9ee7fc","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"mixed","score":null,"score_total":null,"note":"Improves source resemblance only modestly: the report rates the clone at approximately 40–50% similarity, says the voice remains noticeably polished and processed, and notes that speaker identity is only partially preserved.","artifact_count":8,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-highquality-input-4dac94a57d04.wav","evidence_url":"https://aidemos.com/evidence/26a0a7d9-d7d9-4bf8-b763-04b96a9ee7fc"},{"id":"8cbbaf13-e580-4da0-b7b6-f8d82727a9a9","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Recreates a clean studio voice with full identity match: the first high-quality output was described as fully matching the uploaded input, and the second stayed equally strong.","artifact_count":4,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-fishaudio-highquality-input-37844222acf0.wav","evidence_url":"https://aidemos.com/evidence/8cbbaf13-e580-4da0-b7b6-f8d82727a9a9"},{"id":"26db4b2a-85a8-4bc5-a1e2-4211467cc2ca","tool":"heygen","tool_name":"Heygen","verdict":"mixed","score":null,"score_total":null,"note":"The first high-quality clone achieved only moderate similarity to the original voice, so it was usable but not a strong identity match.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-high-quality-variant-1-59bbb4407317.wav","evidence_url":"https://aidemos.com/evidence/26db4b2a-85a8-4bc5-a1e2-4211467cc2ca"},{"id":"14343c6c-c79f-4fc3-9d1e-fc972c929209","tool":"inworld","tool_name":"Inworld","verdict":"mixed","score":null,"score_total":null,"note":"On the clean studio sample, the clone did not match the source precisely and sounded more polished and balanced, suggesting the model optimized for pleasant delivery over exact identity.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-highquality-input-5a7af892998c.wav","evidence_url":"https://aidemos.com/evidence/14343c6c-c79f-4fc3-9d1e-fc972c929209"},{"id":"b9f0a7be-3871-41aa-bb76-0d5c5562c30c","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"Even with a clean source recording, matching accuracy still was not fully there; the report rates this run as Fair (~35–45%) and says the output leaned softer than the original voice.","artifact_count":7,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-highquality-input-058115a62ccc.wav","evidence_url":"https://aidemos.com/evidence/b9f0a7be-3871-41aa-bb76-0d5c5562c30c"},{"id":"a928a78c-e71f-4cc2-9000-714d3622ff79","tool":"speechify","tool_name":"Speechify","verdict":"struggled","score":40.0,"score_total":100.0,"note":"The clean source only matched modestly, at about 40% similarity, with roughly 60% of the output sounding noticeably different from the original.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-speechify-highquality-input-cd6fe714e081.wav","evidence_url":"https://aidemos.com/evidence/a928a78c-e71f-4cc2-9000-714d3622ff79"},{"id":"17d6e619-374f-40bd-af46-c8d70522df45","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"Cleaner source audio did not materially improve identity matching; the clone still did not closely track the original voice.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-high-quality-input-recording-f326629d7802.mp4","evidence_url":"https://aidemos.com/evidence/17d6e619-374f-40bd-af46-c8d70522df45"},{"id":"e71e3f67-5a4f-46d3-acea-6bac6c6ceb8f","tool":"vocalai","tool_name":"VocalAI","verdict":"struggled","score":null,"score_total":null,"note":"Even with a clean studio sample, similarity stayed at only about 10–15%, so the cleaner input produced minimal improvement in identity preservation.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-highquality-input-d2d56f461f81.wav","evidence_url":"https://aidemos.com/evidence/e71e3f67-5a4f-46d3-acea-6bac6c6ceb8f"}],"other_criteria":[{"id":"0e63740f-6710-4e59-896e-d26271ddc32f","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen+ stays stable through long-form generation on the clean sample, with no voice breaks or instability observed.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/0e63740f-6710-4e59-896e-d26271ddc32f"},{"id":"8462742c-8b90-408f-942e-50c73e381675","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD stays consistent across longer scripts on the clean sample, with no noticeable pronunciation issues, interruptions, or degradation.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/8462742c-8b90-408f-942e-50c73e381675"},{"id":"eedd48b7-de42-4024-a225-dfe870fd6e87","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen remains consistent through the generated narration on the clean sample, with no interruptions or instability observed.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/eedd48b7-de42-4024-a225-dfe870fd6e87"},{"id":"29181eec-a14d-4508-97c6-e0a5e0f9345c","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Gen still sounds robotic on the clean source sample, making it less convincing than HD.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/29181eec-a14d-4508-97c6-e0a5e0f9345c"},{"id":"ff84cb60-0eed-46f9-9f33-90c410ef58f3","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen+ sounds more natural than Gen, but the weak identity preservation keeps the output from feeling fully convincing.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/ff84cb60-0eed-46f9-9f33-90c410ef58f3"},{"id":"13dae80b-2bff-4c46-bed3-b2429db82ad6","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD is the most human-like and realistic output on the clean sample, with better emotional delivery and conversational flow.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/13dae80b-2bff-4c46-bed3-b2429db82ad6"},{"id":"b62a2ccb-2b14-4585-b4de-0b510c9887c2","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen keeps the script intelligible on the clean sample; the robotic character is a delivery issue, not misread words.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/b62a2ccb-2b14-4585-b4de-0b510c9887c2"},{"id":"91aabbbb-5272-4dcf-9cb8-bd369fbfa7a2","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen+ keeps the script intelligible on the clean sample despite the gender shift, so pronunciation remains intact.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/91aabbbb-5272-4dcf-9cb8-bd369fbfa7a2"},{"id":"d72cf668-6291-4489-b905-cbb5deacabb6","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD has no misread or garbled words on the clean sample and is the cleanest pronunciation result across all nine English generations tested.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/d72cf668-6291-4489-b905-cbb5deacabb6"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"e536c0ef-9074-43e8-998e-b304828aeb04","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"The cloned voice closely matches the clean source speaker, with the report calling the match about 95% and saying tone and vocal characteristics were closely preserved."},{"id":"26a0a7d9-d7d9-4bf8-b763-04b96a9ee7fc","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"mixed","score":null,"score_total":null,"note":"Improves source resemblance only modestly: the report rates the clone at approximately 40–50% similarity, says the voice remains noticeably polished and processed, and notes that speaker identity is only partially preserved."},{"id":"8cbbaf13-e580-4da0-b7b6-f8d82727a9a9","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Recreates a clean studio voice with full identity match: the first high-quality output was described as fully matching the uploaded input, and the second stayed equally strong."},{"id":"26db4b2a-85a8-4bc5-a1e2-4211467cc2ca","tool":"heygen","tool_name":"Heygen","verdict":"mixed","score":null,"score_total":null,"note":"The first high-quality clone achieved only moderate similarity to the original voice, so it was usable but not a strong identity match."},{"id":"14343c6c-c79f-4fc3-9d1e-fc972c929209","tool":"inworld","tool_name":"Inworld","verdict":"mixed","score":null,"score_total":null,"note":"On the clean studio sample, the clone did not match the source precisely and sounded more polished and balanced, suggesting the model optimized for pleasant delivery over exact identity."},{"id":"b9f0a7be-3871-41aa-bb76-0d5c5562c30c","tool":"minimax","tool_name":"MiniMax","verdict":"mixed","score":null,"score_total":null,"note":"Even with a clean source recording, matching accuracy still was not fully there; the report rates this run as Fair (~35–45%) and says the output leaned softer than the original voice."},{"id":"a928a78c-e71f-4cc2-9000-714d3622ff79","tool":"speechify","tool_name":"Speechify","verdict":"struggled","score":40.0,"score_total":100.0,"note":"The clean source only matched modestly, at about 40% similarity, with roughly 60% of the output sounding noticeably different from the original."},{"id":"17d6e619-374f-40bd-af46-c8d70522df45","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"Cleaner source audio did not materially improve identity matching; the clone still did not closely track the original voice."},{"id":"e71e3f67-5a4f-46d3-acea-6bac6c6ceb8f","tool":"vocalai","tool_name":"VocalAI","verdict":"struggled","score":null,"score_total":null,"note":"Even with a clean studio sample, similarity stayed at only about 10–15%, so the cleaner input produced minimal improvement in identity preservation."}]}