{"observation":{"id":"91aabbbb-5272-4dcf-9cb8-bd369fbfa7a2","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","criterion":"pronunciation-accuracy","criterion_name":"Pronunciation Accuracy","criterion_definition":"How well the tool handles names, technical terms, numbers, acronyms, and other tricky pronunciations.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If a cloned voice cannot handle names, numbers, acronyms, and technical terms correctly, it fails at producing usable voiceover. (3 of 3 judges)","scenario":"high-quality-voice-sample","scenario_name":"High-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","role":"input","filename":"heygen-input-high-quality.wav"}],"stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"],"verdict":"worked","score":null,"score_total":null,"note":"Gen+ keeps the script intelligible on the clean sample despite the gender shift, so pronunciation remains intact.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-highquality-input-6c26a9323e7b.wav","role":"input","alt":"Research media topmediai highquality input.wav"},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-output-high-quality-genplus-5e1ea54f0a66.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","filename":"heygen-input-high-quality.wav","alt":"High-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"]},"tool_page_slug":"topmediai-voice-cloning","tool_url":"https://aidemos.com/tools/topmediai-voice-cloning","permalink":"https://aidemos.com/evidence/91aabbbb-5272-4dcf-9cb8-bd369fbfa7a2","api_url":"https://ai.aidemos.com/v1/observations/91aabbbb-5272-4dcf-9cb8-bd369fbfa7a2"},"peers":[{"id":"a7f33961-1c6f-4dbe-b305-e2aba9be8aa1","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"mixed","score":null,"score_total":null,"note":"Pronunciation is not perfect even when identity matching is strong: the report notes minor pronunciation issues in the clean-voice output.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-input-high-quality-1c369166e815.wav","evidence_url":"https://aidemos.com/evidence/a7f33961-1c6f-4dbe-b305-e2aba9be8aa1"},{"id":"a5510a96-7de7-461e-b9d0-fb42048de7a3","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Pronounces words correctly in the long-form English output, with pronunciation reported as stable across the generated script.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-highquality-input-4dac94a57d04.wav","evidence_url":"https://aidemos.com/evidence/a5510a96-7de7-461e-b9d0-fb42048de7a3"},{"id":"e5f193fd-54a7-4c7b-bcb4-a7615a65bfd6","tool":"fish-audio","tool_name":"Fish Audio","verdict":"mixed","score":null,"score_total":null,"note":"Can still make isolated pronunciation mistakes on clean English input: one high-quality output had a single mispronounced word, about 2–3% of the generation, while the paired output had no such issue.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-high-quality-input-output-1-868896ae57ec.mp3","evidence_url":"https://aidemos.com/evidence/e5f193fd-54a7-4c7b-bcb4-a7615a65bfd6"},{"id":"bec90023-a4d6-4a2d-8587-0e2eb94efd8f","tool":"inworld","tool_name":"Inworld","verdict":"worked","score":null,"score_total":null,"note":"No misread or garbled words were noted in the high-quality pass, and no dedicated pronunciation stress test was run on that scenario.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-highquality-input-5a7af892998c.wav","evidence_url":"https://aidemos.com/evidence/bec90023-a4d6-4a2d-8587-0e2eb94efd8f"},{"id":"f6ef9e8a-c48a-4142-9267-dfec1a4dfc05","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"No misread or garbled words were noted on the clean-input generation.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-output-high-quality-64a65a87ed76.mp3","evidence_url":"https://aidemos.com/evidence/f6ef9e8a-c48a-4142-9267-dfec1a4dfc05"},{"id":"ceae86e2-aaab-4ec2-a167-9f3a6b449892","tool":"speechify","tool_name":"Speechify","verdict":"worked","score":null,"score_total":null,"note":"The clean-output pass also stayed pronunciation-clear, with no flagged mispronounced words.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-speechify-highquality-input-cd6fe714e081.wav","evidence_url":"https://aidemos.com/evidence/ceae86e2-aaab-4ec2-a167-9f3a6b449892"},{"id":"7d5bd84f-3369-49b9-9539-32d3849c8a66","tool":"uberduck","tool_name":"Uberduck","verdict":"worked","score":null,"score_total":null,"note":"English words remained intelligible, and the cleaner source did not change that ceiling on this run.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-high-quality-input-recording-f326629d7802.mp4","evidence_url":"https://aidemos.com/evidence/7d5bd84f-3369-49b9-9539-32d3849c8a66"},{"id":"771a8879-2655-4892-bbf7-9da8f72df423","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The clean-sample run showed no major pronunciation errors.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-output-high-quality-1583ecfff017.wav","evidence_url":"https://aidemos.com/evidence/771a8879-2655-4892-bbf7-9da8f72df423"}],"other_criteria":[{"id":"0e63740f-6710-4e59-896e-d26271ddc32f","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen+ stays stable through long-form generation on the clean sample, with no voice breaks or instability observed.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/0e63740f-6710-4e59-896e-d26271ddc32f"},{"id":"8462742c-8b90-408f-942e-50c73e381675","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD stays consistent across longer scripts on the clean sample, with no noticeable pronunciation issues, interruptions, or degradation.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/8462742c-8b90-408f-942e-50c73e381675"},{"id":"eedd48b7-de42-4024-a225-dfe870fd6e87","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Gen remains consistent through the generated narration on the clean sample, with no interruptions or instability observed.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/eedd48b7-de42-4024-a225-dfe870fd6e87"},{"id":"ff84cb60-0eed-46f9-9f33-90c410ef58f3","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen+ sounds more natural than Gen, but the weak identity preservation keeps the output from feeling fully convincing.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/ff84cb60-0eed-46f9-9f33-90c410ef58f3"},{"id":"13dae80b-2bff-4c46-bed3-b2429db82ad6","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD is the most human-like and realistic output on the clean sample, with better emotional delivery and conversational flow.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/13dae80b-2bff-4c46-bed3-b2429db82ad6"},{"id":"29181eec-a14d-4508-97c6-e0a5e0f9345c","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Gen still sounds robotic on the clean source sample, making it less convincing than HD.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/29181eec-a14d-4508-97c6-e0a5e0f9345c"},{"id":"078f1ee7-15c1-4a69-870b-ecfb34672ef0","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Gen keeps some similarity to the clean source voice, but speaker identity remains only moderately accurate and still sounds robotic.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/078f1ee7-15c1-4a69-870b-ecfb34672ef0"},{"id":"7dd3ad3b-020a-4688-beaa-d24097791ff2","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"HD has the highest similarity to the clean source voice and preserves speaker identity best among the high-quality outputs.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/7dd3ad3b-020a-4688-beaa-d24097791ff2"},{"id":"8167b0ac-f9dd-4ff4-aab8-d1bbda394f44","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Gen+ reduces similarity to the clean source voice by shifting toward a feminine tone, so speaker representation is inaccurate.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/8167b0ac-f9dd-4ff4-aab8-d1bbda394f44"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"a7f33961-1c6f-4dbe-b305-e2aba9be8aa1","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"mixed","score":null,"score_total":null,"note":"Pronunciation is not perfect even when identity matching is strong: the report notes minor pronunciation issues in the clean-voice output."},{"id":"a5510a96-7de7-461e-b9d0-fb42048de7a3","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Pronounces words correctly in the long-form English output, with pronunciation reported as stable across the generated script."},{"id":"e5f193fd-54a7-4c7b-bcb4-a7615a65bfd6","tool":"fish-audio","tool_name":"Fish Audio","verdict":"mixed","score":null,"score_total":null,"note":"Can still make isolated pronunciation mistakes on clean English input: one high-quality output had a single mispronounced word, about 2–3% of the generation, while the paired output had no such issue."},{"id":"bec90023-a4d6-4a2d-8587-0e2eb94efd8f","tool":"inworld","tool_name":"Inworld","verdict":"worked","score":null,"score_total":null,"note":"No misread or garbled words were noted in the high-quality pass, and no dedicated pronunciation stress test was run on that scenario."},{"id":"f6ef9e8a-c48a-4142-9267-dfec1a4dfc05","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"No misread or garbled words were noted on the clean-input generation."},{"id":"ceae86e2-aaab-4ec2-a167-9f3a6b449892","tool":"speechify","tool_name":"Speechify","verdict":"worked","score":null,"score_total":null,"note":"The clean-output pass also stayed pronunciation-clear, with no flagged mispronounced words."},{"id":"7d5bd84f-3369-49b9-9539-32d3849c8a66","tool":"uberduck","tool_name":"Uberduck","verdict":"worked","score":null,"score_total":null,"note":"English words remained intelligible, and the cleaner source did not change that ceiling on this run."},{"id":"771a8879-2655-4892-bbf7-9da8f72df423","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The clean-sample run showed no major pronunciation errors."}]}