{"observation":{"id":"3e841b26-1213-464f-a700-cbcad3e2d9d4","tool":"heygen","tool_name":"Heygen","criterion":"naturalness","criterion_name":"Naturalness & Human Quality","criterion_definition":"How realistic and human-like the output sounds, including pacing, breathing, and micro-pauses.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"A voiceover tool has to sound human and listenable; unnatural pacing or robotic delivery undermines the main job. (3 of 3 judges)","scenario":"high-quality-voice-sample","scenario_name":"High-Quality Voice Sample","group_tag":"voice-cloning","scenario_description":"A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.","modality":"mixed","input_text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","role":"input","filename":"heygen-input-high-quality.wav"}],"stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"],"verdict":"mixed","score":null,"score_total":null,"note":"The first high-quality clone was around 80% human-like, but it still retained synthetic characteristics.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-high-quality-variant-1-59bbb4407317.wav","role":"output","alt":null},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-output-high-quality-variant-3-bes-3b44d7550666.wav","role":"output","alt":null}],"run_id":"46222c41-0046-41cc-bfaa-5f7ba6aa4933","study_title":"Clone Your Voice and Generate Voiceover from Text","study_kind":"generation","research_task":"86ba42bx1","tested_at":null,"completeness":"input-and-output","input":{"state":"text-and-files","text":"Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.","files":[{"url":"https://d3epheqghktydj.cloudfront.net/research-media-heygen-input-high-quality-884aedc4e603.wav","filename":"heygen-input-high-quality.wav","alt":"High-Quality Voice Sample","role":"input"}],"modality":"mixed","stresses":["Maximum voice-cloning accuracy","Naturalness with optimal source quality","Long-form consistency","Pronunciation stability","Voice preservation under ideal conditions"]},"tool_page_slug":"heygen","tool_url":"https://aidemos.com/tools/heygen","permalink":"https://aidemos.com/evidence/3e841b26-1213-464f-a700-cbcad3e2d9d4","api_url":"https://ai.aidemos.com/v1/observations/3e841b26-1213-464f-a700-cbcad3e2d9d4"},"peers":[{"id":"d6ee67e8-0fca-4af9-a652-d06892592855","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"The output sounds highly natural, with human-like delivery and speech rhythm, and the report says it is suitable for short-form content generation.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-input-high-quality-1c369166e815.wav","evidence_url":"https://aidemos.com/evidence/d6ee67e8-0fca-4af9-a652-d06892592855"},{"id":"b37ecb0b-4316-4f1c-a17f-6f589943e44a","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Sounds smoother and more human-like than the noisy-source run, though pacing inconsistencies still appear with occasional fast and slow delivery.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-elevenlabs-highquality-input-4dac94a57d04.wav","evidence_url":"https://aidemos.com/evidence/b37ecb0b-4316-4f1c-a17f-6f589943e44a"},{"id":"4257691f-67f2-4ab1-beb7-c82a2b63cab9","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Produces natural-sounding delivery on clean input: the first high-quality output had genuinely human pitch, pauses, and overall delivery, and the second was reported as equally human-like.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-high-quality-input-output-1-868896ae57ec.mp3","evidence_url":"https://aidemos.com/evidence/4257691f-67f2-4ab1-beb7-c82a2b63cab9"},{"id":"0c7ae9e9-dc06-4aeb-af49-4664697fc567","tool":"inworld","tool_name":"Inworld","verdict":"worked","score":null,"score_total":null,"note":"The high-quality output landed at roughly 70–80% human-like and was judged to have solid naturalness overall.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-inworld-highquality-input-5a7af892998c.wav","evidence_url":"https://aidemos.com/evidence/0c7ae9e9-dc06-4aeb-af49-4664697fc567"},{"id":"840cd150-37f6-4b15-9fd5-6e04d13282bd","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"The clean-input output remained human-like and good overall, though it still sounded softer than the source speaker.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-minimax-output-high-quality-64a65a87ed76.mp3","evidence_url":"https://aidemos.com/evidence/840cd150-37f6-4b15-9fd5-6e04d13282bd"},{"id":"431959bc-0829-4537-91cb-df92cd43a08c","tool":"speechify","tool_name":"Speechify","verdict":"mixed","score":70.0,"score_total":100.0,"note":"The clean source sounded fairly natural, but the researcher still heard noticeable robotic coloration, estimating it at about 70% natural and 30% AI-sounding.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-speechify-highquality-input-cd6fe714e081.wav","evidence_url":"https://aidemos.com/evidence/431959bc-0829-4537-91cb-df92cd43a08c"},{"id":"29181eec-a14d-4508-97c6-e0a5e0f9345c","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"struggled","score":null,"score_total":null,"note":"Gen still sounds robotic on the clean source sample, making it less convincing than HD.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-topmediai-highquality-input-6c26a9323e7b.wav","evidence_url":"https://aidemos.com/evidence/29181eec-a14d-4508-97c6-e0a5e0f9345c"},{"id":"36d63d33-707c-4265-95d6-5bbea891d148","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"The cleaner input still produced a robotic delivery with awkward pauses breaking up the speech.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-high-quality-input-recording-f326629d7802.mp4","evidence_url":"https://aidemos.com/evidence/36d63d33-707c-4265-95d6-5bbea891d148"},{"id":"ef698e18-aed2-4143-8946-47bbd6e3ea4b","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The output sounded smooth and pleasant, with the report estimating it at about 70–80% human-like.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/research-media-vocalai-output-high-quality-1583ecfff017.wav","evidence_url":"https://aidemos.com/evidence/ef698e18-aed2-4143-8946-47bbd6e3ea4b"}],"other_criteria":[{"id":"8ff9b798-baad-4160-8b98-7a90b05a1ca0","criterion":"long-form-consistency","criterion_name":"Long-Form Consistency","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"The longer-script high-quality generation had flow interruptions, word mispronunciations, inconsistent delivery, and the report says multiple regenerations may be required for production-ready results.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/8ff9b798-baad-4160-8b98-7a90b05a1ca0"},{"id":"c208f7da-4963-4cfc-a24d-dc6532da9051","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"The second high-quality clone showed lower accuracy because of a noticeable vocal shift and was described as sounding closer to a female voice profile.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/c208f7da-4963-4cfc-a24d-dc6532da9051"},{"id":"26db4b2a-85a8-4bc5-a1e2-4211467cc2ca","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The first high-quality clone achieved only moderate similarity to the original voice, so it was usable but not a strong identity match.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/26db4b2a-85a8-4bc5-a1e2-4211467cc2ca"},{"id":"27d76bf7-8e01-4067-ab1a-0aaa49bec331","criterion":"voice-match-accuracy","criterion_name":"Voice Match Accuracy","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The best high-quality clone still reached only about 70% similarity to the original voice, so it remained below a strong match despite being the best of the three.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/27d76bf7-8e01-4067-ab1a-0aaa49bec331"}],"appears_in":[{"page_type":"ranking","slug":"voice-cloning-tools","title":"Best AI Tools to Clone Your Voice and Generate Voiceovers from Text","url":"https://aidemos.com/best/voice-cloning-tools","binding":"run"}],"same_scenario":[{"id":"d6ee67e8-0fca-4af9-a652-d06892592855","tool":"aiclonevoicefree-com","tool_name":"AICloneVoiceFree.com","verdict":"worked","score":null,"score_total":null,"note":"The output sounds highly natural, with human-like delivery and speech rhythm, and the report says it is suitable for short-form content generation."},{"id":"b37ecb0b-4316-4f1c-a17f-6f589943e44a","tool":"elevenlabs","tool_name":"ElevenLabs","verdict":"worked","score":null,"score_total":null,"note":"Sounds smoother and more human-like than the noisy-source run, though pacing inconsistencies still appear with occasional fast and slow delivery."},{"id":"4257691f-67f2-4ab1-beb7-c82a2b63cab9","tool":"fish-audio","tool_name":"Fish Audio","verdict":"worked","score":null,"score_total":null,"note":"Produces natural-sounding delivery on clean input: the first high-quality output had genuinely human pitch, pauses, and overall delivery, and the second was reported as equally human-like."},{"id":"0c7ae9e9-dc06-4aeb-af49-4664697fc567","tool":"inworld","tool_name":"Inworld","verdict":"worked","score":null,"score_total":null,"note":"The high-quality output landed at roughly 70–80% human-like and was judged to have solid naturalness overall."},{"id":"840cd150-37f6-4b15-9fd5-6e04d13282bd","tool":"minimax","tool_name":"MiniMax","verdict":"worked","score":null,"score_total":null,"note":"The clean-input output remained human-like and good overall, though it still sounded softer than the source speaker."},{"id":"431959bc-0829-4537-91cb-df92cd43a08c","tool":"speechify","tool_name":"Speechify","verdict":"mixed","score":70.0,"score_total":100.0,"note":"The clean source sounded fairly natural, but the researcher still heard noticeable robotic coloration, estimating it at about 70% natural and 30% AI-sounding."},{"id":"29181eec-a14d-4508-97c6-e0a5e0f9345c","tool":"topmediai-voice-cloning","tool_name":"TopMediai Voice Cloning","verdict":"struggled","score":null,"score_total":null,"note":"Gen still sounds robotic on the clean source sample, making it less convincing than HD."},{"id":"36d63d33-707c-4265-95d6-5bbea891d148","tool":"uberduck","tool_name":"Uberduck","verdict":"failed","score":null,"score_total":null,"note":"The cleaner input still produced a robotic delivery with awkward pauses breaking up the speech."},{"id":"ef698e18-aed2-4143-8946-47bbd6e3ea4b","tool":"vocalai","tool_name":"VocalAI","verdict":"worked","score":null,"score_total":null,"note":"The output sounded smooth and pleasant, with the report estimating it at about 70–80% human-like."}]}