{"observation":{"id":"f82441b8-d817-4477-9666-b22eeb247efa","tool":"invideo-ai","tool_name":"InVideo AI","criterion":"output-resolution","criterion_name":"Output resolution","criterion_definition":"Does it maintain the input resolution or downscale?","criterion_evidence_type":"transformation","criterion_rank_role":"context","criterion_rank_role_reason":"Keeping full resolution is important for delivery quality, but it is a downstream output constraint rather than the main background-removal task itself. (3 of 3 judges)","scenario":"indoor-talking-head","scenario_name":"Indoor Talking Head","group_tag":"remove-or-replace-video-backgrounds-using-ai","scenario_description":"An indoor talking-head video of a person speaking directly to the camera with a static background, natural hand gestures, and facial expressions. It was used to test how well tools preserve the subject while replacing an indoor background.","modality":"video","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/remove-or-replace-video-backgrounds-usin-input-2-0e8b7fa1a596.mp4","role":"input","filename":"Input 2.mp4"}],"stresses":["Static camera performance","Face and body segmentation","Facial expression preservation","Background replacement accuracy in an indoor environment"],"verdict":"worked","score":null,"score_total":null,"note":"It produced 4320×7672 output, exceeding the requested 4K vertical ask.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/26e86b41c14c48d394fcd4b01fec4c66.mp4?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/b4ebfbdee34742e792628c2e65916214.mp4?v=1","role":"input","alt":null}],"run_id":"10155228-7ce8-438a-a210-547331ef080b","study_title":"Remove or Replace Video Backgrounds Using AI","study_kind":"generation","research_task":"86ba42c2d","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://d3epheqghktydj.cloudfront.net/remove-or-replace-video-backgrounds-usin-input-2-0e8b7fa1a596.mp4","filename":"Input 2.mp4","alt":"Indoor Talking Head","role":"input"}],"modality":"video","stresses":["Static camera performance","Face and body segmentation","Facial expression preservation","Background replacement accuracy in an indoor environment"]},"tool_page_slug":"invideo-ai","tool_url":"https://aidemos.com/tools/invideo-ai","permalink":"https://aidemos.com/evidence/f82441b8-d817-4477-9666-b22eeb247efa","api_url":"https://ai.aidemos.com/v1/observations/f82441b8-d817-4477-9666-b22eeb247efa"},"peers":[{"id":"b0926e01-7f7c-4c16-839a-6289f9c0cda5","tool":"bria-ai","tool_name":"Bria.ai","verdict":"mixed","score":null,"score_total":null,"note":"The output kept the 3840×2160 frame size, but the subject content is rotated about 90° inside that frame, so the clip is not a clean upright export.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/e18ebb9e427c471d9c6128e260623df1.mp4?v=1","evidence_url":"https://aidemos.com/evidence/b0926e01-7f7c-4c16-839a-6289f9c0cda5"},{"id":"ff772bac-3024-4bd9-954f-5c103d3ed755","tool":"cutout-pro","tool_name":"Cutout.Pro","verdict":"struggled","score":null,"score_total":null,"note":"Downscales the roughly 2160×3840 source to 360×640 and truncates a 15.35s clip to 5.021s.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/7a80c1fcbc4f413c83f0613d6ba48430.webm?v=1","evidence_url":"https://aidemos.com/evidence/ff772bac-3024-4bd9-954f-5c103d3ed755"},{"id":"74e15c82-514d-483b-bce2-15ba2a3e9488","tool":"descript","tool_name":"Descript","verdict":"failed","score":null,"score_total":null,"note":"It converts a 3840×2160 landscape source into a 720×1280 portrait export, changing orientation as well as downscaling.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/36e7879566c64a3aaf2d1d5881a9db11.png?v=1","evidence_url":"https://aidemos.com/evidence/74e15c82-514d-483b-bce2-15ba2a3e9488"},{"id":"8d017e50-e0e9-4f53-a5bd-bea26f8c89ae","tool":"fotor","tool_name":"Fotor","verdict":"failed","score":null,"score_total":null,"note":"A vertical 1080×1920 input was rendered as a 2462×1080 landscape canvas, instead of preserving the source orientation.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/180e2c765f5d48aea4b60b93bc4b1310.mp4?v=1","evidence_url":"https://aidemos.com/evidence/8d017e50-e0e9-4f53-a5bd-bea26f8c89ae"},{"id":"7130bee3-ddf4-4106-9b6b-43d87055e154","tool":"veed","tool_name":"VEED","verdict":"mixed","score":null,"score_total":null,"note":"Downscales a 4K-class talking-head source from effective 2160×3840 to 1080×1920 on export, while preserving the 15.35s duration exactly.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/dfa254fe0189444886a119de2d6e6f1e.mp4?v=1","evidence_url":"https://aidemos.com/evidence/7130bee3-ddf4-4106-9b6b-43d87055e154"}],"other_criteria":[{"id":"3ca44948-0ba2-4ed5-938b-6d3dda14d259","criterion":"edge-quality","criterion_name":"Edge quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It preserved hair and clothing edges cleanly against the replacement set, with no visible halo on close inspection.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/3ca44948-0ba2-4ed5-938b-6d3dda14d259"},{"id":"f19d7ae7-dbee-4cbc-aa17-fca161e36730","criterion":"hair-and-fine-detail","criterion_name":"Hair and fine detail","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"It preserved hands and face detail well, but it distorted a subtle curved glasses highlight into a straight, disconnected bar floating above the frame line.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/f19d7ae7-dbee-4cbc-aa17-fca161e36730"},{"id":"cf321f3a-d418-4ea0-8862-5187c75b73ea","criterion":"temporal-consistency","criterion_name":"Temporal consistency","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The generated studio background stayed stable across the clip, with identical shelf positions and colors at 0:03, 0:08, and 0:13 and no visible flicker or drift.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/cf321f3a-d418-4ea0-8862-5187c75b73ea"}],"appears_in":[{"page_type":"ranking","slug":"video-background-removers","title":"Best AI Tools to Remove or Replace Video Backgrounds with AI","url":"https://aidemos.com/best/video-background-removers","binding":"run"}],"same_scenario":[{"id":"b0926e01-7f7c-4c16-839a-6289f9c0cda5","tool":"bria-ai","tool_name":"Bria.ai","verdict":"mixed","score":null,"score_total":null,"note":"The output kept the 3840×2160 frame size, but the subject content is rotated about 90° inside that frame, so the clip is not a clean upright export."},{"id":"ff772bac-3024-4bd9-954f-5c103d3ed755","tool":"cutout-pro","tool_name":"Cutout.Pro","verdict":"struggled","score":null,"score_total":null,"note":"Downscales the roughly 2160×3840 source to 360×640 and truncates a 15.35s clip to 5.021s."},{"id":"74e15c82-514d-483b-bce2-15ba2a3e9488","tool":"descript","tool_name":"Descript","verdict":"failed","score":null,"score_total":null,"note":"It converts a 3840×2160 landscape source into a 720×1280 portrait export, changing orientation as well as downscaling."},{"id":"8d017e50-e0e9-4f53-a5bd-bea26f8c89ae","tool":"fotor","tool_name":"Fotor","verdict":"failed","score":null,"score_total":null,"note":"A vertical 1080×1920 input was rendered as a 2462×1080 landscape canvas, instead of preserving the source orientation."},{"id":"7130bee3-ddf4-4106-9b6b-43d87055e154","tool":"veed","tool_name":"VEED","verdict":"mixed","score":null,"score_total":null,"note":"Downscales a 4K-class talking-head source from effective 2160×3840 to 1080×1920 on export, while preserving the 15.35s duration exactly."}]}