{"observation":{"id":"1e1e3fce-4d6f-4c95-9b95-f418b907d759","tool":"imagineart","tool_name":"ImagineArt","criterion":"expression-accuracy","criterion_name":"Expression accuracy","criterion_definition":"Whether the emotional tone and facial expression match what was explicitly prompted.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the character’s intended emotion or facial expression does not match the prompt, the generated character is not being controlled reliably across scenes. (3 of 3 judges)","scenario":"full-frontal-portrait","scenario_name":"Full frontal portrait","group_tag":null,"scenario_description":"Full frontal portrait reference image with fair skin, curly dark hair, bindi, gold jhumka earrings, and a green stone necklace. All features are clearly visible in good natural lighting, making it the easiest identity anchor for the tools.","modality":"image","input_text":null,"input_artifact_refs":[],"stresses":["Baseline identity preservation","Accessory retention","Best-case frontal face matching","Consistent character reuse across varied scenes"],"verdict":"failed","score":null,"score_total":null,"note":"The desert-horse output does not match the prompted intensity and determination; the expression is neutral and composed instead.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/imagineart-imagineart-input1-horseride-79c6e9aedd61.jpg","role":"output","alt":null}],"run_id":"1dfb8fa4-f7f0-47dc-911b-7db3af467e9a","study_title":"Generate Consistent AI Characters Across Different Scenes and Poses","study_kind":"generation","research_task":"86b96df11","tested_at":null,"completeness":"output-only","input":{"state":"not-captured","text":null,"files":[],"modality":"image","stresses":["Baseline identity preservation","Accessory retention","Best-case frontal face matching","Consistent character reuse across varied scenes"]},"tool_page_slug":"imagineart","tool_url":"https://aidemos.com/tools/imagineart","permalink":"https://aidemos.com/evidence/1e1e3fce-4d6f-4c95-9b95-f418b907d759","api_url":"https://ai.aidemos.com/v1/observations/1e1e3fce-4d6f-4c95-9b95-f418b907d759"},"peers":[{"id":"6636ced6-0dd3-4d55-aed4-228102bc03f6","tool":"chatgpt","tool_name":"ChatGPT","verdict":"failed","score":null,"score_total":null,"note":"Misses the prompted brave/determined emotional tone and instead outputs a soft neutral expression, leaving the scene with essentially no emotional alignment.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/chatgpt-chatgpt-input1-horseride-dbbcd9f8517e.png","evidence_url":"https://aidemos.com/evidence/6636ced6-0dd3-4d55-aed4-228102bc03f6"},{"id":"d08a8783-fc20-4cf5-8b89-47e25f9f73b7","tool":"gemini","tool_name":"Gemini","verdict":"worked","score":null,"score_total":null,"note":"The tool can capture a prompted angry and guarded mood, with direct eye contact and an intense expression in the interrogation shot.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/gemini-input-1-abf743cbaaf0.png","evidence_url":"https://aidemos.com/evidence/d08a8783-fc20-4cf5-8b89-47e25f9f73b7"},{"id":"b57aa471-7b05-4433-a4f2-f1a621c5aaf3","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"It does not translate an explicitly angry or guarded prompt into facial expression, defaulting to calm neutrality.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/leonardo-ai-leonardo-input1-interrogation-77ba6593d6a0.jpg","evidence_url":"https://aidemos.com/evidence/b57aa471-7b05-4433-a4f2-f1a621c5aaf3"},{"id":"6a1796e6-76aa-4a88-84d2-ba870e97737e","tool":"scenario","tool_name":"Scenario","verdict":"failed","score":null,"score_total":null,"note":"It fails to map an angry or guarded prompt onto the face; the output stays neutral or calm and even reads with a slight smile.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/scenario-scenario-input1-interrogation-8a626f0b95c8.png","evidence_url":"https://aidemos.com/evidence/6a1796e6-76aa-4a88-84d2-ba870e97737e"}],"other_criteria":[{"id":"68d74788-e0d6-4548-b1b3-483e48229716","criterion":"accessory-and-detail-retention","criterion_name":"Accessory & detail retention","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The warm-cafe output changes a specific facial detail: the eye colour shifts from black in the reference to brown in the output.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/68d74788-e0d6-4548-b1b3-483e48229716"},{"id":"6650d49c-d987-49c9-861d-9cd900685ae5","criterion":"accessory-and-detail-retention","criterion_name":"Accessory & detail retention","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output retains the bindi and tight back hair correctly, and the report says harsh lighting reveals more skin texture with less over-smoothing than the other Input 1 scenes.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/6650d49c-d987-49c9-861d-9cd900685ae5"},{"id":"669f1e69-857a-4f41-88cc-29961c0ee861","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output is the strongest Input 1 match, with the eyes, nose, lips, and face shape closest to the reference across the Input 1 scenes.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/669f1e69-857a-4f41-88cc-29961c0ee861"},{"id":"a9483897-1a64-47e5-ac48-2e943f7a2286","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"The desert-horse output shows clear identity drift, with the eye shape, nose structure, jawline, and face proportions all differing from the reference.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/a9483897-1a64-47e5-ac48-2e943f7a2286"},{"id":"e43e297f-c4bf-436b-85b8-14f273cd82dd","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The warm-cafe output keeps the subject recognisable but not exact; the report says eye colour shifts from black to brown and facial proportions are slightly altered.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/e43e297f-c4bf-436b-85b8-14f273cd82dd"},{"id":"f5b274d6-3bb7-4d09-a3de-66be170e2054","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The warm-cafe output follows the prompt well, with the café environment, window lighting, outfit, and hairstyle all rendered correctly.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/f5b274d6-3bb7-4d09-a3de-66be170e2054"},{"id":"a763eca8-4dde-4cee-8542-14e894e45257","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The desert-horse output follows the scene prompt well, including the black horse, desert setting, riding outfit, scarf, braided hair, and dynamic motion.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/a763eca8-4dde-4cee-8542-14e894e45257"},{"id":"99d02e39-0b2c-4869-b966-fb655c676bfc","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The interrogation-room output follows the harsh lighting, serious expression, formal clothing, tight back hair, and bindi, but it shows less body than the prompt specified.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/99d02e39-0b2c-4869-b966-fb655c676bfc"}],"appears_in":[{"page_type":"ranking","slug":"consistent-ai-characters","title":"Best AI Tools for Consistent AI Characters Across Scenes and Poses","url":"https://aidemos.com/best/consistent-ai-characters","binding":"run"}],"same_scenario":[{"id":"6636ced6-0dd3-4d55-aed4-228102bc03f6","tool":"chatgpt","tool_name":"ChatGPT","verdict":"failed","score":null,"score_total":null,"note":"Misses the prompted brave/determined emotional tone and instead outputs a soft neutral expression, leaving the scene with essentially no emotional alignment."},{"id":"d08a8783-fc20-4cf5-8b89-47e25f9f73b7","tool":"gemini","tool_name":"Gemini","verdict":"worked","score":null,"score_total":null,"note":"The tool can capture a prompted angry and guarded mood, with direct eye contact and an intense expression in the interrogation shot."},{"id":"b57aa471-7b05-4433-a4f2-f1a621c5aaf3","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"It does not translate an explicitly angry or guarded prompt into facial expression, defaulting to calm neutrality."},{"id":"6a1796e6-76aa-4a88-84d2-ba870e97737e","tool":"scenario","tool_name":"Scenario","verdict":"failed","score":null,"score_total":null,"note":"It fails to map an angry or guarded prompt onto the face; the output stays neutral or calm and even reads with a slight smile."}]}