{"observation":{"id":"9f6dd6f3-855d-4f0a-9535-a5219634595e","tool":"imagineart","tool_name":"ImagineArt","criterion":"expression-accuracy","criterion_name":"Expression accuracy","criterion_definition":"Whether the emotional tone and facial expression match what was explicitly prompted.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the character’s intended emotion or facial expression does not match the prompt, the generated character is not being controlled reliably across scenes. (3 of 3 judges)","scenario":"three-quarter-face-portrait","scenario_name":"Three-quarter face portrait","group_tag":null,"scenario_description":"Three-quarter face reference image with medium-dark skin, tight curly hair in an updo, bindi, and a floral dress. The partial angle and softer lighting make it a harder reference than the frontal portrait and are meant to stress identity consistency.","modality":"image","input_text":null,"input_artifact_refs":[],"stresses":["Identity preservation with partial face angle","Hair texture and updo retention","Skin tone fidelity","Reference-image difficulty under softer lighting"],"verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output preserves the reference expression well.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/imagineart-imagineart-input2-interrogation-ce155eab95e8.jpg","role":"output","alt":null}],"run_id":"1dfb8fa4-f7f0-47dc-911b-7db3af467e9a","study_title":"Generate Consistent AI Characters Across Different Scenes and Poses","study_kind":"generation","research_task":"86b96df11","tested_at":null,"completeness":"output-only","input":{"state":"not-captured","text":null,"files":[],"modality":"image","stresses":["Identity preservation with partial face angle","Hair texture and updo retention","Skin tone fidelity","Reference-image difficulty under softer lighting"]},"tool_page_slug":"imagineart","tool_url":"https://aidemos.com/tools/imagineart","permalink":"https://aidemos.com/evidence/9f6dd6f3-855d-4f0a-9535-a5219634595e","api_url":"https://ai.aidemos.com/v1/observations/9f6dd6f3-855d-4f0a-9535-a5219634595e"},"peers":[{"id":"8b496d76-afd9-4658-b0a0-2b73533010ff","tool":"chatgpt","tool_name":"ChatGPT","verdict":"worked","score":null,"score_total":null,"note":"Delivers the requested cold, guarded, unsmiling expression with the exact vibe the prompt asked for.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/chatgpt-chatgpt-input2-interrogation-c78a75b73445.png","evidence_url":"https://aidemos.com/evidence/8b496d76-afd9-4658-b0a0-2b73533010ff"},{"id":"7d21adab-f4e6-4f0a-9975-37fabe9dc1c2","tool":"gemini","tool_name":"Gemini","verdict":"failed","score":null,"score_total":null,"note":"The tool can miss the requested emotion entirely, producing a neutral and emotionless face instead of the prompted angry and guarded look.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/gemini-input-2-1c4e066a03c3.png","evidence_url":"https://aidemos.com/evidence/7d21adab-f4e6-4f0a-9975-37fabe9dc1c2"},{"id":"ca959663-ea8a-4b8b-9cee-60d17e3305d5","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"The same neutral-expression failure repeats on a second reference, so the tool still misses the prompt's angry, guarded tone.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/leonardo-ai-image-95c2a8e685d7.png","evidence_url":"https://aidemos.com/evidence/ca959663-ea8a-4b8b-9cee-60d17e3305d5"}],"other_criteria":[{"id":"a92e28aa-2d3c-4c18-a0dd-fe683af1dfde","criterion":"accessory-and-detail-retention","criterion_name":"Accessory & detail retention","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"The interrogation-room output flattens the dense curly hair substantially and makes the eyebrows look shorter and slightly uneven.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/a92e28aa-2d3c-4c18-a0dd-fe683af1dfde"},{"id":"0c32686e-670a-4740-98f0-c7b63a4fa0dd","criterion":"accessory-and-detail-retention","criterion_name":"Accessory & detail retention","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"The market output lightens the skin tone, removes visible skin marks and scars through smoothing, and makes the lips look less full and more evenly coloured.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/0c32686e-670a-4740-98f0-c7b63a4fa0dd"},{"id":"21c5b413-356f-4640-b66f-fe29ed449e6c","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output stays close to the reference in face shape, skin tone, expression, posture, and overall appearance.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/21c5b413-356f-4640-b66f-fe29ed449e6c"},{"id":"5ec13c6e-a74e-4f10-abd3-dc5f4dfb070c","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The market output keeps the subject reasonably close overall, but the report still calls the identity match moderate because the skin tone shifts and other facial details change slightly.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/5ec13c6e-a74e-4f10-abd3-dc5f4dfb070c"},{"id":"a16d1328-7e31-4ecf-ad70-5f32d9129b28","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output follows the clothing, lighting, pose, and environment accurately.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/a16d1328-7e31-4ecf-ad70-5f32d9129b28"},{"id":"ee4d502e-3819-4940-8013-0e2217e9d1d7","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The market output renders the market environment realistically and follows the prompt well.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/ee4d502e-3819-4940-8013-0e2217e9d1d7"}],"appears_in":[{"page_type":"ranking","slug":"consistent-ai-characters","title":"Best AI Tools for Consistent AI Characters Across Scenes and Poses","url":"https://aidemos.com/best/consistent-ai-characters","binding":"run"}],"same_scenario":[{"id":"8b496d76-afd9-4658-b0a0-2b73533010ff","tool":"chatgpt","tool_name":"ChatGPT","verdict":"worked","score":null,"score_total":null,"note":"Delivers the requested cold, guarded, unsmiling expression with the exact vibe the prompt asked for."},{"id":"7d21adab-f4e6-4f0a-9975-37fabe9dc1c2","tool":"gemini","tool_name":"Gemini","verdict":"failed","score":null,"score_total":null,"note":"The tool can miss the requested emotion entirely, producing a neutral and emotionless face instead of the prompted angry and guarded look."},{"id":"ca959663-ea8a-4b8b-9cee-60d17e3305d5","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"The same neutral-expression failure repeats on a second reference, so the tool still misses the prompt's angry, guarded tone."}]}