{"observation":{"id":"8b496d76-afd9-4658-b0a0-2b73533010ff","tool":"chatgpt","tool_name":"ChatGPT","criterion":"expression-accuracy","criterion_name":"Expression accuracy","criterion_definition":"Whether the emotional tone and facial expression match what was explicitly prompted.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the character’s intended emotion or facial expression does not match the prompt, the generated character is not being controlled reliably across scenes. (3 of 3 judges)","scenario":"three-quarter-face-portrait","scenario_name":"Three-quarter face portrait","group_tag":null,"scenario_description":"Three-quarter face reference image with medium-dark skin, tight curly hair in an updo, bindi, and a floral dress. The partial angle and softer lighting make it a harder reference than the frontal portrait and are meant to stress identity consistency.","modality":"image","input_text":null,"input_artifact_refs":[],"stresses":["Identity preservation with partial face angle","Hair texture and updo retention","Skin tone fidelity","Reference-image difficulty under softer lighting"],"verdict":"worked","score":null,"score_total":null,"note":"Delivers the requested cold, guarded, unsmiling expression with the exact vibe the prompt asked for.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/chatgpt-chatgpt-input2-interrogation-c78a75b73445.png","role":"output","alt":null}],"run_id":"1dfb8fa4-f7f0-47dc-911b-7db3af467e9a","study_title":"Generate Consistent AI Characters Across Different Scenes and Poses","study_kind":"generation","research_task":"86b96df11","tested_at":null,"completeness":"output-only","input":{"state":"not-captured","text":null,"files":[],"modality":"image","stresses":["Identity preservation with partial face angle","Hair texture and updo retention","Skin tone fidelity","Reference-image difficulty under softer lighting"]},"tool_page_slug":"chatgpt","tool_url":"https://aidemos.com/tools/chatgpt","permalink":"https://aidemos.com/evidence/8b496d76-afd9-4658-b0a0-2b73533010ff","api_url":"https://ai.aidemos.com/v1/observations/8b496d76-afd9-4658-b0a0-2b73533010ff"},"peers":[{"id":"7d21adab-f4e6-4f0a-9975-37fabe9dc1c2","tool":"gemini","tool_name":"Gemini","verdict":"failed","score":null,"score_total":null,"note":"The tool can miss the requested emotion entirely, producing a neutral and emotionless face instead of the prompted angry and guarded look.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/gemini-input-2-1c4e066a03c3.png","evidence_url":"https://aidemos.com/evidence/7d21adab-f4e6-4f0a-9975-37fabe9dc1c2"},{"id":"9f6dd6f3-855d-4f0a-9535-a5219634595e","tool":"imagineart","tool_name":"ImagineArt","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output preserves the reference expression well.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/imagineart-imagineart-input2-interrogation-ce155eab95e8.jpg","evidence_url":"https://aidemos.com/evidence/9f6dd6f3-855d-4f0a-9535-a5219634595e"},{"id":"ca959663-ea8a-4b8b-9cee-60d17e3305d5","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"The same neutral-expression failure repeats on a second reference, so the tool still misses the prompt's angry, guarded tone.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/leonardo-ai-image-95c2a8e685d7.png","evidence_url":"https://aidemos.com/evidence/ca959663-ea8a-4b8b-9cee-60d17e3305d5"}],"other_criteria":[{"id":"e922f661-f08e-4080-b888-372f905ec1c7","criterion":"accessory-and-detail-retention","criterion_name":"Accessory & detail retention","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Retains the bindi and strong brows in the frontal output, while only slightly darkening skin tone relative to the reference.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/e922f661-f08e-4080-b888-372f905ec1c7"},{"id":"65f7903c-08c2-4f33-987b-43e5e8b93aa7","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Becomes very hard to verify once the face turns too far, and the report says the bindi disappears entirely, leaving the identity lock extremely weak.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/65f7903c-08c2-4f33-987b-43e5e8b93aa7"},{"id":"0d13a121-3d00-4ff1-a1f2-e9f788a60b06","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Preserves face shape and structure well in a frontal composition, though the report notes a slight darkening and a marginally wider, rounder face than the reference.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/0d13a121-3d00-4ff1-a1f2-e9f788a60b06"},{"id":"8e6c6b91-aa81-4f53-bbf7-a4fb3b34dbef","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Cleansly follows the interrogation-room prompt with the navy shirt, hands flat on the table, and bare room composition all in place.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/8e6c6b91-aa81-4f53-bbf7-a4fb3b34dbef"},{"id":"c3fba2ff-eb7e-4ef1-9658-7af7cce5d52c","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Executes the market scene strongly: the crowd feels authentic, the mustard sari and red blouse match the prompt, and the jute bag with vegetables is present.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/c3fba2ff-eb7e-4ef1-9658-7af7cce5d52c"}],"appears_in":[{"page_type":"ranking","slug":"consistent-ai-characters","title":"Best AI Tools for Consistent AI Characters Across Scenes and Poses","url":"https://aidemos.com/best/consistent-ai-characters","binding":"run"}],"same_scenario":[{"id":"7d21adab-f4e6-4f0a-9975-37fabe9dc1c2","tool":"gemini","tool_name":"Gemini","verdict":"failed","score":null,"score_total":null,"note":"The tool can miss the requested emotion entirely, producing a neutral and emotionless face instead of the prompted angry and guarded look."},{"id":"9f6dd6f3-855d-4f0a-9535-a5219634595e","tool":"imagineart","tool_name":"ImagineArt","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output preserves the reference expression well."},{"id":"ca959663-ea8a-4b8b-9cee-60d17e3305d5","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"The same neutral-expression failure repeats on a second reference, so the tool still misses the prompt's angry, guarded tone."}]}