{"observation":{"id":"d08a8783-fc20-4cf5-8b89-47e25f9f73b7","tool":"gemini","tool_name":"Gemini","criterion":"expression-accuracy","criterion_name":"Expression accuracy","criterion_definition":"Whether the emotional tone and facial expression match what was explicitly prompted.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the character’s intended emotion or facial expression does not match the prompt, the generated character is not being controlled reliably across scenes. (3 of 3 judges)","scenario":"full-frontal-portrait","scenario_name":"Full frontal portrait","group_tag":null,"scenario_description":"Full frontal portrait reference image with fair skin, curly dark hair, bindi, gold jhumka earrings, and a green stone necklace. All features are clearly visible in good natural lighting, making it the easiest identity anchor for the tools.","modality":"image","input_text":null,"input_artifact_refs":[],"stresses":["Baseline identity preservation","Accessory retention","Best-case frontal face matching","Consistent character reuse across varied scenes"],"verdict":"worked","score":null,"score_total":null,"note":"The tool can capture a prompted angry and guarded mood, with direct eye contact and an intense expression in the interrogation shot.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/gemini-input-1-abf743cbaaf0.png","role":"input","alt":"Chatgpt input 1.png"},{"url":"https://d3epheqghktydj.cloudfront.net/gemini-gemini-input1-interrogation-c316928f5e50.png","role":"output","alt":null}],"run_id":"1dfb8fa4-f7f0-47dc-911b-7db3af467e9a","study_title":"Generate Consistent AI Characters Across Different Scenes and Poses","study_kind":"generation","research_task":"86b96df11","tested_at":null,"completeness":"input-and-output","input":{"state":"not-captured","text":null,"files":[],"modality":"image","stresses":["Baseline identity preservation","Accessory retention","Best-case frontal face matching","Consistent character reuse across varied scenes"]},"tool_page_slug":"gemini","tool_url":"https://aidemos.com/tools/gemini","permalink":"https://aidemos.com/evidence/d08a8783-fc20-4cf5-8b89-47e25f9f73b7","api_url":"https://ai.aidemos.com/v1/observations/d08a8783-fc20-4cf5-8b89-47e25f9f73b7"},"peers":[{"id":"6636ced6-0dd3-4d55-aed4-228102bc03f6","tool":"chatgpt","tool_name":"ChatGPT","verdict":"failed","score":null,"score_total":null,"note":"Misses the prompted brave/determined emotional tone and instead outputs a soft neutral expression, leaving the scene with essentially no emotional alignment.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/chatgpt-chatgpt-input1-horseride-dbbcd9f8517e.png","evidence_url":"https://aidemos.com/evidence/6636ced6-0dd3-4d55-aed4-228102bc03f6"},{"id":"a463c854-4e75-436e-bef8-df02f6bcfcff","tool":"imagineart","tool_name":"ImagineArt","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output matches the requested serious, guarded expression.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/imagineart-imagineart-input1-interrogation-cfe1542ae993.jpg","evidence_url":"https://aidemos.com/evidence/a463c854-4e75-436e-bef8-df02f6bcfcff"},{"id":"b57aa471-7b05-4433-a4f2-f1a621c5aaf3","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"It does not translate an explicitly angry or guarded prompt into facial expression, defaulting to calm neutrality.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/leonardo-ai-leonardo-input1-interrogation-77ba6593d6a0.jpg","evidence_url":"https://aidemos.com/evidence/b57aa471-7b05-4433-a4f2-f1a621c5aaf3"},{"id":"6a1796e6-76aa-4a88-84d2-ba870e97737e","tool":"scenario","tool_name":"Scenario","verdict":"failed","score":null,"score_total":null,"note":"It fails to map an angry or guarded prompt onto the face; the output stays neutral or calm and even reads with a slight smile.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/scenario-scenario-input1-interrogation-8a626f0b95c8.png","evidence_url":"https://aidemos.com/evidence/6a1796e6-76aa-4a88-84d2-ba870e97737e"}],"other_criteria":[{"id":"e124c476-ed64-4388-883d-129c53045500","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Against the full-frontal reference, the warm cafe output can drift into a different character: it was judged very weak and the report says multiple facial features changed.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/e124c476-ed64-4388-883d-129c53045500"},{"id":"af4776d3-f013-4967-a4a8-bd0db677c8e4","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"In the horse-riding scene, the tool can erase the reference face entirely: the report calls the identity match very weak and says the output is a completely different character with changed face shape, eyes, and eyebrows.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/af4776d3-f013-4967-a4a8-bd0db677c8e4"},{"id":"dca3c251-2404-443c-b298-4ef168e9c42c","criterion":"identity-preservation","criterion_name":"Identity preservation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The plain interrogation setup is the best identity anchor for this input, keeping the eyes, face shape, nose, and overall facial structure close to the reference.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/dca3c251-2404-443c-b298-4ef168e9c42c"},{"id":"f9bd9056-0c1f-477c-8cb4-f0a2e6428dbc","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The horse-riding prompt is followed strongly, including the dusty desert setting, sunset lighting, horse motion, riding costume, gloves, boots, scarf, and believable action pose.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/f9bd9056-0c1f-477c-8cb4-f0a2e6428dbc"},{"id":"d332e2b5-664e-482e-88e9-cfe3eeadd988","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room prompt is rendered cleanly, with a plain room, metal table, overhead lighting, formal clothing, and uncluttered environment all matching the scene.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/d332e2b5-664e-482e-88e9-cfe3eeadd988"},{"id":"4478b970-4ebb-48c8-9921-315ce84fd9fe","criterion":"scene-compliance","criterion_name":"Scene compliance","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The warm cafe prompt is followed strongly, with the cozy cafe setting, warm lighting, background blur, sweater, braided hairstyle, and natural pose all rendered correctly.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/4478b970-4ebb-48c8-9921-315ce84fd9fe"}],"appears_in":[{"page_type":"ranking","slug":"consistent-ai-characters","title":"Best AI Tools for Consistent AI Characters Across Scenes and Poses","url":"https://aidemos.com/best/consistent-ai-characters","binding":"run"}],"same_scenario":[{"id":"6636ced6-0dd3-4d55-aed4-228102bc03f6","tool":"chatgpt","tool_name":"ChatGPT","verdict":"failed","score":null,"score_total":null,"note":"Misses the prompted brave/determined emotional tone and instead outputs a soft neutral expression, leaving the scene with essentially no emotional alignment."},{"id":"a463c854-4e75-436e-bef8-df02f6bcfcff","tool":"imagineart","tool_name":"ImagineArt","verdict":"worked","score":null,"score_total":null,"note":"The interrogation-room output matches the requested serious, guarded expression."},{"id":"b57aa471-7b05-4433-a4f2-f1a621c5aaf3","tool":"leonardo-ai","tool_name":"Leonardo AI","verdict":"failed","score":null,"score_total":null,"note":"It does not translate an explicitly angry or guarded prompt into facial expression, defaulting to calm neutrality."},{"id":"6a1796e6-76aa-4a88-84d2-ba870e97737e","tool":"scenario","tool_name":"Scenario","verdict":"failed","score":null,"score_total":null,"note":"It fails to map an angry or guarded prompt onto the face; the output stays neutral or calm and even reads with a slight smile."}]}