{"observation":{"id":"90dbd4a6-001a-4c33-a064-e8d46d67b736","tool":"cognee","tool_name":"Cognee","criterion":"correct-application","criterion_name":"Correct Application","criterion_definition":"Checks whether the agent actually uses the retrieved memory correctly.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"For agent memory, it is not enough to retrieve facts; the system must help the agent use them correctly in task execution. (2 of 3 judges)","scenario":"personal-work-brain-memory","scenario_name":"Personal Work Brain Memory","group_tag":"memory-for-ai-agents","scenario_description":"A multi-session personal assistant memory test where the user first sets working-style preferences, then asks for an internal update, and finally requests a formal partner email to check whether the assistant applies memory selectively and appropriately across different writing tasks.","modality":"text","input_text":"Session 1:\nUse this under user_id: founder_001\n\nI run a small AI product/research team. When you help me, remember how I work:\n- Keep outputs short, direct, and copy-paste ready.\n- Do not make writing sound too polished or motivational.\n- Always mention what proof or artifact is needed before making a strong claim.\n- If a task is risky or unclear, tell me the safest next step instead of guessing.\n\nSession 2:\nUse this under user_id: founder_001\n\nToday I am testing tools for an AI memory use case. I want to show users that memory is not just \"remember my favorite color.\" It should help an assistant continue real work across days, remember my working style, and avoid repeating the same explanation again.\n\nCreate a short internal update for my team about what I worked on today and what we should test next.\n\nSession 3:\nUse this under user_id: founder_001\n\nNow write a formal email to a potential enterprise partner asking if they are open to a product demo next week. Keep it professional.","input_artifact_refs":[],"stresses":["work-style preference memory","cross-session retrieval","tone adaptation by task","proof-first behavior","avoiding overgeneralization of memory"],"verdict":"worked","score":null,"score_total":null,"note":"Applies the stored working style to a new internal update by keeping the reply short and proof-oriented while blending in the current memory-testing project context.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-cognee-input1-session2-internal-update-g-55d56096a007.png","role":"output","alt":null}],"run_id":"6e31afbb-34d7-459a-b688-68ef76fc615a","study_title":"Memory for AI Agents","study_kind":"generation","research_task":"86ba16xrp","tested_at":null,"completeness":"input-and-output","input":{"state":"text","text":"Session 1:\nUse this under user_id: founder_001\n\nI run a small AI product/research team. When you help me, remember how I work:\n- Keep outputs short, direct, and copy-paste ready.\n- Do not make writing sound too polished or motivational.\n- Always mention what proof or artifact is needed before making a strong claim.\n- If a task is risky or unclear, tell me the safest next step instead of guessing.\n\nSession 2:\nUse this under user_id: founder_001\n\nToday I am testing tools for an AI memory use case. I want to show users that memory is not just \"remember my favorite color.\" It should help an assistant continue real work across days, remember my working style, and avoid repeating the same explanation again.\n\nCreate a short internal update for my team about what I worked on today and what we should test next.\n\nSession 3:\nUse this under user_id: founder_001\n\nNow write a formal email to a potential enterprise partner asking if they are open to a product demo next week. Keep it professional.","files":[],"modality":"text","stresses":["work-style preference memory","cross-session retrieval","tone adaptation by task","proof-first behavior","avoiding overgeneralization of memory"]},"tool_page_slug":"cognee","tool_url":"https://aidemos.com/tools/cognee","permalink":"https://aidemos.com/evidence/90dbd4a6-001a-4c33-a064-e8d46d67b736","api_url":"https://ai.aidemos.com/v1/observations/90dbd4a6-001a-4c33-a064-e8d46d67b736"},"peers":[{"id":"8408dacc-1dde-47db-b719-ad9d25f7423e","tool":"hindsight","tool_name":"Hindsight","verdict":"worked","score":null,"score_total":null,"note":"It keeps the formal partner email professional and does not over-apply the internal-update style to a different writing task.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-hindsight-input1-formal-email-boundary-c-239af8501813.png","evidence_url":"https://aidemos.com/evidence/8408dacc-1dde-47db-b719-ad9d25f7423e"},{"id":"d06d2bc8-e02f-461d-8b1c-77607ff736e0","tool":"mem0","tool_name":"Mem0","verdict":"mixed","score":null,"score_total":null,"note":"Uses the retrieved style memory only partially: the internal update still reads a bit generic and polished instead of fully short, direct, and copy-paste ready.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-mem0-input1-internal-update-memory-retri-2f188d5c8964.png","evidence_url":"https://aidemos.com/evidence/d06d2bc8-e02f-461d-8b1c-77607ff736e0"},{"id":"eea1f5a6-af74-4381-ac83-b115365c18d3","tool":"supermemory","tool_name":"Supermemory","verdict":"mixed","score":null,"score_total":null,"note":"The tool applied the memory selectively: the internal update was useful but a bit more structured and polished than the requested short, direct style, while the formal partner email stayed professional and did not leak the internal terse style into a different writing task.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-supermemory-input1-session2-internal-upd-1de723d7dbfc.png","evidence_url":"https://aidemos.com/evidence/eea1f5a6-af74-4381-ac83-b115365c18d3"},{"id":"4ea6ccca-2fff-49db-bac2-81b4260855bd","tool":"zep","tool_name":"Zep","verdict":"worked","score":null,"score_total":null,"note":"Applied the remembered style selectively across two writing tasks: the internal update stayed short and practical, while the later partner email stayed professional instead of turning into an internal-note style response.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/zep-zep-input1-internal-update-memory-applie-107498140287.png","evidence_url":"https://aidemos.com/evidence/4ea6ccca-2fff-49db-bac2-81b4260855bd"}],"other_criteria":[{"id":"90d46d99-526b-4942-8825-01f7e7257c0f","criterion":"memory-capture-quality","criterion_name":"Memory Capture Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Stores a compact working-style profile as durable graph-backed memory: concise/direct output, no over-polished tone, proof-before-claims, and safest-next-step behavior for risky tasks.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/90d46d99-526b-4942-8825-01f7e7257c0f"},{"id":"22461c92-39c7-4585-a0b6-7210ea110012","criterion":"observability-and-debugging","criterion_name":"Observability and Debugging","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Exposes recall provenance directly in the UI: the Last Recall Response panel shows the source graph completion, evidence chunks, dataset ID, and document/chunk IDs.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/22461c92-39c7-4585-a0b6-7210ea110012"},{"id":"c216c674-b0bf-4841-a588-fa1f4f037883","criterion":"relevant-retrieval","criterion_name":"Relevant Retrieval","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Recalls prior working-style memory in later sessions through GRAPH_COMPLETION, exposing evidence chunks plus dataset and document/chunk IDs rather than a simple memory-hit flag.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/c216c674-b0bf-4841-a588-fa1f4f037883"}],"appears_in":[{"page_type":"ranking","slug":"ai-agent-memory-tools","title":"Best AI Tools for Memory for AI Agents","url":"https://aidemos.com/best/ai-agent-memory-tools","binding":"run"}],"same_scenario":[{"id":"8408dacc-1dde-47db-b719-ad9d25f7423e","tool":"hindsight","tool_name":"Hindsight","verdict":"worked","score":null,"score_total":null,"note":"It keeps the formal partner email professional and does not over-apply the internal-update style to a different writing task."},{"id":"d06d2bc8-e02f-461d-8b1c-77607ff736e0","tool":"mem0","tool_name":"Mem0","verdict":"mixed","score":null,"score_total":null,"note":"Uses the retrieved style memory only partially: the internal update still reads a bit generic and polished instead of fully short, direct, and copy-paste ready."},{"id":"eea1f5a6-af74-4381-ac83-b115365c18d3","tool":"supermemory","tool_name":"Supermemory","verdict":"mixed","score":null,"score_total":null,"note":"The tool applied the memory selectively: the internal update was useful but a bit more structured and polished than the requested short, direct style, while the formal partner email stayed professional and did not leak the internal terse style into a different writing task."},{"id":"4ea6ccca-2fff-49db-bac2-81b4260855bd","tool":"zep","tool_name":"Zep","verdict":"worked","score":null,"score_total":null,"note":"Applied the remembered style selectively across two writing tasks: the internal update stayed short and practical, while the later partner email stayed professional instead of turning into an internal-note style response."}]}