{"observation":{"id":"9a9dd5b9-dd5d-49ba-9bce-af4ae801dcf2","tool":"zep","tool_name":"Zep","criterion":"memory-capture-quality","criterion_name":"Memory Capture Quality","criterion_definition":"Checks whether the tool stores useful durable context, not random conversation noise.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the tool does not store useful durable context instead of noise, it is not doing the core memory job. (3 of 3 judges)","scenario":"personal-work-brain-memory","scenario_name":"Personal Work Brain Memory","group_tag":"memory-for-ai-agents","scenario_description":"A multi-session personal assistant memory test where the user first sets working-style preferences, then asks for an internal update, and finally requests a formal partner email to check whether the assistant applies memory selectively and appropriately across different writing tasks.","modality":"text","input_text":"Session 1:\nUse this under user_id: founder_001\n\nI run a small AI product/research team. When you help me, remember how I work:\n- Keep outputs short, direct, and copy-paste ready.\n- Do not make writing sound too polished or motivational.\n- Always mention what proof or artifact is needed before making a strong claim.\n- If a task is risky or unclear, tell me the safest next step instead of guessing.\n\nSession 2:\nUse this under user_id: founder_001\n\nToday I am testing tools for an AI memory use case. I want to show users that memory is not just \"remember my favorite color.\" It should help an assistant continue real work across days, remember my working style, and avoid repeating the same explanation again.\n\nCreate a short internal update for my team about what I worked on today and what we should test next.\n\nSession 3:\nUse this under user_id: founder_001\n\nNow write a formal email to a potential enterprise partner asking if they are open to a product demo next week. Keep it professional.","input_artifact_refs":[],"stresses":["work-style preference memory","cross-session retrieval","tone adaptation by task","proof-first behavior","avoiding overgeneralization of memory"],"verdict":"worked","score":null,"score_total":null,"note":"Captured four durable work-style preferences for the user: keep outputs short and direct, make them copy-paste ready, avoid over-polished or motivational writing, require proof/artifacts for strong claims, and choose the safest next step when something is unclear.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/zep-zep-input1-founder-work-preference-creat-83bebe96b8a3.png","role":"output","alt":null}],"run_id":"6e31afbb-34d7-459a-b688-68ef76fc615a","study_title":"Memory for AI Agents","study_kind":"generation","research_task":"86ba16xrp","tested_at":null,"completeness":"input-and-output","input":{"state":"text","text":"Session 1:\nUse this under user_id: founder_001\n\nI run a small AI product/research team. When you help me, remember how I work:\n- Keep outputs short, direct, and copy-paste ready.\n- Do not make writing sound too polished or motivational.\n- Always mention what proof or artifact is needed before making a strong claim.\n- If a task is risky or unclear, tell me the safest next step instead of guessing.\n\nSession 2:\nUse this under user_id: founder_001\n\nToday I am testing tools for an AI memory use case. I want to show users that memory is not just \"remember my favorite color.\" It should help an assistant continue real work across days, remember my working style, and avoid repeating the same explanation again.\n\nCreate a short internal update for my team about what I worked on today and what we should test next.\n\nSession 3:\nUse this under user_id: founder_001\n\nNow write a formal email to a potential enterprise partner asking if they are open to a product demo next week. Keep it professional.","files":[],"modality":"text","stresses":["work-style preference memory","cross-session retrieval","tone adaptation by task","proof-first behavior","avoiding overgeneralization of memory"]},"tool_page_slug":"zep","tool_url":"https://aidemos.com/tools/zep","permalink":"https://aidemos.com/evidence/9a9dd5b9-dd5d-49ba-9bce-af4ae801dcf2","api_url":"https://ai.aidemos.com/v1/observations/9a9dd5b9-dd5d-49ba-9bce-af4ae801dcf2"},"peers":[{"id":"90d46d99-526b-4942-8825-01f7e7257c0f","tool":"cognee","tool_name":"Cognee","verdict":"worked","score":null,"score_total":null,"note":"Stores a compact working-style profile as durable graph-backed memory: concise/direct output, no over-polished tone, proof-before-claims, and safest-next-step behavior for risky tasks.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-cognee-input1-session1-working-style-mem-7bfd6578718b.png","evidence_url":"https://aidemos.com/evidence/90d46d99-526b-4942-8825-01f7e7257c0f"},{"id":"90662cab-d0f4-453a-8504-80a8f4ffed74","tool":"hindsight","tool_name":"Hindsight","verdict":"worked","score":null,"score_total":null,"note":"It durably stores work-style preferences such as short, direct output, proof-before-claim behavior, and safest-next-step handling, rather than treating them as throwaway chat noise.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-hindsight-input1-founder-work-preference-23e99e40f8f0.png","evidence_url":"https://aidemos.com/evidence/90662cab-d0f4-453a-8504-80a8f4ffed74"},{"id":"91b32059-a56e-4bad-b69f-9e48850f45da","tool":"mem0","tool_name":"Mem0","verdict":"worked","score":null,"score_total":null,"note":"Captures a durable work-style preference as compact memory cards rather than raw chat history, including short, direct output expectations and proof-first guidance.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-mem0-input1-founder-work-preference-crea-35e7b82616e2.png","evidence_url":"https://aidemos.com/evidence/91b32059-a56e-4bad-b69f-9e48850f45da"},{"id":"0db4e3ed-3f70-4b20-8747-078fe9c03570","tool":"supermemory","tool_name":"Supermemory","verdict":"worked","score":null,"score_total":null,"note":"The tool captured a reusable working-style profile, not just a one-off fact: it stored the user's short, direct output preference, the anti-hype writing style, the requirement to mention proof or artifacts before strong claims, the safe-next-step rule for risky tasks, and the small AI product/research team context.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-supermemory-input1-session1-working-styl-0f1fbd209476.png","evidence_url":"https://aidemos.com/evidence/0db4e3ed-3f70-4b20-8747-078fe9c03570"}],"other_criteria":[{"id":"4ea6ccca-2fff-49db-bac2-81b4260855bd","criterion":"correct-application","criterion_name":"Correct Application","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Applied the remembered style selectively across two writing tasks: the internal update stayed short and practical, while the later partner email stayed professional instead of turning into an internal-note style response.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/4ea6ccca-2fff-49db-bac2-81b4260855bd"},{"id":"bc958718-f130-443c-9ca9-046dbccae3c4","criterion":"relevant-retrieval","criterion_name":"Relevant Retrieval","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Retrieved the saved founder work style in a later task and produced a short internal update without requiring the user to restate the preference block.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/bc958718-f130-443c-9ca9-046dbccae3c4"}],"appears_in":[{"page_type":"ranking","slug":"ai-agent-memory-tools","title":"Best AI Tools for Memory for AI Agents","url":"https://aidemos.com/best/ai-agent-memory-tools","binding":"run"}],"same_scenario":[{"id":"90d46d99-526b-4942-8825-01f7e7257c0f","tool":"cognee","tool_name":"Cognee","verdict":"worked","score":null,"score_total":null,"note":"Stores a compact working-style profile as durable graph-backed memory: concise/direct output, no over-polished tone, proof-before-claims, and safest-next-step behavior for risky tasks."},{"id":"90662cab-d0f4-453a-8504-80a8f4ffed74","tool":"hindsight","tool_name":"Hindsight","verdict":"worked","score":null,"score_total":null,"note":"It durably stores work-style preferences such as short, direct output, proof-before-claim behavior, and safest-next-step handling, rather than treating them as throwaway chat noise."},{"id":"91b32059-a56e-4bad-b69f-9e48850f45da","tool":"mem0","tool_name":"Mem0","verdict":"worked","score":null,"score_total":null,"note":"Captures a durable work-style preference as compact memory cards rather than raw chat history, including short, direct output expectations and proof-first guidance."},{"id":"0db4e3ed-3f70-4b20-8747-078fe9c03570","tool":"supermemory","tool_name":"Supermemory","verdict":"worked","score":null,"score_total":null,"note":"The tool captured a reusable working-style profile, not just a one-off fact: it stored the user's short, direct output preference, the anti-hype writing style, the requirement to mention proof or artifacts before strong claims, the safe-next-step rule for risky tasks, and the small AI product/research team context."}]}