{"observation":{"id":"a0d5ce57-a837-4383-95b6-2d7e9cfdd208","tool":"hindsight","tool_name":"Hindsight","criterion":"memory-capture-quality","criterion_name":"Memory Capture Quality","criterion_definition":"Checks whether the tool stores useful durable context, not random conversation noise.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the tool does not store useful durable context instead of noise, it is not doing the core memory job. (3 of 3 judges)","scenario":"team-handoff-project-continuity-memory","scenario_name":"Team Handoff / Project Continuity Memory","group_tag":"memory-for-ai-agents","scenario_description":"A project continuity and handoff test where the assistant must remember project goals, project rules, and a changed testing direction, then produce a useful handoff note for an unavailable team member without leaking context into an unrelated project.","modality":"text","input_text":"Session 1:\nUse this under project_id: ai_demos_memory_use_case\n\nWe are working on an AI Demos use case called Memory for AI Agents. The goal is to help users understand which memory tools are actually useful for real agent workflows. We are not promoting any tool. We are testing whether memory can help with real continuity: personal work brain, client relationship memory, and team handoff.\n\nSession 2:\nUse this under project_id: ai_demos_memory_use_case\n\nImportant project rules:\n- No observation without proof.\n- Screenshots and artifacts are primary evidence.\n- Inputs must help rank tools, not just prove that tools can store one fact.\n- Memory should be checked for retrieval, update handling, scope control, deletion or retirement, and observability.\n- The page should stay practical and user-facing, not only technical.\n\nSession 3:\nUse this under project_id: ai_demos_memory_use_case\n\nProject direction changed slightly. The old input set was too QA-style and not relatable enough. The new direction is to use real workflows: personal work brain memory, client relationship memory, and team handoff/project continuity memory.\n\nSession 4:\nUse this under project_id: ai_demos_memory_use_case\n\nI am unavailable tomorrow. Create a handoff note for an intern who needs to continue this use case. The note should explain:\n1. What this use case is about.\n2. What the current testing direction is.\n3. What rules they must follow before writing observations.\n4. What artifacts they need to capture while testing.\n\nSession 5:\nUse this under project_id: unrelated_sales_agent_project\n\nWe are building a sales email agent for a different project. Create a short kickoff note for the team.","input_artifact_refs":[],"stresses":["project-level memory","decision and rule retention","changed direction handling","handoff continuity","scope separation across projects"],"verdict":"worked","score":null,"score_total":null,"note":"It stores the project goal, proof-first rules, evaluation dimensions, and the shift toward real workflows instead of QA-style inputs.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-hindsight-input3-project-context-created-d08633418fc6.png","role":"output","alt":null}],"run_id":"6e31afbb-34d7-459a-b688-68ef76fc615a","study_title":"Memory for AI Agents","study_kind":"generation","research_task":"86ba16xrp","tested_at":null,"completeness":"input-and-output","input":{"state":"text","text":"Session 1:\nUse this under project_id: ai_demos_memory_use_case\n\nWe are working on an AI Demos use case called Memory for AI Agents. The goal is to help users understand which memory tools are actually useful for real agent workflows. We are not promoting any tool. We are testing whether memory can help with real continuity: personal work brain, client relationship memory, and team handoff.\n\nSession 2:\nUse this under project_id: ai_demos_memory_use_case\n\nImportant project rules:\n- No observation without proof.\n- Screenshots and artifacts are primary evidence.\n- Inputs must help rank tools, not just prove that tools can store one fact.\n- Memory should be checked for retrieval, update handling, scope control, deletion or retirement, and observability.\n- The page should stay practical and user-facing, not only technical.\n\nSession 3:\nUse this under project_id: ai_demos_memory_use_case\n\nProject direction changed slightly. The old input set was too QA-style and not relatable enough. The new direction is to use real workflows: personal work brain memory, client relationship memory, and team handoff/project continuity memory.\n\nSession 4:\nUse this under project_id: ai_demos_memory_use_case\n\nI am unavailable tomorrow. Create a handoff note for an intern who needs to continue this use case. The note should explain:\n1. What this use case is about.\n2. What the current testing direction is.\n3. What rules they must follow before writing observations.\n4. What artifacts they need to capture while testing.\n\nSession 5:\nUse this under project_id: unrelated_sales_agent_project\n\nWe are building a sales email agent for a different project. Create a short kickoff note for the team.","files":[],"modality":"text","stresses":["project-level memory","decision and rule retention","changed direction handling","handoff continuity","scope separation across projects"]},"tool_page_slug":"hindsight","tool_url":"https://aidemos.com/tools/hindsight","permalink":"https://aidemos.com/evidence/a0d5ce57-a837-4383-95b6-2d7e9cfdd208","api_url":"https://ai.aidemos.com/v1/observations/a0d5ce57-a837-4383-95b6-2d7e9cfdd208"},"peers":[{"id":"03307abc-be1a-4a93-b0ca-88806eaf21c5","tool":"cognee","tool_name":"Cognee","verdict":"mixed","score":null,"score_total":null,"note":"Stores project context, rules, and direction changes under the project container, but the visible replies were only generic acknowledgments rather than a visible restatement of the stored content.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-cognee-input3-session1-project-context-g-d0e636fb3164.png","evidence_url":"https://aidemos.com/evidence/03307abc-be1a-4a93-b0ca-88806eaf21c5"},{"id":"93919730-5b52-4713-966a-cafd867a5e3c","tool":"mem0","tool_name":"Mem0","verdict":"worked","score":null,"score_total":null,"note":"Stores project continuity context as memory, including the use case goal, proof-first rules, and the direction change toward real workflows.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-mem0-input3-project-context-created-525d36a7efb2.png","evidence_url":"https://aidemos.com/evidence/93919730-5b52-4713-966a-cafd867a5e3c"},{"id":"4716a0a7-a960-4908-9ee4-7ee174e009c7","tool":"supermemory","tool_name":"Supermemory","verdict":"mixed","score":null,"score_total":null,"note":"The tool stored the raw project-session content, but the report flags a granularity problem: it often kept full conversational turns rather than only distilled reusable project facts, which increases retrieval noise as the memory set grows.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-supermemory-input3-session1-project-cont-39812fd361e8.png","evidence_url":"https://aidemos.com/evidence/4716a0a7-a960-4908-9ee4-7ee174e009c7"},{"id":"724dae08-39dc-4526-8d82-66495d788467","tool":"zep","tool_name":"Zep","verdict":"worked","score":null,"score_total":null,"note":"Captured project continuity context including the Memory for AI Agents goal, the three real-work scenarios, and proof-first evaluation rules covering screenshots, ranking value, retrieval, update handling, scope control, deletion or retirement, and observability.","artifact_count":3,"thumbnail":"https://d3epheqghktydj.cloudfront.net/zep-zep-input3-project-direction-updated-e2d8864deb8e.png","evidence_url":"https://aidemos.com/evidence/724dae08-39dc-4526-8d82-66495d788467"}],"other_criteria":[{"id":"83a06854-4c5c-4fcd-96f3-17d4ddff262e","criterion":"correct-application","criterion_name":"Correct Application","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It turns the project memory into a usable handoff note that explains the use case, the current direction, the required rules, and the artifact-capture needs.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/83a06854-4c5c-4fcd-96f3-17d4ddff262e"},{"id":"2cb22d1a-734c-4a92-b9b9-783b1320d1f7","criterion":"scope-control","criterion_name":"Scope Control","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It keeps the unrelated sales-agent project separate, so the AI Demos memory rules do not bleed into a different project bank.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/2cb22d1a-734c-4a92-b9b9-783b1320d1f7"}],"appears_in":[{"page_type":"ranking","slug":"ai-agent-memory-tools","title":"Best AI Tools for Memory for AI Agents","url":"https://aidemos.com/best/ai-agent-memory-tools","binding":"run"}],"same_scenario":[{"id":"03307abc-be1a-4a93-b0ca-88806eaf21c5","tool":"cognee","tool_name":"Cognee","verdict":"mixed","score":null,"score_total":null,"note":"Stores project context, rules, and direction changes under the project container, but the visible replies were only generic acknowledgments rather than a visible restatement of the stored content."},{"id":"93919730-5b52-4713-966a-cafd867a5e3c","tool":"mem0","tool_name":"Mem0","verdict":"worked","score":null,"score_total":null,"note":"Stores project continuity context as memory, including the use case goal, proof-first rules, and the direction change toward real workflows."},{"id":"4716a0a7-a960-4908-9ee4-7ee174e009c7","tool":"supermemory","tool_name":"Supermemory","verdict":"mixed","score":null,"score_total":null,"note":"The tool stored the raw project-session content, but the report flags a granularity problem: it often kept full conversational turns rather than only distilled reusable project facts, which increases retrieval noise as the memory set grows."},{"id":"724dae08-39dc-4526-8d82-66495d788467","tool":"zep","tool_name":"Zep","verdict":"worked","score":null,"score_total":null,"note":"Captured project continuity context including the Memory for AI Agents goal, the three real-work scenarios, and proof-first evaluation rules covering screenshots, ranking value, retrieval, update handling, scope control, deletion or retirement, and observability."}]}