{"observation":{"id":"c42eea97-46ba-43ef-b287-93886b03407c","tool":"hindsight","tool_name":"Hindsight","criterion":"relevant-retrieval","criterion_name":"Relevant Retrieval","criterion_definition":"Checks whether the tool retrieves the right memory for the current task.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"The main value of memory is surfacing the right context when needed, so retrieval quality directly determines usefulness. (3 of 3 judges)","scenario":"personal-work-brain-memory","scenario_name":"Personal Work Brain Memory","group_tag":"memory-for-ai-agents","scenario_description":"A multi-session personal assistant memory test where the user first sets working-style preferences, then asks for an internal update, and finally requests a formal partner email to check whether the assistant applies memory selectively and appropriately across different writing tasks.","modality":"text","input_text":"Session 1:\nUse this under user_id: founder_001\n\nI run a small AI product/research team. When you help me, remember how I work:\n- Keep outputs short, direct, and copy-paste ready.\n- Do not make writing sound too polished or motivational.\n- Always mention what proof or artifact is needed before making a strong claim.\n- If a task is risky or unclear, tell me the safest next step instead of guessing.\n\nSession 2:\nUse this under user_id: founder_001\n\nToday I am testing tools for an AI memory use case. I want to show users that memory is not just \"remember my favorite color.\" It should help an assistant continue real work across days, remember my working style, and avoid repeating the same explanation again.\n\nCreate a short internal update for my team about what I worked on today and what we should test next.\n\nSession 3:\nUse this under user_id: founder_001\n\nNow write a formal email to a potential enterprise partner asking if they are open to a product demo next week. Keep it professional.","input_artifact_refs":[],"stresses":["work-style preference memory","cross-session retrieval","tone adaptation by task","proof-first behavior","avoiding overgeneralization of memory"],"verdict":"worked","score":null,"score_total":null,"note":"On a later task, it retrieved the earlier work-style and memory-testing context for the internal update instead of falling back to unrelated conversation history.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-hindsight-input1-internal-update-memory--1dc2911e7f3a.png","role":"output","alt":null}],"run_id":"6e31afbb-34d7-459a-b688-68ef76fc615a","study_title":"Memory for AI Agents","study_kind":"generation","research_task":"86ba16xrp","tested_at":null,"completeness":"input-and-output","input":{"state":"text","text":"Session 1:\nUse this under user_id: founder_001\n\nI run a small AI product/research team. When you help me, remember how I work:\n- Keep outputs short, direct, and copy-paste ready.\n- Do not make writing sound too polished or motivational.\n- Always mention what proof or artifact is needed before making a strong claim.\n- If a task is risky or unclear, tell me the safest next step instead of guessing.\n\nSession 2:\nUse this under user_id: founder_001\n\nToday I am testing tools for an AI memory use case. I want to show users that memory is not just \"remember my favorite color.\" It should help an assistant continue real work across days, remember my working style, and avoid repeating the same explanation again.\n\nCreate a short internal update for my team about what I worked on today and what we should test next.\n\nSession 3:\nUse this under user_id: founder_001\n\nNow write a formal email to a potential enterprise partner asking if they are open to a product demo next week. Keep it professional.","files":[],"modality":"text","stresses":["work-style preference memory","cross-session retrieval","tone adaptation by task","proof-first behavior","avoiding overgeneralization of memory"]},"tool_page_slug":"hindsight","tool_url":"https://aidemos.com/tools/hindsight","permalink":"https://aidemos.com/evidence/c42eea97-46ba-43ef-b287-93886b03407c","api_url":"https://ai.aidemos.com/v1/observations/c42eea97-46ba-43ef-b287-93886b03407c"},"peers":[{"id":"c216c674-b0bf-4841-a588-fa1f4f037883","tool":"cognee","tool_name":"Cognee","verdict":"worked","score":null,"score_total":null,"note":"Recalls prior working-style memory in later sessions through GRAPH_COMPLETION, exposing evidence chunks plus dataset and document/chunk IDs rather than a simple memory-hit flag.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-cognee-input1-session2-internal-update-g-55d56096a007.png","evidence_url":"https://aidemos.com/evidence/c216c674-b0bf-4841-a588-fa1f4f037883"},{"id":"71a73a3b-261d-43ba-a309-932ceb3789b0","tool":"mem0","tool_name":"Mem0","verdict":"worked","score":null,"score_total":null,"note":"Retrieves the previously stored work-style memory in a later session so the assistant can reuse it for a follow-up internal update.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-mem0-input1-internal-update-memory-retri-2f188d5c8964.png","evidence_url":"https://aidemos.com/evidence/71a73a3b-261d-43ba-a309-932ceb3789b0"},{"id":"95da3a54-d778-448f-8d00-2488e9aab719","tool":"supermemory","tool_name":"Supermemory","verdict":"worked","score":null,"score_total":null,"note":"Later turns could draw on the stored working preferences without the user restating them, showing that prior context remained available across sessions and influenced follow-up responses.","artifact_count":2,"thumbnail":"https://d3epheqghktydj.cloudfront.net/memory-for-ai-agents-supermemory-input1-session2-internal-upd-1de723d7dbfc.png","evidence_url":"https://aidemos.com/evidence/95da3a54-d778-448f-8d00-2488e9aab719"},{"id":"bc958718-f130-443c-9ca9-046dbccae3c4","tool":"zep","tool_name":"Zep","verdict":"worked","score":null,"score_total":null,"note":"Retrieved the saved founder work style in a later task and produced a short internal update without requiring the user to restate the preference block.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/zep-zep-input1-internal-update-memory-applie-107498140287.png","evidence_url":"https://aidemos.com/evidence/bc958718-f130-443c-9ca9-046dbccae3c4"}],"other_criteria":[{"id":"237256dc-81fd-4381-89e5-63302db008cf","criterion":"correct-application","criterion_name":"Correct Application","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"It applied the retrieved work-style memory to the internal update, but the report says the reply was slightly more structured and polished than the user's strict short, direct preference.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/237256dc-81fd-4381-89e5-63302db008cf"},{"id":"8408dacc-1dde-47db-b719-ad9d25f7423e","criterion":"correct-application","criterion_name":"Correct Application","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It keeps the formal partner email professional and does not over-apply the internal-update style to a different writing task.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/8408dacc-1dde-47db-b719-ad9d25f7423e"},{"id":"90662cab-d0f4-453a-8504-80a8f4ffed74","criterion":"memory-capture-quality","criterion_name":"Memory Capture Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It durably stores work-style preferences such as short, direct output, proof-before-claim behavior, and safest-next-step handling, rather than treating them as throwaway chat noise.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/90662cab-d0f4-453a-8504-80a8f4ffed74"}],"appears_in":[{"page_type":"ranking","slug":"ai-agent-memory-tools","title":"Best AI Tools for Memory for AI Agents","url":"https://aidemos.com/best/ai-agent-memory-tools","binding":"run"}],"same_scenario":[{"id":"c216c674-b0bf-4841-a588-fa1f4f037883","tool":"cognee","tool_name":"Cognee","verdict":"worked","score":null,"score_total":null,"note":"Recalls prior working-style memory in later sessions through GRAPH_COMPLETION, exposing evidence chunks plus dataset and document/chunk IDs rather than a simple memory-hit flag."},{"id":"71a73a3b-261d-43ba-a309-932ceb3789b0","tool":"mem0","tool_name":"Mem0","verdict":"worked","score":null,"score_total":null,"note":"Retrieves the previously stored work-style memory in a later session so the assistant can reuse it for a follow-up internal update."},{"id":"95da3a54-d778-448f-8d00-2488e9aab719","tool":"supermemory","tool_name":"Supermemory","verdict":"worked","score":null,"score_total":null,"note":"Later turns could draw on the stored working preferences without the user restating them, showing that prior context remained available across sessions and influenced follow-up responses."},{"id":"bc958718-f130-443c-9ca9-046dbccae3c4","tool":"zep","tool_name":"Zep","verdict":"worked","score":null,"score_total":null,"note":"Retrieved the saved founder work style in a later task and produced a short internal update without requiring the user to restate the preference block."}]}