{"observation":{"id":"87d23b55-bc59-4d3e-8547-9b80fc1107f6","tool":"granola","tool_name":"Granola","criterion":"transcription-accuracy","criterion_name":"Transcription Accuracy","criterion_definition":"Word accuracy on the shared call, especially names, tools, numbers, and jargon.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the transcript gets names, numbers, and jargon wrong, the note-taker has failed at the core job of capturing the call accurately. (3 of 3 judges)","scenario":"ai-demos-daily-standup-31-july-2026","scenario_name":"AI Demos Daily Standup — 31 July 2026","group_tag":"ai-meeting-notetaker","scenario_description":"A real 25-minute technical engineering daily standup with 14 attendees and about 10 active speakers, used as the single parallel-capture meeting for evaluating AI meeting notetakers on transcription, diarization, summaries, action items, search/chat, and collaboration features.","modality":"image","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/547dd6f13e4a420fa8ad7bf2c88c7612.png?v=1","role":"input","filename":"31-july-meeting-screenshot.png"}],"stresses":["Transcription accuracy for real names, tool names, numbers, and technical jargon","Speaker diarization across multiple active speakers","Robustness to overlapping speech, crosstalk, and rapid turn-taking","Join reliability for bot-based and botless capture","Summary quality on identical source material","Action-item extraction with correct owners and commitments","Topic segmentation of standup updates","Search and chat grounded in the meeting content","Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage"],"verdict":"failed","score":null,"score_total":null,"note":"On this 25-minute, multi-speaker standup, Granola’s transcript quality is unreliable: the published excerpt shows garbled phrasing and mistranscribed wording, and the report says the mishearing pattern recurs across early, middle, and late sections rather than being isolated to one moment.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/d3ec039f484d40328057665e51ba34de.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/022dc79e225247179c80a0c7e8edae2e.png?v=1","role":"output","alt":null}],"run_id":"ace58582-3d1e-48ee-996c-9b3cd03f27a2","study_title":"AI Meeting Notetakers — Capture Accurate Transcripts, Summaries & Action Items From Live Calls","study_kind":"generation","research_task":"86baxegnv","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/547dd6f13e4a420fa8ad7bf2c88c7612.png?v=1","filename":"31-july-meeting-screenshot.png","alt":"AI Demos Daily Standup — 31 July 2026","role":"input"}],"modality":"image","stresses":["Transcription accuracy for real names, tool names, numbers, and technical jargon","Speaker diarization across multiple active speakers","Robustness to overlapping speech, crosstalk, and rapid turn-taking","Join reliability for bot-based and botless capture","Summary quality on identical source material","Action-item extraction with correct owners and commitments","Topic segmentation of standup updates","Search and chat grounded in the meeting content","Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage"]},"tool_page_slug":"granola","tool_url":"https://aidemos.com/tools/granola","permalink":"https://aidemos.com/evidence/87d23b55-bc59-4d3e-8547-9b80fc1107f6","api_url":"https://ai.aidemos.com/v1/observations/87d23b55-bc59-4d3e-8547-9b80fc1107f6"},"peers":[{"id":"1c4f4fe3-3ed8-490d-992b-a25ddd604669","tool":"fathom","tool_name":"Fathom","verdict":"mixed","score":null,"score_total":null,"note":"Fathom's transcript mostly preserves the meeting's names and technical content, but the report records one confirmed name-level error: \"Mahreen\" was rendered as \"Meryl.\"","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/19b062d60f8e45a6a5fa364948287362.png?v=1","evidence_url":"https://aidemos.com/evidence/1c4f4fe3-3ed8-490d-992b-a25ddd604669"},{"id":"d4b52ad2-b536-4580-99fa-644f4b772a93","tool":"fellow","tool_name":"Fellow","verdict":"worked","score":null,"score_total":null,"note":"The transcript was near-clean: the tool captured nearly all names, technical jargon, and numbers correctly, with no significant misheard terms or hallucinations observed in the tested meeting.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/7476ae6a377e411094940f94f34b21f9.png?v=1","evidence_url":"https://aidemos.com/evidence/d4b52ad2-b536-4580-99fa-644f4b772a93"},{"id":"bac222c7-f7cf-4041-8449-df5aa22c88ac","tool":"fireflies-ai","tool_name":"Fireflies.ai","verdict":"worked","score":null,"score_total":null,"note":"Generated a timestamped transcript view, and the report says the full transcript was very accurate: nearly all names, tools, jargon, and numbers were captured correctly with no significant misheard terms or hallucinations.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/7aecac3bebb2479aa0f18cf6c9f1f0db.png?v=1","evidence_url":"https://aidemos.com/evidence/bac222c7-f7cf-4041-8449-df5aa22c88ac"},{"id":"624ffd2e-c9da-4fb4-acc3-041a5b115a50","tool":"happyscribe","tool_name":"HappyScribe","verdict":"mixed","score":null,"score_total":null,"note":"On this ~25-minute multi-speaker standup, HappyScribe captured the vast majority of names, tools, and jargon correctly, and the report records only 1–2 misheard words.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/768c8f926786411984134803cb5396a5.png?v=1","evidence_url":"https://aidemos.com/evidence/624ffd2e-c9da-4fb4-acc3-041a5b115a50"},{"id":"5a191ed7-02a9-4979-931d-9219e0b75fce","tool":"meetgeek","tool_name":"MeetGeek","verdict":"worked","score":null,"score_total":null,"note":"It transcribes a normal ~25-minute, ~10-active-speaker engineering standup mostly accurately, with only minor proper-noun/term drift noted in the report; one example given is \"Madin\" being misheard for \"Mahreen\".","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/f3fbc51e628c42c19595b82c5e481a25.png?v=1","evidence_url":"https://aidemos.com/evidence/5a191ed7-02a9-4979-931d-9219e0b75fce"},{"id":"893f5543-bd03-4e2a-ba30-f6426940628b","tool":"notta","tool_name":"Notta","verdict":"worked","score":null,"score_total":null,"note":"Notta’s transcript capture was accurate on the evaluated standup: the report says it correctly captured names, tool names, numbers, and engineering jargon with no significant word-level errors, silent hallucinations, or misheard terms.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/86173a6a33c7494b96de40392d0740b6.png?v=1","evidence_url":"https://aidemos.com/evidence/893f5543-bd03-4e2a-ba30-f6426940628b"},{"id":"0efa5738-c840-44ba-a72a-b35d289e70fb","tool":"otter-ai","tool_name":"Otter.ai","verdict":"worked","score":null,"score_total":null,"note":"Otter generated a full transcript for the standup and, per the report, captured names, tool names, jargon, and numbers correctly with minimal errors, making the transcript reliable for reference.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/0db942cf24dd4c3f9383ac4f5208e840.png?v=1","evidence_url":"https://aidemos.com/evidence/0efa5738-c840-44ba-a72a-b35d289e70fb"}],"other_criteria":[{"id":"9af81f24-1ac8-4a6a-b543-1959eafc22b0","criterion":"action-item-extraction","criterion_name":"Action-Item Extraction","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Granola extracts concrete next steps with ownership: the visible action item says to create a subtask and add details, names Mahreen Fathima as owner, and marks the item for same-day follow-up; the report says the extracted list contained no false positives.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/9af81f24-1ac8-4a6a-b543-1959eafc22b0"},{"id":"633e805b-8efa-4b9f-8667-7b79bb93072c","criterion":"chat-with-notes-ask-questions","criterion_name":"Chat with Notes / Ask Questions","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"The chat/Q&A surface gives grounded answers from the meeting record: on the tool-access question it says access was confirmed that day, cites both the notes and transcript, and identifies rerunning testing as the next step.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/633e805b-8efa-4b9f-8667-7b79bb93072c"},{"id":"e181aa66-75ab-4b14-a804-0bb3f6481466","criterion":"editability","criterion_name":"Editability","rank_role":"context","verdict":"mixed","score":null,"score_total":null,"note":"Granola supports editing the summary layer, but transcript text is not editable in-app; the report describes transcript correction as locked by design, so users can fix notes and action items but not the underlying transcript.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/e181aa66-75ab-4b14-a804-0bb3f6481466"},{"id":"c83c6126-1c92-4f50-a102-cbdfe3713f0f","criterion":"join-method-reliability","criterion_name":"Join Method & Reliability","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The botless desktop capture recorded the full ~25-minute meeting end-to-end without visible dropouts; the transcript reaches the call’s closing lines, indicating uninterrupted capture rather than a mid-call failure.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/c83c6126-1c92-4f50-a102-cbdfe3713f0f"},{"id":"a5c429f5-458e-4d15-807a-a871c71e38be","criterion":"search-across-notes","criterion_name":"Search Across Notes","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Granola’s note search returns exact-match results with navigation: a query for 'api' produced a 1/1 hit and highlighted the matched word in the transcript, so keyword lookup works directly inside the meeting note.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/a5c429f5-458e-4d15-807a-a871c71e38be"},{"id":"73b827a9-4822-4d1e-a8ad-b6853c5eab4d","criterion":"speaker-diarization","criterion_name":"Speaker Diarization","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Granola’s default capture does not attribute speakers: the settings panel shows Speaker tags switched off, and the transcript excerpt is a plain text wall with no speaker labels, so diarization is absent unless the user manually enables it.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/73b827a9-4822-4d1e-a8ad-b6853c5eab4d"},{"id":"d9b8082d-37a5-4e94-8412-e2a6d59d4aa3","criterion":"summary-quality","criterion_name":"Summary Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Granola produces skimmable summaries with named sections; the meeting output is organized into at least three top-level sections, including Use Case Status and Review Progress, Tool Research and Publishing, and Access Tracker Updates.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/d9b8082d-37a5-4e94-8412-e2a6d59d4aa3"},{"id":"e10b1e9f-289c-4e45-bb24-24d7b3063e1e","criterion":"topic-segmentation","criterion_name":"Topic Segmentation","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"The notes are split into named topic sections instead of one long blob; the published section header 'Diagram Animation and Other Use Cases' and the report’s multi-section summary structure show logical breakpoints for navigation.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/e10b1e9f-289c-4e45-bb24-24d7b3063e1e"}],"appears_in":[{"page_type":"ranking","slug":"ai-meeting-notetakers","title":"Best AI Meeting Notetakers for Accurate Transcripts, Summaries, and Action Items","url":"https://aidemos.com/best/ai-meeting-notetakers","binding":"run"}],"same_scenario":[{"id":"1c4f4fe3-3ed8-490d-992b-a25ddd604669","tool":"fathom","tool_name":"Fathom","verdict":"mixed","score":null,"score_total":null,"note":"Fathom's transcript mostly preserves the meeting's names and technical content, but the report records one confirmed name-level error: \"Mahreen\" was rendered as \"Meryl.\""},{"id":"d4b52ad2-b536-4580-99fa-644f4b772a93","tool":"fellow","tool_name":"Fellow","verdict":"worked","score":null,"score_total":null,"note":"The transcript was near-clean: the tool captured nearly all names, technical jargon, and numbers correctly, with no significant misheard terms or hallucinations observed in the tested meeting."},{"id":"bac222c7-f7cf-4041-8449-df5aa22c88ac","tool":"fireflies-ai","tool_name":"Fireflies.ai","verdict":"worked","score":null,"score_total":null,"note":"Generated a timestamped transcript view, and the report says the full transcript was very accurate: nearly all names, tools, jargon, and numbers were captured correctly with no significant misheard terms or hallucinations."},{"id":"624ffd2e-c9da-4fb4-acc3-041a5b115a50","tool":"happyscribe","tool_name":"HappyScribe","verdict":"mixed","score":null,"score_total":null,"note":"On this ~25-minute multi-speaker standup, HappyScribe captured the vast majority of names, tools, and jargon correctly, and the report records only 1–2 misheard words."},{"id":"5a191ed7-02a9-4979-931d-9219e0b75fce","tool":"meetgeek","tool_name":"MeetGeek","verdict":"worked","score":null,"score_total":null,"note":"It transcribes a normal ~25-minute, ~10-active-speaker engineering standup mostly accurately, with only minor proper-noun/term drift noted in the report; one example given is \"Madin\" being misheard for \"Mahreen\"."},{"id":"893f5543-bd03-4e2a-ba30-f6426940628b","tool":"notta","tool_name":"Notta","verdict":"worked","score":null,"score_total":null,"note":"Notta’s transcript capture was accurate on the evaluated standup: the report says it correctly captured names, tool names, numbers, and engineering jargon with no significant word-level errors, silent hallucinations, or misheard terms."},{"id":"0efa5738-c840-44ba-a72a-b35d289e70fb","tool":"otter-ai","tool_name":"Otter.ai","verdict":"worked","score":null,"score_total":null,"note":"Otter generated a full transcript for the standup and, per the report, captured names, tool names, jargon, and numbers correctly with minimal errors, making the transcript reliable for reference."}]}