{"observation":{"id":"295ea299-5f4d-480e-8ff7-36708e2bc9f7","tool":"otter-ai","tool_name":"Otter.ai","criterion":"chat-with-notes-ask-questions","criterion_name":"Chat with Notes / Ask Questions","criterion_definition":"Gives grounded answers with the cited moment and admits when unknown.","criterion_evidence_type":"transformation","criterion_rank_role":"context","criterion_rank_role_reason":"Q&A over notes is useful, but it is an add-on to the capture and summarization job rather than a core measure of it. (3 of 3 judges)","scenario":"ai-demos-daily-standup-31-july-2026","scenario_name":"AI Demos Daily Standup — 31 July 2026","group_tag":"ai-meeting-notetaker","scenario_description":"A real 25-minute technical engineering daily standup with 14 attendees and about 10 active speakers, used as the single parallel-capture meeting for evaluating AI meeting notetakers on transcription, diarization, summaries, action items, search/chat, and collaboration features.","modality":"image","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/547dd6f13e4a420fa8ad7bf2c88c7612.png?v=1","role":"input","filename":"31-july-meeting-screenshot.png"}],"stresses":["Transcription accuracy for real names, tool names, numbers, and technical jargon","Speaker diarization across multiple active speakers","Robustness to overlapping speech, crosstalk, and rapid turn-taking","Join reliability for bot-based and botless capture","Summary quality on identical source material","Action-item extraction with correct owners and commitments","Topic segmentation of standup updates","Search and chat grounded in the meeting content","Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage"],"verdict":"worked","score":null,"score_total":null,"note":"Otter’s AI Chat answered meeting questions with a grounded response and a specific timestamp, and the report says the answers were cited and free of hallucinations in the tested queries.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/e5aefd3684124771a718737d9c8d1754.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/e7e2cfb5a5c04734bad1c2aacb423607.png?v=1","role":"output","alt":null}],"run_id":"ace58582-3d1e-48ee-996c-9b3cd03f27a2","study_title":"AI Meeting Notetakers — Capture Accurate Transcripts, Summaries & Action Items From Live Calls","study_kind":"generation","research_task":"86baxegnv","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/547dd6f13e4a420fa8ad7bf2c88c7612.png?v=1","filename":"31-july-meeting-screenshot.png","alt":"AI Demos Daily Standup — 31 July 2026","role":"input"}],"modality":"image","stresses":["Transcription accuracy for real names, tool names, numbers, and technical jargon","Speaker diarization across multiple active speakers","Robustness to overlapping speech, crosstalk, and rapid turn-taking","Join reliability for bot-based and botless capture","Summary quality on identical source material","Action-item extraction with correct owners and commitments","Topic segmentation of standup updates","Search and chat grounded in the meeting content","Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage"]},"tool_page_slug":"otter-ai","tool_url":"https://aidemos.com/tools/otter-ai","permalink":"https://aidemos.com/evidence/295ea299-5f4d-480e-8ff7-36708e2bc9f7","api_url":"https://ai.aidemos.com/v1/observations/295ea299-5f4d-480e-8ff7-36708e2bc9f7"},"peers":[{"id":"070f7f5f-046c-4139-a132-6e00a3b57acc","tool":"fathom","tool_name":"Fathom","verdict":"worked","score":null,"score_total":null,"note":"Ask Fathom answers direct factual questions from the meeting notes with grounded references; for one query it answered that a call was scheduled for 6th August and linked the supporting transcript mention.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/cc0b830d51cd48a2a6c9121c331b1626.png?v=1","evidence_url":"https://aidemos.com/evidence/070f7f5f-046c-4139-a132-6e00a3b57acc"},{"id":"b8fe3f78-ae35-4ccf-b887-6188bac5ad2f","tool":"fellow","tool_name":"Fellow","verdict":"worked","score":null,"score_total":null,"note":"Ask Fellow returned grounded answers to natural-language questions against the meeting notes, and the tested query produced a cited response rather than an unsupported hallucination.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/553c970153974fc9a12d999ec7a3b888.png?v=1","evidence_url":"https://aidemos.com/evidence/b8fe3f78-ae35-4ccf-b887-6188bac5ad2f"},{"id":"9080d7d0-7707-4182-b68c-0e1f61c09488","tool":"fireflies-ai","tool_name":"Fireflies.ai","verdict":"worked","score":null,"score_total":null,"note":"AskFred answered a natural-language question with a specific grounded response ('August 6th') and relevant context, with no hallucination reported in the tested query.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/5d6339008cea4142b2733ccea5fe3c23.png?v=1","evidence_url":"https://aidemos.com/evidence/9080d7d0-7707-4182-b68c-0e1f61c09488"},{"id":"633e805b-8efa-4b9f-8667-7b79bb93072c","tool":"granola","tool_name":"Granola","verdict":"worked","score":null,"score_total":null,"note":"The chat/Q&A surface gives grounded answers from the meeting record: on the tool-access question it says access was confirmed that day, cites both the notes and transcript, and identifies rerunning testing as the next step.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/8b17d8ed350645bfb4a942128365177f.png?v=1","evidence_url":"https://aidemos.com/evidence/633e805b-8efa-4b9f-8667-7b79bb93072c"},{"id":"09e56c77-5be8-4485-90bc-76a453f3197a","tool":"happyscribe","tool_name":"HappyScribe","verdict":"worked","score":null,"score_total":null,"note":"AI chat answers a meeting question with a grounded transcript-backed response, returning that the call was scheduled for '6th August' and explicitly indicating it is reading the transcription.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/04410e535a0e4f95add9df3c5cd7e79c.png?v=1","evidence_url":"https://aidemos.com/evidence/09e56c77-5be8-4485-90bc-76a453f3197a"},{"id":"75aaf0b0-c9d5-4376-8b81-5b6f2574c0fb","tool":"meetgeek","tool_name":"MeetGeek","verdict":"mixed","score":null,"score_total":null,"note":"It answers direct grounded questions correctly, but the report records an incorrect answer on a speaker-dependent scheduling question, so chat is reliable for simple queries but weaker when attribution/context matters.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/e84f2cebcccd4af092aff8637baab736.png?v=1","evidence_url":"https://aidemos.com/evidence/75aaf0b0-c9d5-4376-8b81-5b6f2574c0fb"},{"id":"b58378a0-5d37-4751-a711-f2d028fa4401","tool":"notta","tool_name":"Notta","verdict":"worked","score":null,"score_total":null,"note":"The Q&A interface answered a natural-language question with a grounded response from the meeting record, including the specific date \"6th August,\" and the report observed no hallucinations.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/a822259cb8bf403a96a2d5814c1b7036.png?v=1","evidence_url":"https://aidemos.com/evidence/b58378a0-5d37-4751-a711-f2d028fa4401"}],"other_criteria":[{"id":"914e606f-d39b-4941-8b42-0f3b3653ee40","criterion":"action-item-extraction","criterion_name":"Action-Item Extraction","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Otter extracted action items, including at least one due-today API-related task with an assignee, but the report says most items were left without an owner and duplicate entries also appeared, so the output needed manual cleanup before delegation.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/914e606f-d39b-4941-8b42-0f3b3653ee40"},{"id":"0364fbbc-9644-47cc-97a6-60174baab699","criterion":"editability","criterion_name":"Editability","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Otter exposes inline editing controls for transcript and summary outputs before sharing, so wrong content can be corrected in-product rather than only exported as-is.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/0364fbbc-9644-47cc-97a6-60174baab699"},{"id":"1ec8c4c4-913b-4e17-8150-83589c5c2f7b","criterion":"join-method-reliability","criterion_name":"Join Method & Reliability","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Otter’s bot joined Google Meet successfully and stayed connected through the full ~25-minute call with no mid-call dropout or ejection, so the meeting was captured end to end.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/1ec8c4c4-913b-4e17-8150-83589c5c2f7b"},{"id":"68c18886-3c73-4ed3-bd68-92814dc4b10b","criterion":"search-across-notes","criterion_name":"Search Across Notes","rank_role":"context","verdict":"mixed","score":null,"score_total":null,"note":"Otter does not show a direct transcript keyword-search workflow in the transcript view; lookup is routed through AI Chat instead, where a natural-language timestamp question returned a specific answer at 0:07:06.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/68c18886-3c73-4ed3-bd68-92814dc4b10b"},{"id":"bc9e6d91-ebd3-4a9e-99eb-5cc194791f45","criterion":"sharing-without-registration","criterion_name":"Sharing Without Registration","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"A shared meeting transcript opened without requiring sign-in, exposing the meeting title, metadata, and transcript snippets; the share dialog also offers restricted access and link-copy controls.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/bc9e6d91-ebd3-4a9e-99eb-5cc194791f45"},{"id":"9552a2b6-0bc3-4802-8f1b-f9cf98c13b46","criterion":"speaker-diarization","criterion_name":"Speaker Diarization","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Otter’s diarization was effectively unusable in this multi-speaker standup: only 1 of about 10 active speakers was identified by name, while the other 9 were left as generic labels or unattributed, which the report summarizes as a 90% failure rate.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/9552a2b6-0bc3-4802-8f1b-f9cf98c13b46"},{"id":"85b52298-f079-4410-b519-3fc1d34593a0","criterion":"summary-quality","criterion_name":"Summary Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Otter produced a clear, structured meeting summary that the report says covered the key decisions and discussion points without dropping anything important, and the summary page loaded with organized sections like Overview and Action Items.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/85b52298-f079-4410-b519-3fc1d34593a0"},{"id":"eedc69a7-0929-4cbd-a001-6122eed42379","criterion":"topic-segmentation","criterion_name":"Topic Segmentation","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Otter broke the standup into useful topic sections rather than one blob, with named headings such as Issue Task Assignments and Status Updates and Error Resolution and Task Link Sharing, making the summary skimmable.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/eedc69a7-0929-4cbd-a001-6122eed42379"},{"id":"0efa5738-c840-44ba-a72a-b35d289e70fb","criterion":"transcription-accuracy","criterion_name":"Transcription Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Otter generated a full transcript for the standup and, per the report, captured names, tool names, jargon, and numbers correctly with minimal errors, making the transcript reliable for reference.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/0efa5738-c840-44ba-a72a-b35d289e70fb"}],"appears_in":[{"page_type":"ranking","slug":"ai-meeting-notetakers","title":"Best AI Meeting Notetakers for Accurate Transcripts, Summaries, and Action Items","url":"https://aidemos.com/best/ai-meeting-notetakers","binding":"run"}],"same_scenario":[{"id":"070f7f5f-046c-4139-a132-6e00a3b57acc","tool":"fathom","tool_name":"Fathom","verdict":"worked","score":null,"score_total":null,"note":"Ask Fathom answers direct factual questions from the meeting notes with grounded references; for one query it answered that a call was scheduled for 6th August and linked the supporting transcript mention."},{"id":"b8fe3f78-ae35-4ccf-b887-6188bac5ad2f","tool":"fellow","tool_name":"Fellow","verdict":"worked","score":null,"score_total":null,"note":"Ask Fellow returned grounded answers to natural-language questions against the meeting notes, and the tested query produced a cited response rather than an unsupported hallucination."},{"id":"9080d7d0-7707-4182-b68c-0e1f61c09488","tool":"fireflies-ai","tool_name":"Fireflies.ai","verdict":"worked","score":null,"score_total":null,"note":"AskFred answered a natural-language question with a specific grounded response ('August 6th') and relevant context, with no hallucination reported in the tested query."},{"id":"633e805b-8efa-4b9f-8667-7b79bb93072c","tool":"granola","tool_name":"Granola","verdict":"worked","score":null,"score_total":null,"note":"The chat/Q&A surface gives grounded answers from the meeting record: on the tool-access question it says access was confirmed that day, cites both the notes and transcript, and identifies rerunning testing as the next step."},{"id":"09e56c77-5be8-4485-90bc-76a453f3197a","tool":"happyscribe","tool_name":"HappyScribe","verdict":"worked","score":null,"score_total":null,"note":"AI chat answers a meeting question with a grounded transcript-backed response, returning that the call was scheduled for '6th August' and explicitly indicating it is reading the transcription."},{"id":"75aaf0b0-c9d5-4376-8b81-5b6f2574c0fb","tool":"meetgeek","tool_name":"MeetGeek","verdict":"mixed","score":null,"score_total":null,"note":"It answers direct grounded questions correctly, but the report records an incorrect answer on a speaker-dependent scheduling question, so chat is reliable for simple queries but weaker when attribution/context matters."},{"id":"b58378a0-5d37-4751-a711-f2d028fa4401","tool":"notta","tool_name":"Notta","verdict":"worked","score":null,"score_total":null,"note":"The Q&A interface answered a natural-language question with a grounded response from the meeting record, including the specific date \"6th August,\" and the report observed no hallucinations."}]}