{"observation":{"id":"5928b605-59f7-4372-a4b7-914060123f74","tool":"fellow","tool_name":"Fellow","criterion":"speaker-diarization","criterion_name":"Speaker Diarization","criterion_definition":"Correctly attributes who said what across a multi-speaker standup.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"Correctly attributing who said what is part of making the transcript and notes trustworthy in multi-speaker meetings. (3 of 3 judges)","scenario":"ai-demos-daily-standup-31-july-2026","scenario_name":"AI Demos Daily Standup — 31 July 2026","group_tag":"ai-meeting-notetaker","scenario_description":"A real 25-minute technical engineering daily standup with 14 attendees and about 10 active speakers, used as the single parallel-capture meeting for evaluating AI meeting notetakers on transcription, diarization, summaries, action items, search/chat, and collaboration features.","modality":"image","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://cdn.futuresmart.ai/public/aidemos/547dd6f13e4a420fa8ad7bf2c88c7612.png?v=1","role":"input","filename":"31-july-meeting-screenshot.png"}],"stresses":["Transcription accuracy for real names, tool names, numbers, and technical jargon","Speaker diarization across multiple active speakers","Robustness to overlapping speech, crosstalk, and rapid turn-taking","Join reliability for bot-based and botless capture","Summary quality on identical source material","Action-item extraction with correct owners and commitments","Topic segmentation of standup updates","Search and chat grounded in the meeting content","Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage"],"verdict":"worked","score":null,"score_total":null,"note":"The transcript attributed speaker turns correctly across the standup, with all ~10 speakers labeled by name and no attribution errors or generic labels reported.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/f6c095044c7a4498b82f300b683163dd.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/7476ae6a377e411094940f94f34b21f9.png?v=1","role":"output","alt":null}],"run_id":"ace58582-3d1e-48ee-996c-9b3cd03f27a2","study_title":"AI Meeting Notetakers — Capture Accurate Transcripts, Summaries & Action Items From Live Calls","study_kind":"generation","research_task":"86baxegnv","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://cdn.futuresmart.ai/public/aidemos/547dd6f13e4a420fa8ad7bf2c88c7612.png?v=1","filename":"31-july-meeting-screenshot.png","alt":"AI Demos Daily Standup — 31 July 2026","role":"input"}],"modality":"image","stresses":["Transcription accuracy for real names, tool names, numbers, and technical jargon","Speaker diarization across multiple active speakers","Robustness to overlapping speech, crosstalk, and rapid turn-taking","Join reliability for bot-based and botless capture","Summary quality on identical source material","Action-item extraction with correct owners and commitments","Topic segmentation of standup updates","Search and chat grounded in the meeting content","Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage"]},"tool_page_slug":"fellow","tool_url":"https://aidemos.com/tools/fellow","permalink":"https://aidemos.com/evidence/5928b605-59f7-4372-a4b7-914060123f74","api_url":"https://ai.aidemos.com/v1/observations/5928b605-59f7-4372-a4b7-914060123f74"},"peers":[{"id":"e9695d25-df6e-448b-b185-4ad01e885bdf","tool":"fathom","tool_name":"Fathom","verdict":"mixed","score":null,"score_total":null,"note":"Fathom separates most speakers correctly in a busy multi-speaker standup, but the report observed one rapid-transition segment where two speakers' lines were merged into a single speaker block.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/e6b553ff015d4c309c161b0bb12a1660.png?v=1","evidence_url":"https://aidemos.com/evidence/e9695d25-df6e-448b-b185-4ad01e885bdf"},{"id":"a1eae6d9-4eff-41c0-9357-c4b9c8e6f7fb","tool":"fireflies-ai","tool_name":"Fireflies.ai","verdict":"worked","score":null,"score_total":null,"note":"Attributed consecutive turns to distinct speakers in the transcript, and the report says speaker identification was almost complete with only minor attribution errors.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/04c680b706994807aa25f0135512b6ee.png?v=1","evidence_url":"https://aidemos.com/evidence/a1eae6d9-4eff-41c0-9357-c4b9c8e6f7fb"},{"id":"73b827a9-4822-4d1e-a8ad-b6853c5eab4d","tool":"granola","tool_name":"Granola","verdict":"failed","score":null,"score_total":null,"note":"Granola’s default capture does not attribute speakers: the settings panel shows Speaker tags switched off, and the transcript excerpt is a plain text wall with no speaker labels, so diarization is absent unless the user manually enables it.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/ecd4eac38c5f4f1f9e0d671349a4f71a.png?v=1","evidence_url":"https://aidemos.com/evidence/73b827a9-4822-4d1e-a8ad-b6853c5eab4d"},{"id":"472fb01a-44b0-4713-b649-40ee52da5802","tool":"happyscribe","tool_name":"HappyScribe","verdict":"mixed","score":null,"score_total":null,"note":"It identified most speakers, but the report says multiple transcript lines were assigned to the wrong speaker, so speaker-to-statement mapping was not fully reliable across transitions.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/64855a79c2564122bac93c704797f365.png?v=1","evidence_url":"https://aidemos.com/evidence/472fb01a-44b0-4713-b649-40ee52da5802"},{"id":"4f7e5b21-ff4f-4846-9a07-3219c0659681","tool":"meetgeek","tool_name":"MeetGeek","verdict":"mixed","score":null,"score_total":null,"note":"It identifies most speakers in a multi-speaker standup, but leaves at least one utterance as \"Unknown speaker\" and misattributes some lines to the wrong speaker, so attribution is not fully reliable.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/a909947fdbe64d41b2bc10dca0f44193.png?v=1","evidence_url":"https://aidemos.com/evidence/4f7e5b21-ff4f-4846-9a07-3219c0659681"},{"id":"c1fc0701-9a26-4f73-a1f2-d4dfae6a7109","tool":"notta","tool_name":"Notta","verdict":"struggled","score":null,"score_total":null,"note":"The transcript contained a line labeled with another notetaker’s name (HappyScribe), which indicates cross-tool contamination or labeling error and breaks speaker attribution for that segment.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/665e1b0e39b34e40bb46f742733e7db7.png?v=1","evidence_url":"https://aidemos.com/evidence/c1fc0701-9a26-4f73-a1f2-d4dfae6a7109"},{"id":"9552a2b6-0bc3-4802-8f1b-f9cf98c13b46","tool":"otter-ai","tool_name":"Otter.ai","verdict":"failed","score":null,"score_total":null,"note":"Otter’s diarization was effectively unusable in this multi-speaker standup: only 1 of about 10 active speakers was identified by name, while the other 9 were left as generic labels or unattributed, which the report summarizes as a 90% failure rate.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/c2c822851e2f4e71a1ae3b14247c11d3.png?v=1","evidence_url":"https://aidemos.com/evidence/9552a2b6-0bc3-4802-8f1b-f9cf98c13b46"}],"other_criteria":[{"id":"21587b7d-dc6e-4aa1-8ee4-4e5f13641a3e","criterion":"action-item-extraction","criterion_name":"Action-Item Extraction","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"Action-item extraction was mostly correct, with real commitments and proper owner assignment for most items, but one real action item was misplaced from Mahreen to Anshika; the report states a 95%+ capture rate.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/21587b7d-dc6e-4aa1-8ee4-4e5f13641a3e"},{"id":"b8fe3f78-ae35-4ccf-b887-6188bac5ad2f","criterion":"chat-with-notes-ask-questions","criterion_name":"Chat with Notes / Ask Questions","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Ask Fellow returned grounded answers to natural-language questions against the meeting notes, and the tested query produced a cited response rather than an unsupported hallucination.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/b8fe3f78-ae35-4ccf-b887-6188bac5ad2f"},{"id":"a47ae17f-dfbd-439c-9ba9-e3153156cca8","criterion":"editability","criterion_name":"Editability","rank_role":"context","verdict":"mixed","score":null,"score_total":null,"note":"Summary and action items are editable inline before sharing, but the transcript itself is locked for audit-trail purposes, so editing is only partial.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/a47ae17f-dfbd-439c-9ba9-e3153156cca8"},{"id":"5d9e1f80-f657-4216-8e73-e81047b806aa","criterion":"join-method-reliability","criterion_name":"Join Method & Reliability","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The bot auto-joined via Google Calendar integration and stayed connected for the full meeting capture without drops or disconnections, including the late closing portion of the call.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/5d9e1f80-f657-4216-8e73-e81047b806aa"},{"id":"a8a8aaba-e0e1-44b8-9cce-4fa91c5e5213","criterion":"search-across-notes","criterion_name":"Search Across Notes","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Transcript search works with exact timestamp retrieval for matching terms, letting users jump to precise moments within the meeting; the report notes cross-meeting search was untested.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/a8a8aaba-e0e1-44b8-9cce-4fa91c5e5213"},{"id":"b22ab795-4b0c-4939-80f6-1f9daca9f6e6","criterion":"sharing-without-registration","criterion_name":"Sharing Without Registration","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"Shared recap links can be opened by anyone with the link without creating a Fellow account, and the share modal also offers optional password protection for viewers outside the workspace.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/b22ab795-4b0c-4939-80f6-1f9daca9f6e6"},{"id":"d89239cd-ae18-4967-90d8-e968276badf1","criterion":"summary-quality","criterion_name":"Summary Quality","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The meeting recap was reported as clearly structured and complete, with the key decisions and discussion points preserved and nothing important dropped.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/d89239cd-ae18-4967-90d8-e968276badf1"},{"id":"b69458f2-5b1d-4915-bf53-39ab59f059d2","criterion":"topic-segmentation","criterion_name":"Topic Segmentation","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"The tool broke the standup into logical topic sections with clear headers and separated discussion points, rather than leaving the meeting as one undifferentiated blob.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/b69458f2-5b1d-4915-bf53-39ab59f059d2"},{"id":"d4b52ad2-b536-4580-99fa-644f4b772a93","criterion":"transcription-accuracy","criterion_name":"Transcription Accuracy","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The transcript was near-clean: the tool captured nearly all names, technical jargon, and numbers correctly, with no significant misheard terms or hallucinations observed in the tested meeting.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/d4b52ad2-b536-4580-99fa-644f4b772a93"}],"appears_in":[{"page_type":"ranking","slug":"ai-meeting-notetakers","title":"Best AI Meeting Notetakers for Accurate Transcripts, Summaries, and Action Items","url":"https://aidemos.com/best/ai-meeting-notetakers","binding":"run"}],"same_scenario":[{"id":"e9695d25-df6e-448b-b185-4ad01e885bdf","tool":"fathom","tool_name":"Fathom","verdict":"mixed","score":null,"score_total":null,"note":"Fathom separates most speakers correctly in a busy multi-speaker standup, but the report observed one rapid-transition segment where two speakers' lines were merged into a single speaker block."},{"id":"a1eae6d9-4eff-41c0-9357-c4b9c8e6f7fb","tool":"fireflies-ai","tool_name":"Fireflies.ai","verdict":"worked","score":null,"score_total":null,"note":"Attributed consecutive turns to distinct speakers in the transcript, and the report says speaker identification was almost complete with only minor attribution errors."},{"id":"73b827a9-4822-4d1e-a8ad-b6853c5eab4d","tool":"granola","tool_name":"Granola","verdict":"failed","score":null,"score_total":null,"note":"Granola’s default capture does not attribute speakers: the settings panel shows Speaker tags switched off, and the transcript excerpt is a plain text wall with no speaker labels, so diarization is absent unless the user manually enables it."},{"id":"472fb01a-44b0-4713-b649-40ee52da5802","tool":"happyscribe","tool_name":"HappyScribe","verdict":"mixed","score":null,"score_total":null,"note":"It identified most speakers, but the report says multiple transcript lines were assigned to the wrong speaker, so speaker-to-statement mapping was not fully reliable across transitions."},{"id":"4f7e5b21-ff4f-4846-9a07-3219c0659681","tool":"meetgeek","tool_name":"MeetGeek","verdict":"mixed","score":null,"score_total":null,"note":"It identifies most speakers in a multi-speaker standup, but leaves at least one utterance as \"Unknown speaker\" and misattributes some lines to the wrong speaker, so attribution is not fully reliable."},{"id":"c1fc0701-9a26-4f73-a1f2-d4dfae6a7109","tool":"notta","tool_name":"Notta","verdict":"struggled","score":null,"score_total":null,"note":"The transcript contained a line labeled with another notetaker’s name (HappyScribe), which indicates cross-tool contamination or labeling error and breaks speaker attribution for that segment."},{"id":"9552a2b6-0bc3-4802-8f1b-f9cf98c13b46","tool":"otter-ai","tool_name":"Otter.ai","verdict":"failed","score":null,"score_total":null,"note":"Otter’s diarization was effectively unusable in this multi-speaker standup: only 1 of about 10 active speakers was identified by name, while the other 9 were left as generic labels or unattributed, which the report summarizes as a 90% failure rate."}]}