{"observation":{"id":"f0009fdb-31cf-4281-aceb-887be1538083","tool":"mistral-ai","tool_name":"Mistral AI","criterion":"markdown-quality","criterion_name":"Markdown Quality","criterion_definition":"Produces clean, well-structured, usable markdown rather than a flat text dump.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"The ranking is specifically about producing clean Markdown, so the quality and usability of the Markdown output directly measures success. (3 of 3 judges)","scenario":"scanned-research-paper","scenario_name":"Scanned Research Paper","group_tag":"scanned-research-paper","scenario_description":"An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.","modality":"pdf","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/convert-a-complex-pdf-into-clean-markdow-scanned-research-pdf-7b86de49784d.pdf","role":"input","filename":"Scanned Research PDF.pdf"}],"stresses":["OCR on scanned pages","Multi-column reading order","Figure and chart handling","Table reconstruction from scans","Caption association","Reference extraction","Overall document structure retention"],"verdict":"worked","score":null,"score_total":null,"note":"The export is packaged as usable markdown files in a ZIP, with both overall and page-wise outputs available for inspection.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/45e3533a31c246b29e0ec6aaa98438e4.pdf?v=1","role":"input","alt":"45e3533a31c246b29e0ec6aaa98438e4.pdf"},{"url":"https://d3epheqghktydj.cloudfront.net/research-media-mistral-ai-scanned-pdf-output-zip-c2a22ad51169.zip","role":"output","alt":null}],"run_id":"6e3160de-fe46-4b45-b071-72560b5c5d0e","study_title":"Convert a Complex PDF into Clean Markdown with an API","study_kind":"generation","research_task":"86b9h7t37","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://d3epheqghktydj.cloudfront.net/convert-a-complex-pdf-into-clean-markdow-scanned-research-pdf-7b86de49784d.pdf","filename":"Scanned Research PDF.pdf","alt":"Scanned Research Paper","role":"input"}],"modality":"pdf","stresses":["OCR on scanned pages","Multi-column reading order","Figure and chart handling","Table reconstruction from scans","Caption association","Reference extraction","Overall document structure retention"]},"tool_page_slug":"mistral-ai","tool_url":"https://aidemos.com/tools/mistral-ai","permalink":"https://aidemos.com/evidence/f0009fdb-31cf-4281-aceb-887be1538083","api_url":"https://ai.aidemos.com/v1/observations/f0009fdb-31cf-4281-aceb-887be1538083"},"peers":[{"id":"39ecbd6f-1599-4ba1-ae17-16b29b9c07cb","tool":"reducto","tool_name":"Reducto","verdict":"worked","score":null,"score_total":null,"note":"Uses no invented tags or malformed markdown; the only HTML seen is legitimate <br /> inside table cells, so the syntax stays clean even though the document does not surface real heading markup.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/5eec926a2ff942bf98539291b857234b.png?v=1","evidence_url":"https://aidemos.com/evidence/39ecbd6f-1599-4ba1-ae17-16b29b9c07cb"},{"id":"4fca39e0-3fcb-4357-9bb8-48233637ab9f","tool":"tensorlake","tool_name":"Tensorlake","verdict":"worked","score":null,"score_total":null,"note":"Renders the scanned-paper extraction as structured markdown in the tool workflow, rather than only exposing raw OCR text.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/f6a345881812480ea988f2b8044108ce.mp4?v=1","evidence_url":"https://aidemos.com/evidence/4fca39e0-3fcb-4357-9bb8-48233637ab9f"}],"other_criteria":[{"id":"d572c22d-d6ae-4a81-b848-53c180d00575","criterion":"complex-document-handling","criterion_name":"Complex Document Handling","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The tool processes a scanned multi-column research paper end-to-end and returns OCR text, tables, and embedded chart assets in page-wise markdown output.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/d572c22d-d6ae-4a81-b848-53c180d00575"},{"id":"a70898cf-5c39-40b0-a6b8-0f49144c92b1","criterion":"reading-order-structure","criterion_name":"Reading Order & Structure","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Section hierarchy and reading flow are preserved in the scanned paper, keeping headings and supporting paragraphs correctly connected despite the multi-column layout.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/a70898cf-5c39-40b0-a6b8-0f49144c92b1"},{"id":"cc8096bb-19af-4b8f-901f-657634e78a25","criterion":"reading-order-structure","criterion_name":"Reading Order & Structure","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"The opening page loses the distinction between the document title and the abstract, flattening the semantic organization of the first page.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/cc8096bb-19af-4b8f-901f-657634e78a25"},{"id":"527f7e79-f463-4b6e-b839-c5153a59a9f3","criterion":"table-preservation","criterion_name":"Table Preservation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The multicolumn table is reconstructed without losing its overall layout logic, so the table structure remains readable in the parsed output.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/527f7e79-f463-4b6e-b839-c5153a59a9f3"},{"id":"de43911e-dc72-4aee-a02a-2145d2d7acfb","criterion":"table-preservation","criterion_name":"Table Preservation","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Broken column boundaries and disrupted value alignment make the reconstructed table significantly less faithful to the source.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/de43911e-dc72-4aee-a02a-2145d2d7acfb"},{"id":"932e3296-1f2a-4d58-9bb4-5c7bb1e269a7","criterion":"text-ocr-completeness","criterion_name":"Text & OCR Completeness","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The report says the parser recovers much of the underlying text from the scanned paper, but it does not present a measured completeness rate and the first-page hierarchy is still lossy.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/932e3296-1f2a-4d58-9bb4-5c7bb1e269a7"},{"id":"0f3a42b1-0d18-4374-bbeb-cd6776d7b79e","criterion":"visual-content-retention","criterion_name":"Visual Content Retention","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Charts are exposed through both page-wise markdown files and the extracted visual assets, keeping the visual content linked to its original document location.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/0f3a42b1-0d18-4374-bbeb-cd6776d7b79e"}],"appears_in":[{"page_type":"ranking","slug":"pdf-to-markdown-apis","title":"Best AI Tools to Convert Complex PDFs into Clean Markdown with an API","url":"https://aidemos.com/best/pdf-to-markdown-apis","binding":"run"}],"same_scenario":[{"id":"39ecbd6f-1599-4ba1-ae17-16b29b9c07cb","tool":"reducto","tool_name":"Reducto","verdict":"worked","score":null,"score_total":null,"note":"Uses no invented tags or malformed markdown; the only HTML seen is legitimate <br /> inside table cells, so the syntax stays clean even though the document does not surface real heading markup."},{"id":"4fca39e0-3fcb-4357-9bb8-48233637ab9f","tool":"tensorlake","tool_name":"Tensorlake","verdict":"worked","score":null,"score_total":null,"note":"Renders the scanned-paper extraction as structured markdown in the tool workflow, rather than only exposing raw OCR text."}]}