{"observation":{"id":"1ff9ad63-6629-4bd0-a643-ff827f73c602","tool":"nutrient-io","tool_name":"Nutrient.io","criterion":"text-ocr-completeness","criterion_name":"Text & OCR Completeness","criterion_definition":"Extracts all readable content, including scanned pages, with accurate OCR and minimal omissions.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"If the tool misses readable text or fails on scanned pages, it has not actually converted the PDF faithfully into Markdown. (3 of 3 judges)","scenario":"scanned-research-paper","scenario_name":"Scanned Research Paper","group_tag":"scanned-research-paper","scenario_description":"An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.","modality":"pdf","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/convert-a-complex-pdf-into-clean-markdow-scanned-research-pdf-7b86de49784d.pdf","role":"input","filename":"Scanned Research PDF.pdf"}],"stresses":["OCR on scanned pages","Multi-column reading order","Figure and chart handling","Table reconstruction from scans","Caption association","Reference extraction","Overall document structure retention"],"verdict":"worked","score":null,"score_total":null,"note":"Recovers readable text from scanned pages, including the abstract, keywords, author affiliations, title, and opening paragraphs on the first page.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/94e13a264c0c464fba32bd78b5bb4b27.png?v=1","role":"input","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/404a1c67c3d34041bd71c13d50b7dd3f.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/e3e34d0ef422468ead6dc13cfc32018a.png?v=1","role":"output","alt":null}],"run_id":"6e3160de-fe46-4b45-b071-72560b5c5d0e","study_title":"Convert a Complex PDF into Clean Markdown with an API","study_kind":"generation","research_task":"86b9h7t37","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://d3epheqghktydj.cloudfront.net/convert-a-complex-pdf-into-clean-markdow-scanned-research-pdf-7b86de49784d.pdf","filename":"Scanned Research PDF.pdf","alt":"Scanned Research Paper","role":"input"}],"modality":"pdf","stresses":["OCR on scanned pages","Multi-column reading order","Figure and chart handling","Table reconstruction from scans","Caption association","Reference extraction","Overall document structure retention"]},"tool_page_slug":"nutrient-io","tool_url":"https://aidemos.com/tools/nutrient-io","permalink":"https://aidemos.com/evidence/1ff9ad63-6629-4bd0-a643-ff827f73c602","api_url":"https://ai.aidemos.com/v1/observations/1ff9ad63-6629-4bd0-a643-ff827f73c602"},"peers":[{"id":"ec57bfbf-d0f3-4011-803c-9f42ea53397b","tool":"adobe-api","tool_name":"Adobe API","verdict":"worked","score":null,"score_total":null,"note":"Recovers the visible title, abstract, keywords, and opening paragraphs from a scanned USDA forestry report as dense OCR text.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/28052bbc164e4f51be6903dcaccdb477.png?v=1","evidence_url":"https://aidemos.com/evidence/ec57bfbf-d0f3-4011-803c-9f42ea53397b"},{"id":"6125882c-9dbc-46e0-99d4-07315648ef3c","tool":"extend-ai","tool_name":"Extend AI","verdict":"mixed","score":null,"score_total":null,"note":"Detects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/5582cf26df374cdf98a46ff859ba1239.png?v=1","evidence_url":"https://aidemos.com/evidence/6125882c-9dbc-46e0-99d4-07315648ef3c"},{"id":"d162b0b4-78dd-4ace-81af-d47f25041f26","tool":"landing-ai","tool_name":"Landing AI","verdict":"mixed","score":null,"score_total":null,"note":"OCRs most of the scanned cover-page text, including the agency header, report number/date, title, authors, abstract, and keywords, but inserts noisy tokens such as '186153' and 'USA/-' into the title line.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/5db82c36907f41bc8b9bf656cb3c0a90.png?v=1","evidence_url":"https://aidemos.com/evidence/d162b0b4-78dd-4ace-81af-d47f25041f26"},{"id":"04bda6ca-e954-4d0b-93a9-4220eb0b980a","tool":"llamaparse","tool_name":"LlamaParse","verdict":"worked","score":null,"score_total":null,"note":"Recovers dense readable prose from a scanned page-image source, including the section heading and multiple long paragraphs.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/bb646d6406274eb280aa628c32bfa474.png?v=1","evidence_url":"https://aidemos.com/evidence/04bda6ca-e954-4d0b-93a9-4220eb0b980a"},{"id":"932e3296-1f2a-4d58-9bb4-5c7bb1e269a7","tool":"mistral-ai","tool_name":"Mistral AI","verdict":"mixed","score":null,"score_total":null,"note":"The report says the parser recovers much of the underlying text from the scanned paper, but it does not present a measured completeness rate and the first-page hierarchy is still lossy.","artifact_count":3,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/f84ba080d3254487844fc0a3f51f1873.mp4?v=1","evidence_url":"https://aidemos.com/evidence/932e3296-1f2a-4d58-9bb4-5c7bb1e269a7"},{"id":"d9557fa2-03ec-4057-aeb9-17d994dc635d","tool":"pdfvector","tool_name":"PDFVector","verdict":"worked","score":null,"score_total":null,"note":"Successfully parsed a 12-page scanned research paper and produced a long extracted-text preview in 20.6 s using 48 credits.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/83b8ca4cdbc84286a556237324181bdd.mp4?v=1","evidence_url":"https://aidemos.com/evidence/d9557fa2-03ec-4057-aeb9-17d994dc635d"},{"id":"b0876a41-66cd-455c-b38c-421968bf4e16","tool":"reducto","tool_name":"Reducto","verdict":"worked","score":null,"score_total":null,"note":"Converts all 12 scanned pages with no gaps; the page-11-to-page-12 handoff is preserved verbatim, and a separate page-marker output shows pages 1 through 12 present with no missing markers.","artifact_count":4,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/b752a48186ac4761b694699c4c559228.png?v=1","evidence_url":"https://aidemos.com/evidence/b0876a41-66cd-455c-b38c-421968bf4e16"},{"id":"a1cca8ba-01fc-4f63-ab47-bbe9cbd46915","tool":"upstage-ai","tool_name":"Upstage AI","verdict":"worked","score":null,"score_total":null,"note":"OCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/e536615e5d89406ba15aca0a924432d7.png?v=1","evidence_url":"https://aidemos.com/evidence/a1cca8ba-01fc-4f63-ab47-bbe9cbd46915"}],"other_criteria":[{"id":"b50f09d7-7cee-4268-8441-50b46ddeb000","criterion":"complex-document-handling","criterion_name":"Complex Document Handling","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Processes an image-only scanned research paper end to end and returns a parsed markdown output plus preview, showing it can handle a multi-page scanned document.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/b50f09d7-7cee-4268-8441-50b46ddeb000"},{"id":"269c8673-40ed-4da2-a18e-2a17a92616e3","criterion":"reading-order-structure","criterion_name":"Reading Order & Structure","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Misorders the first page so the abstract is placed before the title, showing a layout-reading-order inversion on the scanned article.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/269c8673-40ed-4da2-a18e-2a17a92616e3"},{"id":"0310cd16-a31c-473f-bf94-1da61a08607c","criterion":"reading-order-structure","criterion_name":"Reading Order & Structure","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"The tool preserves section hierarchy and column order in a dense multi-column scanned section, keeping headings aligned with the correct body text.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/0310cd16-a31c-473f-bf94-1da61a08607c"},{"id":"97dfdbc2-3390-4bf7-9d20-af8a5dd33d01","criterion":"table-preservation","criterion_name":"Table Preservation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Largely preserves grouped-column tables, keeping their internal organization intact in the extracted output.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/97dfdbc2-3390-4bf7-9d20-af8a5dd33d01"},{"id":"d9350bf4-9574-4634-8293-5ced2ab60d9c","criterion":"table-preservation","criterion_name":"Table Preservation","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Breaks more complex multi-level tables, with rows and merged cells misaligned or lost once the hierarchy becomes dense.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/d9350bf4-9574-4634-8293-5ced2ab60d9c"},{"id":"474d7192-6964-4d73-9726-1ca2efcb6edc","criterion":"visual-content-retention","criterion_name":"Visual Content Retention","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Extracts the Figure 3 chart's values, but the chart structure and layout are not preserved, so the visualization is flattened into text-like output.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/474d7192-6964-4d73-9726-1ca2efcb6edc"}],"appears_in":[{"page_type":"ranking","slug":"pdf-to-markdown-apis","title":"Best AI Tools to Convert Complex PDFs into Clean Markdown with an API","url":"https://aidemos.com/best/pdf-to-markdown-apis","binding":"run"}],"same_scenario":[{"id":"ec57bfbf-d0f3-4011-803c-9f42ea53397b","tool":"adobe-api","tool_name":"Adobe API","verdict":"worked","score":null,"score_total":null,"note":"Recovers the visible title, abstract, keywords, and opening paragraphs from a scanned USDA forestry report as dense OCR text."},{"id":"6125882c-9dbc-46e0-99d4-07315648ef3c","tool":"extend-ai","tool_name":"Extend AI","verdict":"mixed","score":null,"score_total":null,"note":"Detects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible."},{"id":"d162b0b4-78dd-4ace-81af-d47f25041f26","tool":"landing-ai","tool_name":"Landing AI","verdict":"mixed","score":null,"score_total":null,"note":"OCRs most of the scanned cover-page text, including the agency header, report number/date, title, authors, abstract, and keywords, but inserts noisy tokens such as '186153' and 'USA/-' into the title line."},{"id":"04bda6ca-e954-4d0b-93a9-4220eb0b980a","tool":"llamaparse","tool_name":"LlamaParse","verdict":"worked","score":null,"score_total":null,"note":"Recovers dense readable prose from a scanned page-image source, including the section heading and multiple long paragraphs."},{"id":"932e3296-1f2a-4d58-9bb4-5c7bb1e269a7","tool":"mistral-ai","tool_name":"Mistral AI","verdict":"mixed","score":null,"score_total":null,"note":"The report says the parser recovers much of the underlying text from the scanned paper, but it does not present a measured completeness rate and the first-page hierarchy is still lossy."},{"id":"d9557fa2-03ec-4057-aeb9-17d994dc635d","tool":"pdfvector","tool_name":"PDFVector","verdict":"worked","score":null,"score_total":null,"note":"Successfully parsed a 12-page scanned research paper and produced a long extracted-text preview in 20.6 s using 48 credits."},{"id":"b0876a41-66cd-455c-b38c-421968bf4e16","tool":"reducto","tool_name":"Reducto","verdict":"worked","score":null,"score_total":null,"note":"Converts all 12 scanned pages with no gaps; the page-11-to-page-12 handoff is preserved verbatim, and a separate page-marker output shows pages 1 through 12 present with no missing markers."},{"id":"a1cca8ba-01fc-4f63-ab47-bbe9cbd46915","tool":"upstage-ai","tool_name":"Upstage AI","verdict":"worked","score":null,"score_total":null,"note":"OCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels."}]}