{"observation":{"id":"0310cd16-a31c-473f-bf94-1da61a08607c","tool":"nutrient-io","tool_name":"Nutrient.io","criterion":"reading-order-structure","criterion_name":"Reading Order & Structure","criterion_definition":"Maintains headings, sections, column order, document hierarchy, and overall flow.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"Clean Markdown requires the original reading order, headings, sections, and hierarchy to be maintained correctly. (3 of 3 judges)","scenario":"scanned-research-paper","scenario_name":"Scanned Research Paper","group_tag":"scanned-research-paper","scenario_description":"An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.","modality":"pdf","input_text":null,"input_artifact_refs":[{"alt":null,"url":"https://d3epheqghktydj.cloudfront.net/convert-a-complex-pdf-into-clean-markdow-scanned-research-pdf-7b86de49784d.pdf","role":"input","filename":"Scanned Research PDF.pdf"}],"stresses":["OCR on scanned pages","Multi-column reading order","Figure and chart handling","Table reconstruction from scans","Caption association","Reference extraction","Overall document structure retention"],"verdict":"worked","score":null,"score_total":null,"note":"The tool preserves section hierarchy and column order in a dense multi-column scanned section, keeping headings aligned with the correct body text.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/9e300becd44b42ebbcd784ae4a79e52a.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/18b2277b09fa42e3ae7a325fba02b272.png?v=1","role":"input","alt":null}],"run_id":"6e3160de-fe46-4b45-b071-72560b5c5d0e","study_title":"Convert a Complex PDF into Clean Markdown with an API","study_kind":"generation","research_task":"86b9h7t37","tested_at":null,"completeness":"input-and-output","input":{"state":"files","text":null,"files":[{"url":"https://d3epheqghktydj.cloudfront.net/convert-a-complex-pdf-into-clean-markdow-scanned-research-pdf-7b86de49784d.pdf","filename":"Scanned Research PDF.pdf","alt":"Scanned Research Paper","role":"input"}],"modality":"pdf","stresses":["OCR on scanned pages","Multi-column reading order","Figure and chart handling","Table reconstruction from scans","Caption association","Reference extraction","Overall document structure retention"]},"tool_page_slug":"nutrient-io","tool_url":"https://aidemos.com/tools/nutrient-io","permalink":"https://aidemos.com/evidence/0310cd16-a31c-473f-bf94-1da61a08607c","api_url":"https://ai.aidemos.com/v1/observations/0310cd16-a31c-473f-bf94-1da61a08607c"},"peers":[{"id":"24968ab8-597c-4c69-9ec3-a0c216320ebc","tool":"adobe-api","tool_name":"Adobe API","verdict":"failed","score":null,"score_total":null,"note":"Dumps the scanned title page as a dense OCR block without section boundaries or other structural cues, so the document hierarchy is lost.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/28052bbc164e4f51be6903dcaccdb477.png?v=1","evidence_url":"https://aidemos.com/evidence/24968ab8-597c-4c69-9ec3-a0c216320ebc"},{"id":"4556b641-0861-4cec-9988-0748cdfa1664","tool":"extend-ai","tool_name":"Extend AI","verdict":"worked","score":null,"score_total":null,"note":"Reconstructs a multi-column research page so the 'STUDY AREA' heading, its paragraphs, and the following 'STAND PRESCRIPTIONS' section stay in order.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/18b2277b09fa42e3ae7a325fba02b272.png?v=1","evidence_url":"https://aidemos.com/evidence/4556b641-0861-4cec-9988-0748cdfa1664"},{"id":"4d96cef6-8e58-4798-b86c-1b626c8d3ae2","tool":"landing-ai","tool_name":"Landing AI","verdict":"worked","score":null,"score_total":null,"note":"Reconstructs a dense two-column section with its hierarchy intact, keeping the section heading attached to the correct body text across the columns.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/18b2277b09fa42e3ae7a325fba02b272.png?v=1","evidence_url":"https://aidemos.com/evidence/4d96cef6-8e58-4798-b86c-1b626c8d3ae2"},{"id":"fae001a3-d13a-4fa3-8415-862dfb85dca9","tool":"llamaparse","tool_name":"LlamaParse","verdict":"worked","score":null,"score_total":null,"note":"Reflows a two-column scanned page into a single coherent reading order while preserving section-to-body flow.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/bb646d6406274eb280aa628c32bfa474.png?v=1","evidence_url":"https://aidemos.com/evidence/fae001a3-d13a-4fa3-8415-862dfb85dca9"},{"id":"cc8096bb-19af-4b8f-901f-657634e78a25","tool":"mistral-ai","tool_name":"Mistral AI","verdict":"failed","score":null,"score_total":null,"note":"The opening page loses the distinction between the document title and the abstract, flattening the semantic organization of the first page.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/dd8e2c3021024a24be3970549c24ccc7.png?v=1","evidence_url":"https://aidemos.com/evidence/cc8096bb-19af-4b8f-901f-657634e78a25"},{"id":"7840c24e-489b-4551-b576-148911ebe640","tool":"reducto","tool_name":"Reducto","verdict":"struggled","score":null,"score_total":null,"note":"Displaces the byline by a full column: the author line that sits above the two-column split in the source is emitted only after the entire left column, and its footnote markers are rendered inconsistently.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/2979812e747e44b2b8b333d7b12c3618.png?v=1","evidence_url":"https://aidemos.com/evidence/7840c24e-489b-4551-b576-148911ebe640"},{"id":"fc3553de-a932-4313-bb83-66645ce7441c","tool":"tensorlake","tool_name":"Tensorlake","verdict":"worked","score":null,"score_total":null,"note":"Retains section-level reading order in a scanned multi-column paper, with headings continuing to guide the flow across columns and into the next section.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/18b2277b09fa42e3ae7a325fba02b272.png?v=1","evidence_url":"https://aidemos.com/evidence/fc3553de-a932-4313-bb83-66645ce7441c"},{"id":"86f4991b-db92-402b-bd03-7a1f3777691e","tool":"upstage-ai","tool_name":"Upstage AI","verdict":"failed","score":null,"score_total":null,"note":"Breaks paragraph-level segmentation in the scanned two-column page, so the extracted text no longer follows the source column order cleanly.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/18b2277b09fa42e3ae7a325fba02b272.png?v=1","evidence_url":"https://aidemos.com/evidence/86f4991b-db92-402b-bd03-7a1f3777691e"}],"other_criteria":[{"id":"b50f09d7-7cee-4268-8441-50b46ddeb000","criterion":"complex-document-handling","criterion_name":"Complex Document Handling","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Processes an image-only scanned research paper end to end and returns a parsed markdown output plus preview, showing it can handle a multi-page scanned document.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/b50f09d7-7cee-4268-8441-50b46ddeb000"},{"id":"97dfdbc2-3390-4bf7-9d20-af8a5dd33d01","criterion":"table-preservation","criterion_name":"Table Preservation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Largely preserves grouped-column tables, keeping their internal organization intact in the extracted output.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/97dfdbc2-3390-4bf7-9d20-af8a5dd33d01"},{"id":"d9350bf4-9574-4634-8293-5ced2ab60d9c","criterion":"table-preservation","criterion_name":"Table Preservation","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Breaks more complex multi-level tables, with rows and merged cells misaligned or lost once the hierarchy becomes dense.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/d9350bf4-9574-4634-8293-5ced2ab60d9c"},{"id":"1ff9ad63-6629-4bd0-a643-ff827f73c602","criterion":"text-ocr-completeness","criterion_name":"Text & OCR Completeness","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"Recovers readable text from scanned pages, including the abstract, keywords, author affiliations, title, and opening paragraphs on the first page.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/1ff9ad63-6629-4bd0-a643-ff827f73c602"},{"id":"474d7192-6964-4d73-9726-1ca2efcb6edc","criterion":"visual-content-retention","criterion_name":"Visual Content Retention","rank_role":"decisive","verdict":"struggled","score":null,"score_total":null,"note":"Extracts the Figure 3 chart's values, but the chart structure and layout are not preserved, so the visualization is flattened into text-like output.","artifact_count":2,"evidence_url":"https://aidemos.com/evidence/474d7192-6964-4d73-9726-1ca2efcb6edc"}],"appears_in":[{"page_type":"ranking","slug":"pdf-to-markdown-apis","title":"Best AI Tools to Convert Complex PDFs into Clean Markdown with an API","url":"https://aidemos.com/best/pdf-to-markdown-apis","binding":"run"}],"same_scenario":[{"id":"24968ab8-597c-4c69-9ec3-a0c216320ebc","tool":"adobe-api","tool_name":"Adobe API","verdict":"failed","score":null,"score_total":null,"note":"Dumps the scanned title page as a dense OCR block without section boundaries or other structural cues, so the document hierarchy is lost."},{"id":"4556b641-0861-4cec-9988-0748cdfa1664","tool":"extend-ai","tool_name":"Extend AI","verdict":"worked","score":null,"score_total":null,"note":"Reconstructs a multi-column research page so the 'STUDY AREA' heading, its paragraphs, and the following 'STAND PRESCRIPTIONS' section stay in order."},{"id":"4d96cef6-8e58-4798-b86c-1b626c8d3ae2","tool":"landing-ai","tool_name":"Landing AI","verdict":"worked","score":null,"score_total":null,"note":"Reconstructs a dense two-column section with its hierarchy intact, keeping the section heading attached to the correct body text across the columns."},{"id":"fae001a3-d13a-4fa3-8415-862dfb85dca9","tool":"llamaparse","tool_name":"LlamaParse","verdict":"worked","score":null,"score_total":null,"note":"Reflows a two-column scanned page into a single coherent reading order while preserving section-to-body flow."},{"id":"cc8096bb-19af-4b8f-901f-657634e78a25","tool":"mistral-ai","tool_name":"Mistral AI","verdict":"failed","score":null,"score_total":null,"note":"The opening page loses the distinction between the document title and the abstract, flattening the semantic organization of the first page."},{"id":"7840c24e-489b-4551-b576-148911ebe640","tool":"reducto","tool_name":"Reducto","verdict":"struggled","score":null,"score_total":null,"note":"Displaces the byline by a full column: the author line that sits above the two-column split in the source is emitted only after the entire left column, and its footnote markers are rendered inconsistently."},{"id":"fc3553de-a932-4313-bb83-66645ce7441c","tool":"tensorlake","tool_name":"Tensorlake","verdict":"worked","score":null,"score_total":null,"note":"Retains section-level reading order in a scanned multi-column paper, with headings continuing to guide the flow across columns and into the next section."},{"id":"86f4991b-db92-402b-bd03-7a1f3777691e","tool":"upstage-ai","tool_name":"Upstage AI","verdict":"failed","score":null,"score_total":null,"note":"Breaks paragraph-level segmentation in the scanned two-column page, so the extracted text no longer follows the source column order cleanly."}]}