{"observation":{"id":"2dddc43a-aa71-4fd9-89a6-c195bdbcd6e6","tool":"spider","tool_name":"Spider","criterion":"output-quality","criterion_name":"Output Quality","criterion_definition":"Conceptually matches the prompt and is visually usable.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"Whether the result matches the page and is actually usable is the main outcome the ranking is trying to measure. (3 of 3 judges)","scenario":"glassdoor-software-engineer-jobs-behind-sign-in-modal","scenario_name":"Glassdoor software engineer jobs behind sign-in modal","group_tag":"web-scraping-benchmark","scenario_description":"A Glassdoor jobs listing page protected by Cloudflare and a sign-in/interstitial overlay, used to test proxy evasion, anti-bot handling, and the ability to dismiss blocking modal UI before extracting listings.","modality":"mixed","input_text":"https://www.glassdoor.com/Job/software-engineer-jobs-SRCH_KO0,17.htm — Dismiss any immediate sign-in or signup modal overlays that block the view. Once cleared, extract the top 5 job listings, including job title, company name, location, and the short summary snippet.","input_artifact_refs":[],"stresses":[],"verdict":"mixed","score":null,"score_total":null,"note":"The extractor preserves the main recipe content accurately, including the ingredients block and directions layout, but the returned markdown is highly unrefined and bloated with boilerplate text.","evidence_state":"observed","source":"first-party","artifacts":[],"run_id":"06e1dbd6-5518-4af8-aa1a-735259a75b4f","study_title":"Scrape Web Pages Into Clean Markdown or Structured Data Using AI","study_kind":"backfill","research_task":"86b9jm3a3","tested_at":"2026-06-23T04:32:57.304000+00:00","completeness":"input-only","input":{"state":"text","text":"https://www.glassdoor.com/Job/software-engineer-jobs-SRCH_KO0,17.htm — Dismiss any immediate sign-in or signup modal overlays that block the view. Once cleared, extract the top 5 job listings, including job title, company name, location, and the short summary snippet.","files":[],"modality":"mixed","stresses":[]},"tool_page_slug":"spider","tool_url":"https://aidemos.com/tools/spider","permalink":"https://aidemos.com/evidence/2dddc43a-aa71-4fd9-89a6-c195bdbcd6e6","api_url":"https://ai.aidemos.com/v1/observations/2dddc43a-aa71-4fd9-89a6-c195bdbcd6e6"},"peers":[{"id":"9ed0263e-7874-4c44-a0d2-c55dd52cca1c","tool":"firecrawl","tool_name":"Firecrawl","verdict":"mixed","score":null,"score_total":null,"note":"Extracts the core job-listing content, but breaks the text structure with navigation buttons, search filter blocks, and internal page links.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/firecrawl-firecrawl-glassdoor-scrape-markdown-output.png","evidence_url":"https://aidemos.com/evidence/9ed0263e-7874-4c44-a0d2-c55dd52cca1c"},{"id":"fc132656-20f1-48c6-9daa-c78d98a6dc3e","tool":"jina-ai-reader","tool_name":"Jina AI Reader","verdict":"mixed","score":null,"score_total":null,"note":"Its output quality is partial on the hydrated ecommerce page: it preserves the SEO header and static price markers, but replaces the transactional product data with a giant global link directory.","artifact_count":0,"thumbnail":null,"evidence_url":"https://aidemos.com/evidence/fc132656-20f1-48c6-9daa-c78d98a6dc3e"},{"id":"6883c340-cb0d-4463-abd0-44546868d454","tool":"skyvern","tool_name":"Skyvern","verdict":"worked","score":null,"score_total":null,"note":"Produces a clean markdown_content payload with readable headings and bold labels for the extracted job listings.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/skyvern-skyvern-job-listings-output-markdown.png","evidence_url":"https://aidemos.com/evidence/6883c340-cb0d-4463-abd0-44546868d454"}],"other_criteria":[{"id":"f027509e-7bee-4fa6-8580-a72e8c2950e4","criterion":"noise-filtering","criterion_name":"Noise Filtering","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"The scraper fails to strip static boilerplate from cluttered pages: the markdown included the global header navigation, social-sharing URLs, cookie-choice notices, and user reviews instead of isolating only the core recipe content.","artifact_count":0,"evidence_url":"https://aidemos.com/evidence/f027509e-7bee-4fa6-8580-a72e8c2950e4"},{"id":"e69ce784-8987-4af8-bba7-1a1e1fe4ee59","criterion":"proxy-evasion","criterion_name":"Proxy Evasion","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"It was stopped by the site's security interstitial and returned only anti-bot warning text instead of the target listings, showing no recovered job content.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/e69ce784-8987-4af8-bba7-1a1e1fe4ee59"},{"id":"509adb45-829b-45f6-b216-f08ae6ac7454","criterion":"proxy-evasion","criterion_name":"Proxy Evasion","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Native proxy handling fails against anti-bot protection, triggering a full 'Humans only' Cloudflare-style block page instead of the target listings.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/509adb45-829b-45f6-b216-f08ae6ac7454"},{"id":"5d31e225-3519-47b9-97f4-420a31f39baa","criterion":"schema-extraction-integrity","criterion_name":"Schema Extraction Integrity","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It can preserve the main recipe content accurately, keeping the central ingredients block and recipe directions layout intact.","artifact_count":0,"evidence_url":"https://aidemos.com/evidence/5d31e225-3519-47b9-97f4-420a31f39baa"},{"id":"95ec5f9f-da1c-43e4-a0c4-32537ef704f0","criterion":"visual-spatial-awareness","criterion_name":"Visual Spatial Awareness","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Its structural-cleanup is weak on cluttered static pages: it leaves global navigation links, social-sharing URLs, cookie notices, and user reviews in the markdown instead of isolating the core page content.","artifact_count":0,"evidence_url":"https://aidemos.com/evidence/95ec5f9f-da1c-43e4-a0c4-32537ef704f0"}],"appears_in":[{"page_type":"ranking","slug":"ai-web-scraping-tools","title":"Best AI Tools to Scrape Web Pages Into Clean Markdown or Structured Data","url":"https://aidemos.com/best/ai-web-scraping-tools","binding":"run"}],"same_scenario":[{"id":"9ed0263e-7874-4c44-a0d2-c55dd52cca1c","tool":"firecrawl","tool_name":"Firecrawl","verdict":"mixed","score":null,"score_total":null,"note":"Extracts the core job-listing content, but breaks the text structure with navigation buttons, search filter blocks, and internal page links."},{"id":"fc132656-20f1-48c6-9daa-c78d98a6dc3e","tool":"jina-ai-reader","tool_name":"Jina AI Reader","verdict":"mixed","score":null,"score_total":null,"note":"Its output quality is partial on the hydrated ecommerce page: it preserves the SEO header and static price markers, but replaces the transactional product data with a giant global link directory."},{"id":"6883c340-cb0d-4463-abd0-44546868d454","tool":"skyvern","tool_name":"Skyvern","verdict":"worked","score":null,"score_total":null,"note":"Produces a clean markdown_content payload with readable headings and bold labels for the extracted job listings."}]}