{"observation":{"id":"f027509e-7bee-4fa6-8580-a72e8c2950e4","tool":"spider","tool_name":"Spider","criterion":"noise-filtering","criterion_name":"Noise Filtering","criterion_definition":"How well the tool strips boilerplate, ads, navigation, comments, and other clutter from a static page.","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"Removing ads, nav, comments, and other clutter is fundamental to producing clean Markdown or structured data. (3 of 3 judges)","scenario":"glassdoor-software-engineer-jobs-behind-sign-in-modal","scenario_name":"Glassdoor software engineer jobs behind sign-in modal","group_tag":"web-scraping-benchmark","scenario_description":"A Glassdoor jobs listing page protected by Cloudflare and a sign-in/interstitial overlay, used to test proxy evasion, anti-bot handling, and the ability to dismiss blocking modal UI before extracting listings.","modality":"mixed","input_text":"https://www.glassdoor.com/Job/software-engineer-jobs-SRCH_KO0,17.htm — Dismiss any immediate sign-in or signup modal overlays that block the view. Once cleared, extract the top 5 job listings, including job title, company name, location, and the short summary snippet.","input_artifact_refs":[],"stresses":[],"verdict":"failed","score":null,"score_total":null,"note":"The scraper fails to strip static boilerplate from cluttered pages: the markdown included the global header navigation, social-sharing URLs, cookie-choice notices, and user reviews instead of isolating only the core recipe content.","evidence_state":"observed","source":"first-party","artifacts":[],"run_id":"06e1dbd6-5518-4af8-aa1a-735259a75b4f","study_title":"Scrape Web Pages Into Clean Markdown or Structured Data Using AI","study_kind":"backfill","research_task":"86b9jm3a3","tested_at":"2026-06-23T04:32:57.304000+00:00","completeness":"input-only","input":{"state":"text","text":"https://www.glassdoor.com/Job/software-engineer-jobs-SRCH_KO0,17.htm — Dismiss any immediate sign-in or signup modal overlays that block the view. Once cleared, extract the top 5 job listings, including job title, company name, location, and the short summary snippet.","files":[],"modality":"mixed","stresses":[]},"tool_page_slug":"spider","tool_url":"https://aidemos.com/tools/spider","permalink":"https://aidemos.com/evidence/f027509e-7bee-4fa6-8580-a72e8c2950e4","api_url":"https://ai.aidemos.com/v1/observations/f027509e-7bee-4fa6-8580-a72e8c2950e4"},"peers":[{"id":"ae382ded-80c3-4f21-9916-62043515403e","tool":"firecrawl","tool_name":"Firecrawl","verdict":"failed","score":null,"score_total":null,"note":"Leaves page scaffolding in the extraction stream, including skip links and global navigation, instead of cleaning the listing output down to the core jobs content.","artifact_count":1,"thumbnail":"https://d3epheqghktydj.cloudfront.net/firecrawl-firecrawl-glassdoor-scrape-markdown-output.png","evidence_url":"https://aidemos.com/evidence/ae382ded-80c3-4f21-9916-62043515403e"},{"id":"1c2ec4aa-ed8e-4021-8ef3-56d35ce08689","tool":"jina-ai-reader","tool_name":"Jina AI Reader","verdict":"mixed","score":null,"score_total":null,"note":"It can recover the page text layer, but the extraction still leaves job data interleaved with French and German translation strings and header redirect text, so heavy post-processing cleanup is still required.","artifact_count":0,"thumbnail":null,"evidence_url":"https://aidemos.com/evidence/1c2ec4aa-ed8e-4021-8ef3-56d35ce08689"}],"other_criteria":[{"id":"ec274907-80f4-4b17-8978-4513d4584e38","criterion":"output-quality","criterion_name":"Output Quality","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The extractor can cleanly capture static structural text and basic marketing attributes, but it misses vital dynamic transactional nodes on hydrated pages.","artifact_count":0,"evidence_url":"https://aidemos.com/evidence/ec274907-80f4-4b17-8978-4513d4584e38"},{"id":"ec071df3-73be-44eb-b169-0f8eb530baee","criterion":"output-quality","criterion_name":"Output Quality","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Returned zero job payload; the output text consisted entirely of multilingual CAPTCHA strings and security warnings.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/ec071df3-73be-44eb-b169-0f8eb530baee"},{"id":"2dddc43a-aa71-4fd9-89a6-c195bdbcd6e6","criterion":"output-quality","criterion_name":"Output Quality","rank_role":"decisive","verdict":"mixed","score":null,"score_total":null,"note":"The extractor preserves the main recipe content accurately, including the ingredients block and directions layout, but the returned markdown is highly unrefined and bloated with boilerplate text.","artifact_count":0,"evidence_url":"https://aidemos.com/evidence/2dddc43a-aa71-4fd9-89a6-c195bdbcd6e6"},{"id":"509adb45-829b-45f6-b216-f08ae6ac7454","criterion":"proxy-evasion","criterion_name":"Proxy Evasion","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Native proxy handling fails against anti-bot protection, triggering a full 'Humans only' Cloudflare-style block page instead of the target listings.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/509adb45-829b-45f6-b216-f08ae6ac7454"},{"id":"e69ce784-8987-4af8-bba7-1a1e1fe4ee59","criterion":"proxy-evasion","criterion_name":"Proxy Evasion","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"It was stopped by the site's security interstitial and returned only anti-bot warning text instead of the target listings, showing no recovered job content.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/e69ce784-8987-4af8-bba7-1a1e1fe4ee59"},{"id":"5d31e225-3519-47b9-97f4-420a31f39baa","criterion":"schema-extraction-integrity","criterion_name":"Schema Extraction Integrity","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It can preserve the main recipe content accurately, keeping the central ingredients block and recipe directions layout intact.","artifact_count":0,"evidence_url":"https://aidemos.com/evidence/5d31e225-3519-47b9-97f4-420a31f39baa"},{"id":"95ec5f9f-da1c-43e4-a0c4-32537ef704f0","criterion":"visual-spatial-awareness","criterion_name":"Visual Spatial Awareness","rank_role":"decisive","verdict":"failed","score":null,"score_total":null,"note":"Its structural-cleanup is weak on cluttered static pages: it leaves global navigation links, social-sharing URLs, cookie notices, and user reviews in the markdown instead of isolating the core page content.","artifact_count":0,"evidence_url":"https://aidemos.com/evidence/95ec5f9f-da1c-43e4-a0c4-32537ef704f0"}],"appears_in":[{"page_type":"ranking","slug":"ai-web-scraping-tools","title":"Best AI Tools to Scrape Web Pages Into Clean Markdown or Structured Data","url":"https://aidemos.com/best/ai-web-scraping-tools","binding":"run"}],"same_scenario":[{"id":"ae382ded-80c3-4f21-9916-62043515403e","tool":"firecrawl","tool_name":"Firecrawl","verdict":"failed","score":null,"score_total":null,"note":"Leaves page scaffolding in the extraction stream, including skip links and global navigation, instead of cleaning the listing output down to the core jobs content."},{"id":"1c2ec4aa-ed8e-4021-8ef3-56d35ce08689","tool":"jina-ai-reader","tool_name":"Jina AI Reader","verdict":"mixed","score":null,"score_total":null,"note":"It can recover the page text layer, but the extraction still leaves job data interleaved with French and German translation strings and header redirect text, so heavy post-processing cleanup is still required."}]}