Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Key Skills for This Role
Full Job Posting
About Chryselys
Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.
Role Summary
Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Responsibilities
- Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
- Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
- Introduce proxy rotation and egress management; retire the single-IP failure mode.
- Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
- Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
- Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
- Containerize and schedule the pipeline; add CI running the offline tests on every change.
- Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Skills
- Web & protocol fundamentals — HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]
- Legacy stack (real mileage) — urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]
- Modern stack — Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting [Must]
- Reverse engineering — Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]
- Methodology breadth — API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness [Must]
- Non-HTML extraction — PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]
- Anti-bot & reliability — Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries [Must]
- Data engineering — Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]
- Testing & observability — vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]
- Build vs. buy — Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]
- Legal & ethical — robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]
Experience — Required
8+ years in data acquisition; 5+ years owning a scraping platform end to end.
Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
Mentored engineers; set crawl standards, review practice, and on-call runbooks.
Degree optional — equivalent practical experience is fully accepted.
Nice-to-Have
US payer policy, formulary, or prior-authorization document domain knowledge.
Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
LLM-assisted extraction at controlled cost — we run AWS Bedrock in verifier/ .
Compliance or legal-review exposure on data acquisition programmes.
Equal Employment Opportunity
Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
About Chryselys
Verified company details for this employer are not available yet.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Chryselys
Senior Generative AI Engineer - Agentic Infrastructure
Hyderabad, IND
Generative AI Agent Engineer - Agent Design & Performance
Hyderabad, IND
Lead Application Architect / Full-Stack Application Engineer
Hyderabad, IND
Senior Generative AI Engineer - Agentic Infrastructure
Hyderabad, IND
Senior Generative AI Engineer - Front End
Hyderabad, IND
Generative AI Agent Engineer - Agent Design & Performance
Hyderabad, IND
Lead Application Architect / Full-Stack Application Engineer
Hyderabad, IND
Graph RAG - Knowledge Engineer
Hyderabad, IND
Senior Associate - Digital Analytics
Hyderabad, IND
Senior Associate - Market Access
Hyderabad, IND
Senior Associate - TA Analytics
Hyderabad, IND