{bc}
zoho_recruit

Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)

Chryselys
Hyderabad, IND
Full-time
Senior · 8+ years experience
Onsite
Discovered 6 days ago
AWS BedrockHTTP/1.1HTTP/2HTTP/3TLSJA3/JA4
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

AWS BedrockHTTP/1.1HTTP/2
Smart Apply

Full Job Posting

About Chryselys

Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.

Role Summary

Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.

Responsibilities

  • Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
  • Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
  • Introduce proxy rotation and egress management; retire the single-IP failure mode.
  • Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
  • Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
  • Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
  • Containerize and schedule the pipeline; add CI running the offline tests on every change.
  • Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.

Skills

  • Web & protocol fundamentals — HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]
  • Legacy stack (real mileage) — urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]
  • Modern stack — Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting [Must]
  • Reverse engineering — Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]
  • Methodology breadth — API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness [Must]
  • Non-HTML extraction — PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]
  • Anti-bot & reliability — Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries [Must]
  • Data engineering — Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]
  • Testing & observability — vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]
  • Build vs. buy — Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]
  • Legal & ethical — robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]

Experience — Required

8+ years in data acquisition; 5+ years owning a scraping platform end to end.

Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.

Rescued a brittle legacy scraper, with before/after reliability and cost numbers.

Mentored engineers; set crawl standards, review practice, and on-call runbooks.

Degree optional — equivalent practical experience is fully accepted.

Nice-to-Have

US payer policy, formulary, or prior-authorization document domain knowledge.

Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.

LLM-assisted extraction at controlled cost — we run AWS Bedrock in verifier/ .

Compliance or legal-review exposure on data acquisition programmes.

Equal Employment Opportunity

Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Chryselys