CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)
Proof of Skill
Proof of Skill

Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)

Hyderabad, India
5+ years exp
Full-time
Posted 5d ago
1 views
Actively Hiring Urgent Opening Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Subject Matter Expert - Web Scraping & Crawling (Web Data Platform) role at Proof of Skill.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

About Chryselys

Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.

Role Summary

Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.

Responsibilities

  • Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
  • Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
  • Introduce proxy rotation and egress management; retire the single-IP failure mode.
  • Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
  • Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
  • Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
  • Containerize and schedule the pipeline; add CI running the offline tests on every change.
  • Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.

Skills

Web & protocol fundamentals

— HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]

Legacy stack (real mileage)

— urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]

Modern stack

— Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting [Must]

Reverse engineering

— Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]

Methodology breadth

— API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness [Must]

Non-HTML extraction

— PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]

Anti-bot & reliability

— Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries [Must]

Data engineering

— Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]

Testing & observability

— vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]

Build vs. buy

— Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]

Legal & ethical

— robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]

Experience — Required

  • 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
  • Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
  • Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
  • Mentored engineers; set crawl standards, review practice, and on-call runbooks.
  • Degree optional — equivalent practical experience is fully accepted.

Nice-to-Have

  • US payer policy, formulary, or prior-authorization document domain knowledge.
  • Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
  • LLM-assisted extraction at controlled cost — we run AWS Bedrock in verifier/.
  • Compliance or legal-review exposure on data acquisition programmes.

Equal Employment Opportunity

Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Key Requirements & Skills

  • 8+ years in data acquisition; 5+ years owning a scraping platform end to end
  • Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production
  • Rescued a brittle legacy scraper, with before/after reliability and cost numbers
  • Mentored engineers; set crawl standards, review practice, and on-call runbooks
  • Web & protocol fundamentals: HTTP/1.1-3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings
  • Legacy scraping stack mileage: urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, XPath/XSLT, Scrapy, Selenium 3, Splash/PhantomJS, SOAP/XML, ASP.NET __VIEWSTATE, sitemap.xml, WARC
  • Modern stack: Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata
  • Reverse engineering private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenges
  • Methodology breadth: API-first vs browser rendering, SPA payloads (NEXT_DATA, Nuxt), BFS/DFS frontier management, URL canonicalization & dedup, incremental/CDC crawling, distributed queues (Redis/
  • Non-HTML extraction: PDF (pdfplumber, PyMuPDF), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction cost & failure modes
  • Anti-bot & reliability: Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; fingerprinting; proxy rotation; CAPTCHAs; honeypots; backoff with jitter, circuit breakers, idempotent retries
  • Data engineering: Python expert (async, typing, profiling), SQL, pydantic validation, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD
  • Testing & observability: vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics
  • Legal & ethical: robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing with Legal partnership
  • Build vs. buy evaluation of Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk
  • US payer policy, formulary, or prior-authorization document domain knowledge
  • Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library
  • LLM-assisted extraction at controlled cost (AWS Bedrock experience)
  • Compliance or legal-review exposure on data acquisition programmes
  • Node/TypeScript useful

Benefits & Perks

vision-assisted extraction and its cost & failure modes [Must]

Frequently Asked Questions

How to apply for Subject Matter Expert - Web Scraping & Crawling (Web Data Platform) at Proof of Skill?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

5+ years of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
Proof of Skill

Proof of Skill

Hire talent faster with verified skill assessments. Streamline your recruitment by validating candidates' abilities.

Visit Company Website

More jobs at Proof of Skill

Mobile App Developer

Bangalore North, India

Social Media Executive

Mumbai, India

UX/UI Designer (Fresher)

Mumbai, India

Share this Opening

Job Alerts for data_engineering

Receive email alerts whenever new data_engineering roles in Hyderabad are posted.

Set Free Alert →

Similar Openings

Explore related active roles in data_engineering

View all
UrgentActively Hiring
Momentum Financial Services Group
Lead Data Engineer
Momentum Financial Services Group Verified
10+ years
Salary not disclosed
Hyderabad (Remote)
data_engineeringFull-timeRemote
Posted 18h ago
Apply Now
UrgentActively Hiring
aecom2
Data Engineer
aecom2 Verified
0-2 Yrs
₹111/mo
Bristol, 3 RIVERGATE, gb
data_engineeringFull-time
Posted 18h ago
Apply Now
UrgentActively Hiring
Magnals
Lead Data Engineer
Magnals Verified
5+ years
₹10.7L – ₹12.1L/mo
Remote
data_engineeringFull-timeRemote
Posted 18h ago
Apply Now

Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)

Proof of Skill · Hyderabad

Apply on Company Website