Pinwheel API is a leading data infrastructure company enabling seamless access to employment and income verification data through secure APIs. The company powers payroll connectivity for fintechs, lenders, HR platforms, and financial services across the US market.
Role Overview
As an AI Web Scraping Engineer (Stealth Data Acquisition), you will be responsible for designing, building, and maintaining resilient, undetectable data collection systems at scale. This is not a "requests + BeautifulSoup" role — we need someone who has operated against modern anti-bot infrastructure (Cloudflare, PerimeterX, DataDome, Akamai, Kasada) and knows how to stay under the radar while extracting structured data reliably and continuously. You will integrate AI/LLM-based parsing to handle dynamic, unstructured, or frequently-changing page layouts and build self-healing scrapers that detect layout/DOM changes and adapt automatically.
Responsibilities
Stealth Scraping Pipeline Development – Design, build, and maintain stealth-based scraping pipelines that evade bot detection, fingerprinting, and rate-limiting defenses.
Distributed Infrastructure Architecture – Architect distributed scraping infrastructure using rotating residential/mobile proxies, headless browser farms, and session/cookie management at scale.
AI-Powered Parsing Integration – Integrate AI/LLM-based parsing to handle dynamic, unstructured, or frequently-changing page layouts without brittle selectors.
Self-Healing Scraper Development – Build self-healing scrapers that detect layout/DOM changes and adapt automatically.
Monitoring & Iteration – Monitor scraper health, success rates, and detection signals; rapidly iterate when sites patch defenses.
Compliance & Legal Guardrails – Ensure data pipelines are compliant with client-defined legal and ethical guardrails (ToS review, robots.txt handling, rate limiting) as directed by Pinwheel's legal/compliance team.
Data Pipeline Collaboration – Collaborate with data engineering to deliver clean, structured, deduplicated datasets into downstream systems.
Requirements
4+ years building production web scraping systems at scale (not just scripts — real infrastructure).
Deep hands-on experience with headless browser automation (Playwright, Puppeteer, Selenium) including stealth plugins/patches.
Practical experience defeating or working around Cloudflare, DataDome, PerimeterX, Akamai Bot Manager, Kasada, or similar.
Strong understanding of TLS fingerprinting, HTTP/2 fingerprinting, canvas/WebGL fingerprinting, and browser fingerprint spoofing.
Experience with residential/mobile/datacenter proxy rotation and session management at scale.
Proficiency in Python and/or Node.js/TypeScript for scraping and orchestration.
Experience using LLMs/AI models for content extraction, parsing, and classification on messy/unstructured HTML.
Familiarity with CAPTCHA-solving approaches and detection-avoidance strategies.
Must be based in and authorized to work in the United States.
Must pass a standard background check prior to start.
Comfortable working in a fast-moving, iterative environment where target sites change defenses frequently.
Nice to Have
Experience with mobile app API reverse engineering (in addition to web).
Background in adversarial ML or anti-bot research (from either side).
Experience with distributed task queues (Celery, Kafka, SQS) for scraping orchestration.
Prior work in a data-as-a-service, alt-data, or web intelligence company.
Benefits
Compensation in USD
Fully remote work
Career growth on an international company
Numbers & Facts
Location
Denver, CO (Remote)
Skills
Apache Kafkaunmatched
Application Programming Interface (API)unmatched
Artificial Intelligence (AI)unmatched
Automationunmatched
Canvasunmatched
DOM (Document Object Model)unmatched
Data Collectionunmatched
Data Managementunmatched
Data Setsunmatched
Financial Servicesunmatched
HTML (HyperText Markup Language)unmatched
HTTP (HyperText Transport Protocol)unmatched
Legalunmatched
Loansunmatched
Mobile Applications Developmentunmatched
Network Operations Centerunmatched
Node.jsunmatched
Production Systemsunmatched
Python Programming/Scripting Languageunmatched
Reverse Engineeringunmatched
SSL-TLS (Secure Socket Layer - Transport Layer Security)unmatched
Scripting (Scripting Languages)unmatched
Seleniumunmatched
Simple Queue Service (SQS)unmatched
Software Patchesunmatched
Structured Dataunmatched
Traffic Shapingunmatched
Web Browsersunmatched
Web Client Plug-insunmatched
Web Productionunmatched
Web Programmingunmatched
Work From Homeunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.