This is a hands-on engineering role focused on building and operating large-scale web crawlers and data acquisition systems that power dataset creation for frontier AI model training. You will work as an extension of AI research teams, helping to collect, clean, and curate web data at a scale that few organizations can match. The work directly shapes the pretraining and inference pipelines used by some of the most advanced AI labs in the world.
What You'll Do
Build and maintain large-scale web crawlers across diverse domains including social media, travel, and multi-language sites.
Design high-throughput, fault-tolerant data collection systems capable of handling millions to billions of URLs per day.
Navigate anti-bot systems, rate limits, and JavaScript-heavy sites, finding creative solutions when standard protocols fall short.
Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB to PB scale.
Construct and maintain datasets for research and model training in close collaboration with research teams.
Monitor crawl performance, coverage, and data quality, iterating quickly as web environments change.
Optimize infrastructure for cost, latency, and reliability across cloud and bare-metal environments.
What We're Looking For
3 or more years building and operating web crawlers at scale (2 or more years considered for candidates with a PhD).
Proficiency in one or more of: Go, Rust, Python, Java, or C++.
Experience running data pipelines at TB or greater scale.
Deep knowledge of HTTP, networking, and browser behavior.
Hands-on experience with distributed systems or parallel processing.
Experience with headless browsers such as Playwright, Puppeteer, or Chrome DevTools Protocol.
Familiarity with proxy systems, IP rotation, or request orchestration.
Experience with data quality evaluation, scoring, or benchmarking at scale.
Experience running crawling or data workloads on cloud platforms (AWS, GCP) or bare-metal infrastructure.
Background in NLP pipelines, ML dataset curation, or AI lab work is a strong plus.
Compensation & Benefits
Salary range: $160,000 to $250,000 USD annually. Visa sponsorship is not available for this role.
Location
This role is fully remote. The primary location is Los Angeles, CA, United States, though candidates based in other major US cities are welcome.
Numbers & Facts
Location
Los Angeles, California (Remote)
Salary
$160,000–$250,000 Per Year
Skills
Amazon Web Services (AWS)unmatched
Artificial Intelligence (AI)unmatched
Benchmarkingunmatched
C++ Programming Languageunmatched
Cloud Computingunmatched
Cost Controlunmatched
Data Analysisunmatched
Data Collectionunmatched
Data Managementunmatched
Data Qualityunmatched
Distributed Computingunmatched
GCP (Good Clinical Practices)unmatched
Go Programming Language (Golang)unmatched
HTTP (HyperText Transport Protocol)unmatched
High Throughputunmatched
IP (Internet Protocol)unmatched
Javaunmatched
JavaScriptunmatched
Laboratoryunmatched
Natural Language Processing (NLP)unmatched
Parallel Computingunmatched
Performance Analysisunmatched
Problem Solving Skillsunmatched
Python Programming/Scripting Languageunmatched
Rust Programming Languageunmatched
Social Mediaunmatched
Traffic Shapingunmatched
Training Data Setsunmatched
Web Browsersunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.