← All jobs
MTS - Pre-Training Data & Acquisition Engineer
San Francisco · Remote · FullTime · R&D
Apply well, not just fast
Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.
About the role
PythonGoRustC++PrometheusDistributed Systems
THE ROLE
We’re hiring a Pre-Training Data & Acquisition Engineer to build the data systems powering Prometheus’s foundation models for the physical world. You’ll work closely with research and infrastructure teams to acquire, process, and deliver large-scale training datasets across engineering, scientific, and multimodal domains. This role spans distributed crawling, source integration, data processing, and production operation, with end-to-end ownership from raw content to training-ready datasets.
WHAT YOU’LL DO
- Identify and integrate valuable data sources across engineering and scientific domains.
- Build distributed crawlers, API integrations, and bulk ingestion systems with effective scheduling, rate limiting, retries, and incremental updates.
- Develop pipelines for parsing, extraction, normalization, deduplication, quality filtering, and tokenization across heterogeneous formats.
- Optimize throughput and cost across networking, compute, storage, and databases as acquisition and processing workloads scale.
- Build monitoring and tooling to track source coverage, ingestion failures, processing throughput, and usable data yield.
- Work closely with pre-training researchers to translate data requirements into reliable pipelines and deliver datasets ready for large-scale training.
- Own dataset reproducibility, versioning, provenance, and recovery from acquisition through delivery.
WHAT WE’RE LOOKING FOR
- Experience building and operating large-scale distributed systems, web crawlers, or data processing pipelines.
- Strong programming ability in Python and Rust, Go, C++, or a comparable systems language.
- A practical understanding of web infrastructure, including HTTP, DNS, concurrency, caching, and common failure modes.
- Hands-on experience with databases, object storage, and distributed batch or streaming processing.
- Experience designing fault-tolerant systems that handle partial failures, resume interrupted work, and prevent unintended duplication or data loss.
- Ability to profile and debug performance across CPU, memory, disk, and network usage.
- Strong technical judgment when integrating unfamiliar sources, APIs, and file formats.
- Bias toward fast iteration and end-to-end ownership, from initial implementation through reliable production operation.
- Experience with search indexing, document extraction, or foundation-model data pipelines is a plus.
WHY JOIN US
- Work with world-class researchers on frontier AI systems for the physical world.
- Build the acquisition and processing systems that supply engineering, scientific, and multimodal data to large-scale model training.
- Competitive compensation and flexible work arrangements.
- High-impact, mission-driven environment.