← All jobs
Research Scientist, Benchmarks & Evals
United States · Remote · FullTime · Engineering
$150K – $250K
Apply well, not just fast
Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.
About the role
PyTorchPythonPostgreSQLMachine Learning
ABOUT KLED
Kled is building the largest opt-in human data network in the world.
We are not a labeling firm. We are not a task marketplace.
We are a consumer application where people upload their real photos, videos, and documents and get paid continuously.
We then filter, standardize, and license that data to frontier AI labs and enterprises that need fresh, rights-aware training data.
Since launching our mobile app in 2026, we have:
• Reached #1 on the App Store (Finance) with 0 paid marketing
• Scaled to 600,000+ active data contributors
• Processed 5M uploads per day
• Raised $15M+ from investors behind SpaceX, Airbnb, Coinbase, xAI, OpenAI, Anthropic, Spotify, Lyft, Uber, and more
Our mission is to let anyone download the app and earn a real living wage from uploading their data.
ABOUT THE ROLE
Research Scientist - Benchmarks & Evals
AI labs buy data when a benchmark shows it works.
You’ll build the benchmarks that show what one of the largest human data networks in the world can do for frontier AI models.
You will:
- Design and ship public benchmarks for image, video and document models
- Build private test sets from real, opt-in data that no model has trained on
- Run data-value experiments: train models with and without Kled data, measure what changes
- Design human evaluation studies and the automatic metrics that track them
- Publish leaderboards and reports that researchers at frontier labs trust
- Find where models fall short and tell our team what data to collect next
- Work directly with the founders and lab buyers to turn results into data deals
WE’RE LOOKING FOR
- 3+ years ML research or research engineering experience
- A benchmark or evaluation you built that other people actually used
- Strong hands-on experience training and fine-tuning image or video models (PyTorch)
- Deep understanding of experimental design and statistics (controls, ablations, human studies)
- Experience running experiments end to end, from GPUs to final report
- Clear writing for both researchers and buyers
Bonus:
- Publications at NeurIPS, ICML, ICLR or CVPR (datasets and benchmarks especially)
- Experience on an evaluation or data team at an AI lab
- Experience selling or licensing data to AI labs
- Experience with data valuation or data attribution research
CURRENT STACK
Research
- Python / PyTorch
- Open-weight image and video models
- Cloud GPUs
Data
- Hundreds of millions of opt-in photos, videos and documents
- PostgreSQL (Supabase)
- S3 storage
COMPENSATION
- Base salary: $150,000 - $250,000
- $350,000 – $750,000 equity
We move fast and work hard.
If you're excited to build the benchmarks that decide what data trains frontier AI, let’s talk!
GROWTH OPPORTUNITY
You’ll join a team operating at the frontier of applied AI data infrastructure.
In this role, you’ll have the opportunity to:
• Own core systems that power one of the largest human data networks in the world
• Design infrastructure that directly influences what data trains next-generation AI models
• Build at real scale - millions of uploads per day, adversarial environments, global contributors
• Ship alongside a team that has built marketplaces, AI systems, and products used by millions
If you’re excited to move fast, build systems that matter, and help define how human data powers frontier AI, let’s talk.