İş RadarıAll jobs
Active On-site San Francisco, California, United States Posted · 24.07.2026 Ashby (US)

Research Engineer, Real Environments

Mercor

ABOUT MERCOR Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.   Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. ABOUT THE ROLE You’ll work with large enterprises to capture their data and transform it into high-fidelity RL environments for capability evaluations and training datasets for frontier labs. We focus on pushing the frontier of world-building, verifier engineering, and more alongside our partners. Your goal will be to automate the process of building evals for real work in the economy. WHAT YOU'LL DO • Ship models for workflow extraction, classification, and grading. • Engineer autonomous task refinement processes which distill data taste into pipelines. • Deliver data to customers and deploy into real engagements. • Help define the future of agentic transformation for enterprises around the world. • Deeply learn about the intricacies of enterprises through building evaluations for all aspects of work. • Build end-to-end environments for labs & enterprises by platformizing sandbox app clones, load real data into the sandboxes, build prompts from real workflows, and write verifiers leveraging enterprise expertise & golden outputs. • Systematize the production of environments to scale throughput while maintaining high-quality worlds and verifiers. WHAT WE'RE LOOKING FOR • Prior experience shipping environments – you’ve contributed to an OSS framework, built environments at previous companies, or worked on agentic evaluations. • Strong full-stack engineering skills – you’ll be responsible for everything from infrastructure to app code to analytics • Bias to action – this team is focused on shipping evals, not just philosophizing about them. • Curiosity – being biased towards understanding and digging deep into model behavior and actually looking at the data. • Sweat the details that make a simulation indistinguishable from the real thing and have systems-level thinking skills that allow you to scale up quality. NICE TO HAVE • Experience with Temporal, Modal, or similar orchestration/compute services • Experience with synthetic data generation for frontier models.Past work auditing and scrutinizing industry-standard evaluations BENEFITS • Semi-annual performance bonus structure • Generous equity grant vested over 4 years • Up to $15k Relocation bonus - $10K housing bonus (if you live within 0.5 miles of our office) - $1.5K monthly stipend for meals • Free Equinox membership - $200 monthly laundry reimbursement - $200 monthly personal wellness reimbursement • Health, Dental, Vision insurance
This job was verified from Ashby (US). Applications are completed on the original source.
Apply on the original listing ↗
Something wrong with this job?