Member of Technical Staff, Deeptune Environments
TLDR
Build sandboxed AI environments with orchestration, tool APIs, verifiers, and pipelines that turn human demonstrations into reproducible training systems.
About Mercor
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
About Deeptune
Deeptune is an RL environments lab within Mercor. We build training gyms for AI agents — high-fidelity simulations where models learn to solve economically valuable problems through reinforcement learning. We work with the leading AI labs to train the next generation of agentic models, and our environments have already contributed to recent breakthroughs in computer use, code generation, and multi-step task completion.
About the Role
This is a backend and infrastructure role, not a research or model-training role.
You'll build the systems our environments run on: sandboxed execution, orchestration for thousands of concurrent rollouts, agent and task harnesses, tool APIs, the grading layer that decides whether an agent actually succeeded, and the pipelines that turn raw human demonstrations into training-ready environments.
You'll do this alongside researchers who own the RL side. Your job is to make their ideas real: fast, correct, and reliable at scale. That means you should know enough about post-training, evals, and reward modeling to push back on a research spec productively. It does not mean you'll be training models. If you want to do model training or research engineering, this isn't the right scope, take a look at our other open roles instead.
The work is high-ownership and lightly specified. You'll set direction, drive outcomes, and stay hands-on. You'll also work directly with AI labs and enterprise partners, which means shipping against real external deadlines rather than internal ones.
What You'll Do
Build environments end to end: the simulated app or system, the agent-facing tool surface, the task definitions, and the verifiers that score them.
Design and operate the infrastructure that runs environments at scale: containers, sandboxing, orchestration, queues, observability
Turn messy human demonstration data into clean, reproducible training environments.
Kill flakiness. A non-deterministic or slow environment is worse than no environment. Own the interface with labs and researchers: translate a research goal into a system that exists next week.
What We're Looking For
We care more about what you've built than how long you've been building it.
4+ years of production backend engineering, including at least 1 year at a startup, ideally as a founding engineer, an early engineer at a company people have heard of, or a founder yourself
Strong generalist. A master of at least one language; we mostly use Python, and quick to pick up whatever the problem requires
Real distributed-systems depth: scaling, concurrency, isolation, failure modes, performance under load
Working fluency with ML/LLM concepts. Post-training, evaluation metrics, reward modeling: deep enough to partner with researchers and execute on novel RL environments.
Sound judgment in ambiguity. You scope your own work, make pragmatic calls, and ship without a spec handed to you.
Ability to raise the bar around you. Coaching engineers and driving execution, while staying in the code
Bonus, not required: sandboxing or virtualization, browser and computer-use automation, CI/build systems, developer tooling, game or simulation infrastructure.
You'll Succeed Here If
Ownership, impact, and building frontier tech are what motivate you
Your work is a craft you want to master
You thrive in ambiguity and like hard problems
You appreciate diverse perspectives and uncommon ideas
You're excited to build in person, 5 days a week, 10 am–8 pm ET, from our office at One World Trade
Benefits
Equity Compensation
Offers Equity
Mercor builds an AI-powered platform that connects human expertise with AI development, streamlining the hiring process while sourcing and vetting talent. We're dedicated to partnering with AI labs and enterprises, leveraging a vast network of over 30,000 experts who contribute valuable knowledge to train advanced AI models. Our unique approach creates a new category of work where human intelligence drives the future of AI innovation.
- Founded
- Founded 2015
- Employees
- 1-10 employees
- Industry
- Capital Markets