[Job Radar](https://jobradar.live/) / [Jobs](https://jobradar.live/ilanlar) / Research Scientist, Scaling RL

Active
On-site
Menlo Park, CA; Montreal, Canada, United States
Posted · 01.10.2026
Ashby (US)

# Research Scientist, Scaling RL

Periodic Labs

ABOUT PERIODIC LABS
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.
ABOUT THE ROLE
We're training frontier models to develop deep scientific knowledge and reasoning for scientific tasks. You’ll study how RL scales with training compute, develop better algorithms, and take ideas from controlled experiments to our largest runs like Periodic Neon https://periodic.com/news/nature-is-our-learning-environment.
WHAT YOU'LL DO
• Design experiments to understand how RL performance scales with compute, model size, data, and reward quality, building on work such as ScaleRL https://arxiv.org/abs/2510.13786
• Develop better RL algorithms, spanning policy optimization, advantage estimation, exploration, and credit assignment for long-horizon RL tasks
• Build adaptive sampling and curriculum methods that adjust task difficulty, problem selection, and the number of rollouts as models improve
• Study bias and stability during RL training, including importance-sampling corrections and methods to tackle policy staleness and training–inference mismatch, as discussed here https://www.youtube.com/watch?v=GH4JCdAAUYg.
• Improve compute efficiency across training and inference through experiments with hyperparameters, such as length penalties, rollout counts, batch sizes, and update schedules.
YOU WILL THRIVE IN THIS ROLE IF YOU HAVE
• Hands-on experience training LLMs with reinforcement learning
• Strong attention to detail and rigorous approach to answer questions scientifically.
• Coming up with small-scale RL setups that transfers to large-scale training runs.
• Comfort working across a complex training stack to implement, debug, and test new research ideas.
Mechanics
Minimum education: Bachelor’s degree or similar experience
Location: Menlo Park, CA or Montreal, Canada
Compensation: $225,000-$350,000 base + equity
Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

This job was verified from Ashby (US). Applications are completed on the original source.

[Apply on the original listing ↗](https://jobradar.live/ilan/94967148-3650-4dd3-8fff-293b9a538be7/git)

## Stop searching one by one for roles like this.

Upload your resume or enter your target roles to see your first 3 matches for free.

[Find jobs for me →](https://jobradar.live/uye/kayit)
Your resume is never shared with employers; it is processed only for matching.

Something wrong with this job?

## Similar jobs

[Periodic Labs Jobs](https://jobradar.live/company/periodic-labs) · [Jobs in California](https://jobradar.live/jobs/california)
· [Data Scientist Jobs in California](https://jobradar.live/jobs/california/data-scientist) · [Jobs by location](https://jobradar.live/jobs) · [Jobs by company](https://jobradar.live/company)

- [Research Scientist](https://jobradar.live/ilan/4c2abb26-e18f-41ae-a75f-0a533f2129f6) Latent · San Francisco, California, United States

- [Machine Learning Engineer](https://jobradar.live/ilan/08dd75c3-ebde-4832-a056-a41193756f39) Latent · San Francisco, California, United States

- [AI Engineer – Decision & Optimization Systems](https://jobradar.live/ilan/1cf2f435-82c8-4402-883d-9d8152f053ac) Gallatin · El Segundo, CA, California, United States

- [AI Engineer, Internal Systems](https://jobradar.live/ilan/a73277e1-dd0f-4ce5-adbb-724351e90ce4) Wispr Flow · San Francisco, California, United States

- [Research Scientist, Condensed Matter Theory](https://jobradar.live/ilan/27026554-8111-415b-b84e-8cbbb8a3f984) Periodic Labs · Menlo Park, CA; Montreal, Canada, United States

- [2027 Internship Behavior ML Engineer, Learned Manipulation Policies](https://jobradar.live/ilan/ca194be6-41e7-4535-b8e9-e00631402ec0) Bedrock Robotics Inc · San Francisco, CA, California, United States

- [Machine Learning Engineer - Content Discovery](https://jobradar.live/ilan/830d8a7c-55f1-4892-9ed4-ed4dc8dd68f9) Suno · San Francisco, California, United States

- [Machine Learning Software Engineer](https://jobradar.live/ilan/0981d1aa-08f1-4d1c-b060-a7a38c20d953) Wayve · Sunnyvale, California USA
