About the Role
Design and build high-fidelity RL training environments for frontier model development. You'll work directly with domain experts to translate real-world workflows into structured state spaces, task sequences, and reward functions.
Requirements
- 3+ years experience in RL environment design or simulation engineering
- Strong Python skills and experience with RL frameworks
- Understanding of reward shaping and curriculum learning
- Experience with multi-step, long-horizon task design
- Bonus: domain expertise in coding, math, or scientific reasoning