About the Role
Lead the development of our proprietary benchmark methodology. You'll design evaluation protocols that measure what academic benchmarks cannot — real-world operational reliability.
Requirements
- PhD or equivalent research experience in ML/AI
- Published work in model evaluation, benchmarking, or related areas
- Experience with frontier model capabilities and limitations
- Strong experimental design and statistical analysis skills