All Positions

AI Evaluation Researcher

Research Remote / Global Full-time

About the Role

Lead the development of our proprietary benchmark methodology. You'll design evaluation protocols that measure what academic benchmarks cannot — real-world operational reliability.

Requirements

- PhD or equivalent research experience in ML/AI
- Published work in model evaluation, benchmarking, or related areas
- Experience with frontier model capabilities and limitations
- Strong experimental design and statistical analysis skills