UK remote
Principal AI Evaluations Platform Engineer
About this role
As Microsoft continues to push the boundaries of AI, we are looking for passionate individuals to work with us on some of the most interesting and challenging AI questions of our time. Our vision is bold and broad: to build systems with true artificial intelligence across agents, applications, services, and infrastructure. It is also inclusive: we aim to make AI accessible to consumers, businesses, and developers so that everyone can realize its benefits.
Microsoft AI (MS AI) is seeking an experienced engineer to help build and operate the evaluation platform that supports large-scale model training and development. We’re looking for someone who combines strong systems engineering fundamentals with operational excellence and who can build reliable evaluation infrastructure at scale. This role will work closely with researchers and training teams across our European offices to ensure evaluations are reliable, efficient, and available when teams need them.
We seek a versatile engineer who can build solutions that stand the test of time and who brings positive energy, empathy, and kindness to the team while remaining highly effective in a fast-paced environment. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees, we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals.
Each day, we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. This role is based in London, U.K., or Zurich, Switzerland. Candidates are expected to be local to the applicable office and work in the office four days per week. Responsibilities Serve as a primary engineer for the evaluation platform during European hours, providing dedicated coverage for the evaluation stack supporting large-scale model training.
Develop and extend core evaluation platform capabilities, including benchmark configurations, problem sets, graders, and evaluation runners. Improve evaluation throughput and scheduling efficiency. Build and maintain monitoring, alerting, and dashboards to proactively detect evaluation degradation. Partner closely with researchers and training teams across European offices to define evaluation requirements, coordinate launches, and unblock experiments.