EU remote
Site Reliability Engineer
About this role
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society.
Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. The Role As a Site Reliability Engineer (SRE) on the Platform team, you will shape the reliability, scalability, and performance of our platform and customer-facing applications. You’ll work closely with software engineers and research teams to ensure our systems meet and exceed the expectations of both internal and external customers.
This role balances day-to-day operations on production systems with long-term software engineering improvements. Your work will reduce operational toil, foster reliability, and ensure high availability for our web services, inference environments, and ML workloads. You’ll enable seamless replication of work environments across multiple HPC clusters, directly impacting the stability and efficiency of our AI platform. What You Will Do Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.
Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters. Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling. Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance. Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
Participate in on-call rotations to respond to incidents and perform root cause analysis. Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform. Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments. Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams.