EU remote
Site Reliability Engineer, Mistral Cloud
About this role
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society.
Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. The Role As a Site Reliability Engineer (SRE) on the Cloud Platform team, you will shape the reliability, scalability, and performance of our Cloud platform and customer-facing applications. You’ll work closely with software engineers and product teams to ensure our systems meet and exceed the expectations of both internal and external customers.
This role is critical in maintaining the stability and efficiency of our infrastructure, enabling seamless experiences for users and developers. Your expertise will directly impact the robustness of our AI platform, ensuring it operates at scale with minimal downtime. What You Will Do Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform. Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, and infrastructure scaling.
Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance. Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging. Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent recurrence. Drive continuous improvement in infrastructure automation, deployment, and orchestration.
Collaborate with software engineers to enable safe and reproducible model-training experiments. Build and enhance a cloud platform that abstracts infrastructure complexities for science and engineering teams. Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance. Ensure infrastructure adheres to security best practices and compliance requirements in collaboration with the security team.