EU remote
Mistral Cloud - Software Engineer, Managed Kubernetes
About this role
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society.
Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. The Role We’re looking for passionate and talented Software Engineers to join our Compute team. In this role, you’ll have the opportunity to shape and deliver on-demand Kubernetes clusters with GPU support for Training, Inference, and general-purpose CPU workloads.
More information on Mistral Compute here: https://mistral.ai/products/compute Location and Office Policy We will prioritize candidates who either reside in one of our main offices (Paris, London, NYC) or are open to relocating. We will also consider remote candidates based in the following countries: EMEA : France, United Kingdom, Germany, Switzerland, Netherlands, Spain, Austria, Poland, Luxembourg Americas : USA, Canada We strongly believe in the value of in-person collaboration to foster strong relationships and seamless communication within our team.
In any case, we ask all new hires to visit our Paris HQ office (accommodation and travelling covered) • for the first week of their onboarding • then at least 3 days every month What You Will Do Software Development: Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes Cluster lifecycle management: Develop automation including provisioning, upgrades, patching, and decommissioning.
Monitoring and Observability: Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime Internal tooling: Create workflows, tools, APIs, and command-line interfaces (CLIs) to empower customers and ML/AI teams to deploy and monitor inference services efficiently. Reliability: Design resilient systems capable of gracefully handling failures in large-scale distributed environments.