EU remote
Datacenter Hardware Engineer, HPC
About this role
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society.
Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. Role Summary Our compute footprint is growing fast to support our science and engineering teams. We’re hiring a Datacenter HW Engineer to maintain, troubleshoot, and scale our GPU/CPU clusters safely and reliably.
You’ll execute hands-on hardware work in our Paris-area datacenter and partner with hardware owners, DC operations, and vendors to keep one of France’s largest GPU clusters healthy. Location: Bruyères-le-Châtel — on-site, field role Impact • Compute is a key lever for Mistral’s success and our largest spend item. • Direct impact on scale: your work keeps one of France’s largest AI clusters healthy as we grow to unprecedented scale.
• Enable breakthrough AI: you unlock our science & engineering teams to deliver groundbreaking AI solutions . What you will do • Diagnose & operate core server/cluster components - Investigate and handle compute/storage hardware issues ( CPU, memory, drives, NICs, GPUs, PSUs ) and interconnect problems ( switches, cables, transceivers; Ethernet/InfiniBand ). Perform safe interventions (power-off/lockout, ESD ) to replace, re-seat, or recable components and restore service.
• Safety & procedures - Apply lockout/tagout (LOTO) and ESD discipline; follow pre/post-work checklists; maintain tidy, safe work areas. • First-line diagnostics - Triage using LEDs, POST, beep codes and basic tests; capture evidence (photos, serials, results); open/update/close tickets with clear notes. • Preventive maintenance - Provide feedback and ideas to improve proactive activities, monitoring, and targeted follow-ups on recurring or specific anomalies; help turn ad-hoc checks into SOPs, alerts, and dashboards.