EU remote
Senior Operations Engineer, MetalDev
About this role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Learn more at www.coreweave.com . We're proud to be a Living Wage accredited Employer. What You'll Do CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines high-performance infrastructure with deep technical expertise to accelerate breakthroughs and turn compute into capability.
About the role As a Senior Operations Engineer on the MetalDev Operations team, you will focus on service reliability, observability, and operational excellence across the Redfish-based services and automation tools powering our hardware cloud. Splitting your time between day-to-day operations (~80%) and continuous reliability engineering (~20%), you will troubleshoot production issues across BMCs, servers, DPUs, power shelves, Cooling Distribution Units (CDUs), and NVLink switches supporting GB200/GB300 Vera Rubin NVL72 systems.
You will perform root-cause analysis, manage vendor-provided hardware/firmware change requests, own Prometheus/Grafana observability dashboards, and participate in an on-call rotation to minimize operational toil and ensure high fleet availability. Who You Are 5+ years of experience in cloud operations, site reliability engineering (SRE), or infrastructure operations. Hands-on experience deploying, supporting, and troubleshooting containerized applications within Kubernetes environments.
Strong Linux system administration and internals knowledge, with proficiency in shell scripting or modern automation languages. Proven experience in incident management, structured troubleshooting, root-cause analysis, and participating in production on-call rotations. Technical proficiency with observability, monitoring, and alerting frameworks using Prometheus, PromQL, and Grafana. Demonstrated ability to author high-quality runbooks, troubleshooting guides, and post-incident documentation.