EU remote
Site Reliability Engineer, Infrastructure Engineering
About this role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Learn more at www.coreweave.com . We're proud to be a Living Wage accredited Employer. What You'll Do The MetalDev team within CoreWeave's Hardware Compute organisation develops software automation tooling and services used to bring up data centre rack systems and manage bare-metal infrastructure. We provide core reliability, availability, and operational stability functions across regional data centres to ensure seamless infrastructure provisioning.
About the role As a Site Reliability Engineer on the MetalDev team, you will split your focus between production operations and reliability (60%) and engineering automation (40%). You will lead incident response, troubleshooting, root-cause analyses, and post-incident reviews while participating in an on-call rotation. In this role, you will write resilient Go code, build Prometheus and Grafana dashboards, and develop automated remediation workflows to reduce manual overhead.
Additionally, you will define SLOs and KPIs, improve CI/CD deployment pipelines, and create self-service tooling for Fleet Operations and Hardware engineering teams. Who You Are Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience). 3+ years of experience in Site Reliability Engineering, production engineering, cloud infrastructure, or software engineering.
Working proficiency in Go with experience developing production-quality software. Hands-on production experience with Kubernetes and containerised microservices. Experience with observability and telemetry stacks, specifically Prometheus and Grafana. Demonstrated track record supporting production services, leading incident management, and participating in on-call rotations. Strong troubleshooting, analytical, and technical documentation skills.