EU remote
Staff Software Engineer, Observability
About this role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Learn more at www.coreweave.com . We're proud to be a Living Wage accredited Employer. What You'll Do: The Observability team is responsible for deploying and maintaining critical infrastructure at CoreWeave, including our global logging, tracing, and metrics platforms as well as the high-throughput pipelines that feed them. We build the foundational telemetry systems that enable engineering teams across CoreWeave to operate our massive GPU cloud reliably and at scale.
About the role: As a Staff Software Engineer, Observability, you will lead our efforts in building, maintaining, and optimising highly scalable, reliable, and secure observability systems. You will take ownership of scaling our core logging, tracing, and metrics platforms to support an expanding global data centre footprint. In this role, you will develop and refine monitoring and alerting frameworks, automate interactions with CoreWeave’s compute infrastructure layer, and manage production clusters to ensure development teams follow deployment best practices.
Additionally, you will advise engineers across the company on optimal observability usage, lead incident management and post-mortem analyses, and mentor engineers to foster a culture of technical excellence and collaboration. Who You Are: Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience). 7+ years of professional experience in Software Engineering, Site Reliability Engineering (SRE), DevOps, or a related technical field.
Deep technical expertise across all observability pillars (logging, metrics, tracing) using ecosystem tools such as ClickHouse, Elastic, Loki, VictoriaMetrics, Prometheus, Thanos, and Grafana. Advanced expertise in Kubernetes, containerisation, and microservices architectures at scale. Proven track record of leading production incident management, root cause analysis, and post-mortem processes. Strong analytical, problem-solving, and cross-functional communication skills.