EU remote
Staff Software Engineer, Cluster Orch (Non SUNK)
About this role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Learn more at www.coreweave.com . We're proud to be a Living Wage accredited Employer. What You'll Do: CoreWeave’s AI Workload Orchestration Platform team builds and operates the core, Kubernetes-native substrate that governs how massive AI workloads are admitted, scheduled, and executed across our global GPU footprint. Serving as a strategic complement to SUNK (Slurm on Kubernetes), this platform underpins both training and inference pipelines across the CoreWeave cloud, ensuring highly efficient resource utilisation for the world's most demanding AI applications.
About the role: As a Staff Software Engineer, Cluster Orch (Non SUNK), you will act as a principal technical leader driving CoreWeave’s Kubernetes-native orchestration strategy. You will own the technical vision and architecture for major portions of the platform, defining how AI workloads are admitted, scheduled, and governed across large GPU clusters using frameworks such as Kueue, Volcano, and Ray. In this high-impact role, you will apply systems thinking to resolve systemic performance, scalability, and fairness issues at scale.
Additionally, you will lead cross-team architecture reviews, drive technical alignment across broader infrastructure, CKS, and managed inference teams, and establish platform-wide standards for reliability, capacity management, and developer experience while mentoring senior engineers across the organisation. Who You Are: Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
8+ years of professional software engineering experience, with deep technical expertise in distributed systems or cloud platforms. Strong software development proficiency in Go, with a proven track record of designing large-scale, long-lived production systems. Deep technical knowledge of Kubernetes internals, scheduling mechanisms, Custom Resource Definitions (CRDs), and controller-based architectures. Demonstrated engineering experience designing, scaling, or evolving orchestration, scheduling, or hardware resource-management platforms.