UK remote
MLOps Engineer
About this role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Learn more at www.coreweave.com . We're proud to be a Living Wage accredited Employer. What You'll Do: CoreWeave’s Physical AI Platform Engineering team builds and scales the data and workflow backbone powering advanced engineering simulation and AI workflows. Our ambition is to become the super-intelligent AI test lab for the engineering industry, delivering the performant, reliable, and trustworthy data foundation trusted by the world’s largest engineering companies.
About the role: As an MLOps Engineer on the Physical AI team, you will serve as the hands-on owner for our machine learning operations surface across the end-to-end model lifecycle—from experimentation and training through to packaging, deployment, serving, and retirement. You will define and roll out MLOps practices, establish operational SLOs/SLAs, and build automated CI/CD and continuous training pipelines to accelerate the path from experiment to supported production deployment.
In this role, you will implement comprehensive model observability, data versioning, and drift monitoring while ensuring robust security and governance controls. Additionally, you will partner closely with product, data science, and core infrastructure teams to optimize GPU compute utilization, resolve cross-boundary platform incidents, and mentor engineers on production-grade ML practices. Who You Are: 5–6+ years of professional experience in MLOps, ML platform engineering, ML infrastructure, or SRE/DevOps for production machine learning systems.
Proven experience building, operating, and automating production ML pipelines covering experiment tracking, model registries, artifact versioning, dataset management, and deployment workflows. Deep hands-on experience implementing observability for ML systems, including monitoring inference availability, latency, throughput, GPU/resource utilization, and data or model drift. Strong background in reliability engineering, including defining operational SLOs, writing runbooks, building automated remediation, and managing incident response for ML workloads.