UK remote
Senior DevOps Engineer
About this role
Build the infrastructure powering the future of construction At XYZ Reality , we’ve created the world’s first engineering-grade Augmented Reality solution for construction. Our technology, The Atom, is used on major projects worldwide to help construction teams build more accurately, efficiently and with fewer mistakes. As a Series B business scaling across the UK, US and Europe, the infrastructure behind our technology needs to scale with us.
We’re looking for a Senior DevOps Engineer to take ownership of the cloud infrastructure and DevOps practices powering our BIM Platform, working across Azure, Kubernetes, CI/CD, observability, security and developer tooling. The Role This is a hands-on senior engineering role with broad ownership across our cloud and platform infrastructure. You’ll work closely with our Engineering, Data and Security teams to build reliable, scalable infrastructure while improving how our developers deploy, monitor and troubleshoot their services.
We want DevOps to be an enabler rather than a gatekeeper. That means building the tooling, automation and self-service capabilities that allow engineering teams to move quickly and safely without creating unnecessary dependencies on DevOps. You’ll have the autonomy to identify where our infrastructure and engineering practices can improve and take ownership of making those improvements happen. The role is based in our London office on a hybrid basis, with 3 days per week in the office.
What You’ll Be Doing Architect, maintain and evolve our Microsoft Azure infrastructure and Kubernetes clusters Build and maintain Infrastructure-as-Code using Terraform, ARM templates and Helm Design and improve CI/CD pipelines using GitHub Actions, enabling fast, safe and repeatable deployments Build observability across our infrastructure and services through monitoring, logging, alerting and distributed tracing Improve platform reliability, scalability and performance through capacity planning, autoscaling and resource optimisation Build and maintain disaster recovery, backup and failover strategies Strengthen infrastructure security across areas including network policies, secrets management and platform hardening Establish and improve incident response processes, on-call practices, runbooks and blameless post-mortems Work with our Data team to support reliable and scalable PostgreSQL and MongoDB environments Build self-service tooling, templates and documentation that reduce unnecessary reliance on DevOps Support infrastructure for data pipelines and AI/ML model deployment as our technology continues to evolve Partner directly with developers on deployment, troubleshooting and infrastructure best practices.