UK remote
Lead DevOps Engineer
About this role
About the Role Reporting to the Head of DevOps, you'll provide technical leadership across cloud infrastructure, platform engineering and automation while staying hands-on in delivery. The Head of DevOps owns the overall function and people leadership across our UK and Pune capabilities. You'll be the senior technical leader, working alongside the engineers. You won't lead from the sidelines. You'll be in the infrastructure, pipelines and code alongside the team, shaping our DevOps practice and turning good engineering principles into solutions that work in production.
You'll design and evolve secure, scalable and cost-effective infrastructure, keeping our platforms reliable, observable and efficient. Working with engineering, security and product teams, you'll take difficult technical challenges from design through to a stable production outcome. If you want real technical depth with genuine leadership, where you can set direction and still know what's happening under the hood, this is built for you.
What You'll Do Own our AWS and Azure infrastructure. You'll set practical standards for security, reliability, scalability, performance, compliance, cost and resource use, working with architects, developers, security and product teams to ensure they work in the real world. Write and maintain the Terraform that runs our infrastructure. Build reusable modules, manage state and environments, and keep documentation up to date.
Use the consoles to diagnose problems, but make lasting changes through code, review and controlled deployment. Get hands-on with our CircleCI pipelines, creating standards practical enough to be adopted, not simply written down. Embed secret detection, dependency and container scanning, software bills of materials, code-quality checks and supply-chain controls, then help teams fix the risks that matter. Run our production Kubernetes platforms day-to-day, working directly with Argo CD, Helm and container tooling to deploy, upgrade, optimise and troubleshoot.
Keep a close eye on availability, performance and resource consumption. Improve metrics, logs, dashboards and alerts so the team can spot trouble coming. When something goes wrong, lead the technical response, root-cause analysis and preventative actions. Ensure we can recover as effectively as we can run. Help maintain and test our backup, resilience and disaster-recovery arrangements, using Python or Bash to remove repetitive work and improve consistency.