EU remote
Sr. Site Reliability Engineer - SRE
About this role
We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance of our mission-critical services. This is an opportunity to shape our SRE practices, drive automation, reduce operational toil, and significantly impact our product's operational excellence.
What You'll Do Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experiences. Serve as a subject matter expert for observability, including monitoring, alerting, logging, tracing, dashboards, and synthetic testing. Develop robust, maintainable software and self-service tooling to automate operational tasks and improve reliability. Identify and eliminate operational toil through automation, process improvements, and systematic problem solving.
Lead incident response, participate in on-call rotations, and drive blameless post-mortems. Define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Leverage infrastructure as code, GitOps practices, and CI/CD automation using Terraform, Flux, and GitHub Actions. Provide reliability expertise during system design reviews and influence architectural decisions. Document processes, build runbooks, and mentor engineers across the organisation.
Leverage AI responsibly to accelerate investigations, improve documentation, reduce toil, and build intelligent operational workflows while maintaining appropriate human oversight, security, and governance. What You'll Bring Core SRE Capabilities Demonstrated experience operating and improving production systems at scale in an SRE, Production Engineering, or Platform Engineering role. Ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability domains.
Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis. Experience defining and using SLIs, SLOs, and error budgets to guide reliability decisions. Excellent written and verbal communication skills. Technical Domains Experience across several of the following areas: Kubernetes platforms, including Amazon EKS, and service mesh technologies such as Istio. Cloud infrastructure and services within AWS.