EU remote
Reliability Operations Specialist
About this role
- Incident, Problem & Change Management, Reporting & Insights - About the rol e We're looking for a Reliability Operations Specialist to help drive operational excellence across incident management, service reliability, observability, and continuous improvement initiatives. This role serves as a central coordinator and subject matter expert for reliability practices, helping teams improve service stability, reduce operational risk, and strengthen operational readiness across the organization.
The Reliability Operations Specialist partners closely with engineering, infrastructure, security, and operations teams to facilitate incident response, oversee post-incident reviews, track corrective actions, and provide visibility into the health and reliability of our platforms. This role does not have direct people management responsibilities but plays a critical role in influencing reliability outcomes through collaboration, process ownership, and data-driven decision making.
Location: Remote Employment type: Permanent, Full-time Pay Range: $85,000 - $100,000 Annually, The final compensation offered will be determined based on factors including location, experience, skills, qualifications, and market conditions. What you'll do Incident Management & Operational Excellence Participate in major incident response activities and serve as an Incident Commander when assigned. Facilitate incident coordination, escalation, stakeholder communications, and status reporting during service-impacting events.
Support ongoing improvement of incident management processes, procedures, and operational readiness. Drive initiatives focused on reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). Maintain and apply the Criticality Matrix to tier services, infrastructure, and customer MRR impact. Mortem Management & Corrective Actions Coordinate and facilitate post-mortem reviews following significant incidents. Ensure post-mortems are completed accurately, consistently, and within established timelines.
Synthesize findings across incidents to identify trends, recurring issues, and systemic risks. Maintain accountability for corrective action tracking and closure. Promote a blameless culture of learning and continuous improvement and proactive/reactive problem management. Reliability Strategy & Observability Partner with engineering teams to define and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) across our product lines and services.