USA remote
Staff Site Reliability Engineer
About this role
BeyondTrust is a place where you can bring your purpose to life through the work that you do, creating a safer world through our cybersecurity SaaS portfolio. Our culture of flexibility, trust, and continual learning means you will be recognized for your growth, and for the impact you make on our success. You will be surrounded by people who challenge, support, and inspire you to be the best version of yourself. The Role We are seeking a Staff Site Reliability Engineer (SRE) to lead the evolution of the Password Safe platform, infrastructure, and deployment ecosystem.
In this role, you will bridge the gap between systems engineering and software development, architecting highly availability highly resilient systems across both cloud and on-premises environments. As a Staff level engineer, you will tackle complex technical challenges and serve as a technical leader, mentoring engineers and influencing our broader engineering roadmap. You will own the reliability, scalability, and efficiency of shared services, CI/CD pipelines, and platform engineering initiatives via an AI first operating model.
What You’ll Do Infrastructure & Platform Engineering Design, scale, and maintain highly available, secure, and resilient systems spanning both cloud (AWS/Azure) and on premises environments. Champion platform engineering initiatives that reduce cognitive load for software engineers and accelerate velocity. Own and optimize foundational services, including api gateways, service meshes, caches, configuration management, and secrets management.
Drive a "Everything as Code" culture, ensuring that all cloud and on-prem infrastructure, CI/CD pipelines, and configurations are declaratively defined, version-controlled, and deployed via automated GitOps workflows. CI/CD & Automation Standardize, secure, and optimize modern CI/CD pipelines to ensure safe, repeatable, and rapid code deployments. Treat infrastructure as code (Terraform, OpenTofu, or Ansible), driving architectural patterns that eliminate configuration drift.
Reliability, Testing & Observability Design and implement automated chaos engineering frameworks and disaster recovery simulations to proactively identify system weaknesses. Architect and mature our telemetry stack (metrics, logs, traces) using tools like Grafana Cloud, Datadog, and OpenTelemetry to ensure deep system visibility. Define, implement, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across critical applications.