USA remote
Site Reliability Engineer
About this role
Here at Ooma we empower people to connect in smarter ways. We do this by creating powerful communication experiences through our cloud-based platform to bring people together at work and at home. Our solutions help small business owners stay connected with their customers and manage their businesses from anywhere. For larger companies we provide customized unified communications solutions to meet their unique needs. At home, we help our customers connect with their loved ones by providing the #1 rated VoIP phone service available.
We also provide them with peace of mind through our innovative smart home security solution. At Ooma, all our products and services are priced competitively, because we believe advanced technology should be accessible to all. About the Role: As a Site Reliability Engineer, you will leverage your extensive expertise in Linux systems, virtualization, containers, Kubernetes clusters, and CI/CD pipelines to ensure the stability and efficiency of our systems, collaborating across teams to implement best practices for infrastructure management, automated deployment, and application performance monitoring.
Deep on-premises experience is a core requirement, not a secondary consideration. Our production environment runs on our own hardware — large data centers built on hundreds of bare metal servers and VMs, with our own storage, virtualization, and physical network beneath them. You will operate comfortably at the hardware and OS layer while also owning the container and delivery platform on top of it. **Location and Onsite Requirement: This role requires onsite work at least once per week at one of our designated data centers in Dallas, TX, Ashburn, VA or San Jose, CA.
Candidates must be able to commute regularly to one of these locations. Relocation assistance and reimbursement for routine commuting, travel, or overnight lodging are not available. What You’ll Do: Provide expert guidance on managing large data centers, including hundreds of bare metal servers and virtual machines (VMs), ensuring optimal configuration and performance. Monitor and troubleshoot system performance, reliability, and availability using modern observability tools and techniques, with strong emphasis on diagnosing and resolving issues in operating systems and bare metal environments.