Remote
Senior Site Reliability Engineering
About this role
📋 Description Collaborate with internal teams to identify sources of instability in distributed systems and drive Support core infrastructure (i.e understand, diagnose, and debug these systems in production) Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews Maintain and document sustainable postmortem/incident response practices Advocate for and implement changes that improve reliability, scalability, and velocity Reduce the burden of toil with iterative development of tooling and automation 🎯 Requirements 5+ years of experience within site reliability engineering/DevOps of a product with millions of Experience identifying and solving issues in large-scale distributed systems Experience with Java, Kotlin, Python or Go An understanding of containerization toolsets and container orchestration technologies (Docker Experience in improving automation and tooling to reduce service maintenance toil Proven experience driving improvements to incident response processes 🎁 Benefits base salary supplemented by equity compensation holistic well-being benefits (see https://careers.duolingo.com/#benefits) accommodations for interviews (contact [email protected]) Equal Employment Opportunity AI-assisted hiring with human review
Source listing: empllo_remote