UK remote
Site Reliability Engineer
About this role
Job Title: Site Reliability Platform Engineer About Luupli: Luupli is a social media app that has equity, diversity, and equality at its heart. We believe that social media can be a force for good, and we are committed to creating a platform that maximizes the value that creators and businesses can gain from it, while making a positive impact on society and the planet. Our team is made up of passionate and dedicated individuals who are committed to making Luupli a success.
Role Description: We are seeking a talented and experienced Site Reliability Engineer (SRE) to join our team. As an SRE, you will play a crucial role in ensuring the reliability, scalability, and performance of our cloud-based infrastructure and services, primarily hosted on AWS. If you have a passion for problem-solving, a deep understanding of AWS services, hands-on experience with Terraform, and proficiency in scripting with Python or Bash, we invite you to apply for this exciting opportunity.
Role and Responsibilities: 1. Infrastructure Design and Automation: - Collaborate with software engineering and operations teams to design, build, and maintain cloud-based infrastructure using AWS and Terraform. - Implement and enhance infrastructure-as-code (IaC) practices using Terraform to ensure reproducibility and scalability of infrastructure components. 2. Monitoring and Incident Management: - Develop and maintain monitoring solutions to proactively identify performance bottlenecks, system outages, and other potential issues.
- Participate in incident response and root cause analysis efforts to drive continuous improvement and prevent future incidents. 3. Reliability and Performance Optimization: - Optimise system performance, reliability, and cost efficiency through continuous monitoring, performance tuning, and capacity planning. - Identify opportunities to automate manual processes and improve system resilience. 4. Scripting and Automation: - Utilise Python or Bash scripting to create and maintain automation tools for various operational tasks and deployments.
- Implement and improve continuous integration and continuous deployment (CI/CD) pipelines. 5. Security and Compliance: - Collaborate with security teams to implement best practices for securing cloud infrastructure and services. - Ensure compliance with relevant industry standards and regulations. 6. Deployment and Release Management: - Support CI/CD pipelines for application deployments and updates. - Contribute to the design and implementation of deployment strategies that promote zero-downtime releases.