EU remote
Site Reliability Engineer, IaaS
About this role
At Algolia, we’re proud to be a pioneer and market leader in AI Search, empowering 18,000+ businesses to deliver blazing-fast, predictive search and browse experiences at internet scale. Every week, we power over 30 billion search requests — four times more than Microsoft Bing, Yahoo, Baidu, Yandex, and DuckDuckGo combined. In 2021, we raised $150 million in Series D funding, quadrupling our valuation to $2.25 billion.
This strong foundation enables us to keep investing in our market-leading platform and serving incredible customers like Under Armour, PetSmart, Stripe, Gymshark, and Walgreens. The team The Infrastructure as a Service team is at the center of one of Algolia’s most consequential engineering transformations. For years, Algolia has operated a production fleet of approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability that our customers expect.
We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come. . We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come. This is not a lift-and-shift project. It is an opportunity to rethink how Algolia provisions, secures, operates, observes, upgrades, and scales production infrastructure and to build it as a platform that engineers can safely consume, rather than a queue of manual requests.
The opportunity As a Site Reliability Engineer in IaaS, you will help build the next generation of Algolia’s production infrastructure. You will contribute to the cloud baseline and reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, automation, reliability, and large-scale production operations. As a P3 engineer, you will be a hands-on contributor.
You will build, operate, and improve production infrastructure while developing deep expertise in cloud, Kubernetes, reliability, and automation. YOU WILL: Build and improve Cloud Baseline capabilities, including identity and access, networking, security, resource inventory, tagging, and auditability. Develop and maintain infrastructure as code and automation for cloud environments and Kubernetes infrastructure. Contribute to reliable, repeatable cloud and cluster lifecycle operations.