EU remote
Senior Platform Engineer - Data Infrastructure
About this role
About the team The Data Infrastructure team looks after the data backbone: how those events are taken in, processed, stored, and served, and everything that keeps them reliable, fast, and affordable to run. When a piece of data infrastructure becomes important enough that several teams depend on it, looking after it properly becomes our job. Today that estate covers Kafka (Amazon MSK), ClickHouse (on Amazon EKS), Amazon Aurora, and MongoDB (Atlas) across several AWS regions.
We are now adding search, vector, and in-memory stores (OpenSearch, Redis, ElastiCache) to support new product work. You would join a small group of experienced engineers who own their areas from the first design through to running them in production. We like DevOps and CI/CD, we automate the things worth automating, and above all we care about doing the work well. Why this role matters Almost everything our customers rely on sits on top of this layer.
When a product team promises customers high availability for what they have built, that promise only holds if the Kafka, the databases, and the streaming underneath hold too. Your work allows Nexthink to move quickly and trust the ground they are standing on. What you will do Design and improve our data infrastructure alongside the Architecture, Product Engineering, and Security teams, following cloud-native good practice.
Build the tooling and automation that provisions and scales it, with a real focus on resilience and elasticity, and make it self-service (Crossplane) so product teams can build on it with confidence. Bring in and run new kinds of data store (search, vector, in-memory and caching) to the same standard as everything else we look after. Spend time with the product and feature teams, understand what they actually need, and bring that back to shape the platform.
Plan for the bad days: disaster recovery and cross-region replication, with clear RPO and RTO targets. Keep an eye on availability, performance, and observability (Datadog) so you spot trouble before it turns into an incident. Handle incidents from start to finish: spot them, work out what happened, fix them to SLA, and write the post-mortem. You will share the on-call rotation with the rest of the team. How we work A few things that are true about this team and, we think, make it a good place to build: It is a small team with real ownership.