Remote
Software Engineer, Infrastructure
About this role
Platform owns the foundation the company runs on: compute and deployment, reliability and on-call, CI/CD and monorepo health, developer environments, the infrastructure that model training and inference run on, and the security boundaries around all of it. AI and agent tooling is the clearest example. You will contribute to how models are trained and served here, as well as how agents work inside our codebase: the environments they run in, the verification that makes their output trustworthy, and the review paths that keep it all legible.
There is no industry standard and you’ll help form our opinions rather than inheriting one. Our users are other engineers. Expect to spend time with all the other engineering teams, understanding their needs and building a roadmap. The scope is large, so you'll be choosing what to leave alone as much as build. Architecture decisions here last: this is a small team covering a large surface. You'll have real room to decide things, and you'll stay close to the systems you decide about.
What you'll own Own our platform: GCP, Kubernetes, Temporal, the GPU fleet behind cloud export, and the deploy and rollback machinery everything ships through. You'll be in the on-call rotation, and we'll expect you to make it quieter and more actionable. The AI enablement substrate: GPU capacity, training and inference pipelines, and the reliability and cost of the systems serving models in production. Cost is an engineering constraint: infrastructure decisions have a number attached and you're expected to make smart trade-offs.
As inference grows with usage, the metering and attribution behind those numbers sit with this team. Security comes with the systems you run : identity and access, secrets management, least-privilege boundaries, and supply-chain integrity. Make what you build legible. Infrastructure-as-code that explains itself. Runbooks and in-repo context written for someone jumping in to help. Observability that tells you what failed and why.
Improve how the team learns and ships. Form hypotheses, instrument your work, release incrementally, and read results honestly. Strengthen the tooling, standards, tests, observability, and release practices that help the team move quickly without compromising quality. Raise the team’s technical ambition. Provide architectural direction, thoughtful reviews, mentoring, and clear human writing. What you bring Required: You have 8+ years building and operating production distributed systems, or equivalent server-side engineering with a heavy infrastructure focus.