USA remote
Senior SRE
About this role
The Role We're building the financial data platform at Accelerant — the premium, claims, and paid data products that underpin financial processing, reserving analysis, and the monthly close — and it needs to stay fast, resilient, and observable as we scale. You'll drive the reliability and observability strategy across the platform and the enterprise systems it depends on: Velocity, MuleSoft, D365, Snowflake, Fabric, and the streaming and integration layers that move data through it.
You are a key decider about what gets measured, how we define reliability, and where engineering needs to invest to keep production healthy. We need someone who can prove a repeatable, define-to-alert observability pipeline, harden it, and scale it into systems that have never had real SLOs — and build modern, AI-assisted operational tooling that lets a small team punch far above its weight. This Is a High-Autonomy, High-Impact Role for Someone Who: • Sees a recurring alert or a fragile deploy path and cannot leave it alone.
Excels at shipping the right fix and the right automation, not the perfect one. • Has run real production systems at scale — not just written runbooks for them. • Has built with Datadog, OpenTelemetry, incident tooling, and AI coding assistants long enough to have strong opinions about what fits our needs. • Can prototype an operational agent in Cursor and iterate as they go. • Operates with autonomy, and can carry a technical discussion on system architecture, failure modes, and tradeoffs.
• Is genuinely curious about applying emerging AI to reliability and operations. What You'll Do Drive the reliability and observability initiative • Own the reliability roadmap end to end. Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure.
Harden the foundational platform • Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired. Expand observability breadth and depth • Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion.