Remote
Principal Software Engineer
About this role
POS-5690 About the Team The Observability team owns the internal platform that gives every HubSpot engineer real visibility into how their systems behave in production. We build and operate the distributed tracing, metrics, alerting, and logging infrastructure that spans hundreds of microservices, billions of daily events, and thousands of engineers who depend on that signal to ship reliably. We are now investing in the next generation of this platform.
As HubSpot deploys AI agents and ML-powered features across the product, the team is building the tracing and telemetry primitives that make it possible to understand, debug, and trust what those systems are doing in production. This is greenfield, technically interesting work at a scale few companies operate at and we are looking for a Principal Engineer to help lead it. About the Role We are seeking a Principal Software Engineer to be the technical anchor for HubSpot’s Observability platform.
This role sits at the intersection of large-scale distributed systems, developer platform design, and AI observability. A big part of this role is working horizontally across a large engineering org and setting patterns and standards that make it easier for teams to instrument, alert on, and reason about their services. You will also shape how we trace and understand our growing fleet of AI agents and ML systems in production: a technically distinct and increasingly critical problem.
Key Expectations Observability Platform Architecture: Define the patterns and evolution of HubSpot’s core telemetry platform — distributed tracing, metrics, and structured logging — at a scale that spans hundreds of services and billions of daily events. Set the standards for how instrumentation is done across a large, polyglot engineering organization. AI & Agentic Observability: Lead the technical strategy for tracing and understanding AI agents and ML-powered systems in production.
Define the primitives, telemetry standards, and debugging workflows that help product engineers understand what their models and agents are doing — and build trust in those systems over time. This is greenfield and consequential work. High-Cardinality, High-Throughput Systems: Architect telemetry pipelines and storage systems that handle high-cardinality data at high throughput without blowing up cost or query latency.