EU remote
Technical Lead, Senior AI Engineer - VonHalsky (m/f/n)
About this role
Why this role exists Von Halsky is InPost's conversational AI shopping assistant, live in production and serving a growing share of our customers. What decides whether it wins is not the model but whether we can tell, at release cadence, that a change made conversations better. That is the Evaluations Platform, and we are hiring the engineer who takes it to the next level. What you will own The LLM-judge pipeline. Our release-gating judge over real and golden conversations, calibrated well enough that people act on its verdicts instead of arguing about it.
Eval datasets and the golden set. Real coverage across intents and categories, including the Polish-language coverage generic benchmarks do not give us. Root-cause analysis on real conversations. Making our conversation-mining stack diagnostic rather than descriptive, on a stable issue taxonomy. The "sus" detector. Abusive, adversarial and anomalous sessions, next to our guardrails and red-team work. Evals-driven development.
The eval comes before the feature, and writing it is as cheap as writing the code. The interface to product and business. Vague asks in, measurable quality definitions out, and results stakeholders can act on. What "leadership aspirations" means here concretely This is a technical lead role, not a people-management role, and you will not carry line-management duties on day one. You will set and defend the technical direction for the platform, act as reviewer of record for the area, mentor other engineers, scope work with our PM and EM, present results to stakeholders, and hold the line against ad-hoc requests crowding out platform work.
Engineering management later, or a Staff-level hands-on track, are both paths we will build with you. Either way we need someone accountable for an area rather than for a ticket. How we define success in this role Judge pass rate is calibrated against human labels and is a number people trust and cite. A stable, versioned issue taxonomy is live and week-over-week trends are comparable. Every release is gated by an eval run the team can reproduce.
The golden set has documented coverage and named blind spots. At least two engineers besides you can operate and extend the platform. What we are looking for Required 5+ years building and running production software, with 2+ years on LLM-based systems that real users hit. Strong engineering fundamentals , plus the habits that go with production ownership: testing, CI/CD, containers, observability, and working in cloud.