UK remote
Machine Learning Engineer
About this role
We pull medical technology from the future to solve human health. Scarlet is authorised to assess and certify medical devices. We combine clinical, technical and regulatory expertise with AI agents and software so rigorous certification can keep pace with product development at the world’s most ambitious technology companies, without lowering the safety bar. Our customers have cut a year or more from their certification timelines for new AI-enabled medical devices and shortened product-update cycles from months to weeks.
You’ll join a team building the infrastructure that makes these outcomes repeatable at scale. About the role The Applied Machine Learning team owns production systems and pursues new ideas, from conception and prototyping through to deployment, evaluation and iterative improvement. Working with clinicians, assessors and engineers, we build reliable ML systems that help bring medical devices to market faster, without compromising safety.
It’s a domain rich in text, data and expert judgment, with little precedent for much of what we’re building. You’ll own meaningful problems in how we assess medical devices: define success with domain experts, decide what to build and test, and take ML systems through deployment and measurable improvement in production. Responsibilities Things you might work on: Agentic document understanding – Build agents that search, parse and visually inspect messy technical files: scanned certificates, tables, architecture diagrams and thousands of pages of evidence.
Find what matters and show exactly where it came from. Harnesses built for evidence – Develop custom agent harnesses for retrieving information across large document collections in varied formats. Preserve source attribution, minimise hallucination, and make deliberate trade-offs between accuracy, latency and cost. Deploy to production. Evals wired into real workflows – Define success with assessors and build datasets and benchmarks that capture a complex, nuanced domain.
Measure retrieval quality, citation correctness and agreement with expert judgment alongside the impact on assessor effort, assessment quality and customer experience, including rework. Use the results to choose what to improve next. Applied AI alignment – Build useful agents that respect the impartiality and objectivity required of a certification body. Help people understand the evidence, recognise uncertainty and retain responsibility for consequential judgments.