Remote
Senior AI Evaluation Engineer
About this role
Joining a dynamic cross-functional development team, the full-time Senior AI Evaluation Engineer will build evaluation harnesses, create golden datasets, and establish production monitoring for AI systems in a remote work environment. Key responsibilities Build the central evaluation harness and templates for consistent testing across teams Create golden datasets with input from business and clinical reviewers, ensuring safety and medical accuracy Establish production monitoring for drift and regression, auditing evaluations and reporting metrics monthly Required qualifications 4+ years of experience in ML/LLM evaluation, QA engineering for AI systems, or applied research engineering Hands-on experience with evaluation frameworks and statistical rigor on small samples Ability to set and defend pass thresholds independently of the development team Experience with healthcare or safety-critical evaluation is desirable Background in red-teaming is a plus
Source listing: virtualvocations_main