About this job
<h3>About the Role</h3><p style="min-height:1.5em">This company is building an AI executive assistant that operates across email, calendars, meetings, and business software. As a <strong>Data Scientist — Agent Evaluations & Quality</strong>, you will own the measurement system that determines whether the assistant is genuinely improving in ambiguous, real-world environments. You'll partner directly with AI Agent Capabilities engineers to generate the evidence that shapes product decisions, model choices, and release quality.</p><p style="min-height:1.5em">This is a high-ownership, deeply technical role at the intersection of applied data science, LLM evaluation, and product quality — ideal for someone who thrives on turning hard, open-ended quality questions into rigorous, actionable answers.</p><h3>What You'll Do</h3><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces.</p></li><li><p style="min-height:1.5em">Translate agent capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks.</p></li><li><p style="min-height:1.5em">Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios.</p></li><li><p style="min-height:1.5em">Define meaningful metrics — task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.</p></li><li><p style="min-height:1.5em">Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and track grader agreement.</p></li><li><p style="min-height:1.5em">Compare models, prompts, and implementations using rigorous offline experiments and production evidence.</p></li><li><p style="min-height:1.5em">Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy.</p></li><li><p style="min-height:1.5em">Turn production failures into regression cases and continuously close gaps in evaluation coverage.</p></li><li><p style="min-height:1.5em">Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership.</p></li><li><p style="min-height:1.5em">Recommend improvements to capability engineers and verify that fixes raise quality without unacceptable regressions.</p></li></ul><h3>What We're Looking For</h3><p style="min-height:1.5em"><strong>Required</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines, or production ML infrastructure.</p></li><li><p style="min-height:1.5em">Experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems.</p></li><li><p style="min-height:1.5em">Production-grade proficiency in <strong>Python and SQL</strong>, with experience building and maintaining automated analytical pipelines on large datasets.</p></li><li><p style="min-height:1.5em">Applied statistical and experimental skills: significance testing, variance analysis, and sampling to evaluate non-deterministic AI/ML systems.</p></li><li><p style="min-height:1.5em">Experience developing labeled datasets, annotation guidelines, and quality-control processes for ground-truth data in dynamic product environments.</p></li><li><p style="min-height:1.5em">Solid understanding of LLM agent behaviors: tool use, multi-step execution, retrieval, and practical failure modes.</p></li><li><p style="min-height:1.5em">Demonstrated ability to analyze model traces, tool calls, and outputs to identify root causes across model, prompt, tool, and data layers.</p></li><li><p style="min-height:1.5em">Experience using production telemetry and observability data to monitor system quality, build dashboards, and analyze real-world user outcomes.</p></li></ul><p style="min-height:1.5em"><strong>Nice to Have</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Hands-on experience with LLM-as-a-judge systems, model-based grading, or AI benchmarking platforms.</p></li><li><p style="min-height:1.5em">Experience shipping or operating production ML products, agentic systems, or customer-facing consumer software.</p></li><li><p style="min-height:1.5em">Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems.</p></li></ul><p style="min-height:1.5em"><strong>What makes you a great fit</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">You're product-oriented — you prioritize metrics tied to real user outcomes, not just convenient measurements.</p></li><li><p style="min-height:1.5em">You drive ambiguous quality questions from evaluation design all the way into product decisions.</p></li><li><p style="min-height:1.5em">You write maintainable, production-quality code — not just ad-hoc notebooks.</p></li><li><p style="min-height:1.5em">You collaborate naturally with engineers and are comfortable digging into traces and system internals.</p></li></ul><h3>Location</h3><p style="min-height:1.5em">This role is <strong>on-site</strong>. Visa sponsorship is <strong>not available</strong> for this position.</p><h3>Compensation & Benefits</h3><p style="min-height:1.5em">Compensation details were not provided for this listing. A competitive package commensurate with experience is expected at this stage of company growth.</p><p>Find <a href="https://www.arbeitnow.co.uk">Jobs in United Kingdom</a> on Arbeitnow</a>