About this job
<p style="min-height:1.5em"><strong>TLDR: We are looking for several ML Engineers to train, post-train, and evaluate the LLMs at the core of our platform. This is hands-on modern model training work: large-scale data pipelines, SFT/RLHF/DPO-style alignment, reward models, distributed multi-GPU training, and evaluation.</strong></p><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>About us</strong></p><p style="min-height:1.5em"><a target="_blank" rel="noopener noreferrer nofollow" href="https://whitecircle.ai/"><u>White Circle</u></a> is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do. We automatically test, enforce, and continuously improve these policies at scale.</p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others</p></li><li><p style="min-height:1.5em">We process over 100M+ API calls every month</p></li><li><p style="min-height:1.5em">We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model</p></li></ul><p style="min-height:1.5em">We’re a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built – you’re the one we need.</p><div style="min-height:1.2em;margin-top:0;margin-bottom:0"> </div><p style="min-height:1.5em"><strong>You will:</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Turn petabytes of unstructured text into a structured, explorable view (topics, clusters, segments, trends, anomalies): iterate from “unknown unknowns” to stable definitions we can track.</p></li><li><p style="min-height:1.5em">Build scalable representation pipelines: sampling strategies, preprocessing/normalization, embeddings at scale, indexing, and retrieval to make the corpus searchable and analyzable.</p></li><li><p style="min-height:1.5em">Use LLMs pragmatically: labeling/classification, weak supervision, data enrichment, summarization, and automated diagnostics of inbound volumes (with cost/quality controls).</p></li><li><p style="min-height:1.5em">Deliver insights that change decisions: translate findings into product and operational actions (what data we have, what’s missing, where quality breaks, what to prioritize next).</p></li><li><p style="min-height:1.5em">Ship self-serve analytics: datasets, data models, and lightweight tools/dashboards so the team can explore and answer questions without ad-hoc requests.</p></li><li><p style="min-height:1.5em">Partner closely with engineering/research: align pipelines with production constraints (latency/cost/privacy), and integrate outputs into workflows.</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>You'll fit right in if you:</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Strong Python + SQL with an engineering mindset: you can build reliable pipelines, not just notebooks.</p></li><li><p style="min-height:1.5em">Solid applied NLP/ML experience on real-world text: embeddings, clustering, topic modeling, semantic search, classification; you understand failure modes and how to debug them.</p></li><li><p style="min-height:1.5em">Comfortable at scale: distributed processing, large-scale storage-querying, and performance-cost tradeoffs.</p></li><li><p style="min-height:1.5em">You know how to evaluate fuzzy problems: offline/online metrics, human-in-the-loop labelling, inter-annotator agreement, drift monitoring, and reproducibility.</p></li><li><p style="min-height:1.5em">Have prior work with safety/moderation datasets, policy/rule systems, or high-volume logging/observability</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>A big plus:</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">A public builder footprint: open-source models, datasets, or training frameworks on HuggingFace/GitHub, benchmarks, papers (workshop or main conference), or technical posts with real usage</p></li><li><p style="min-height:1.5em">Experience training models at a frontier or near-frontier lab, or leading open-source model releases with documented adoption</p></li><li><p style="min-height:1.5em">Experience with RL methods for LLMs beyond standard RLHF: online RL, GRPO-style methods, or novel alignment approaches</p></li><li><p style="min-height:1.5em">Experience with moderation, safety, or classification models at scale</p></li><li><p style="min-height:1.5em">Multilingual model training experience</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>Why White Circle</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Paid time off in line with your local regulations, no matter where you work from</p></li><li><p style="min-height:1.5em">Work from Paris (hybrid) with a relocation package available, or work from London (note: we are currently unable to provide relocation support and medical insurance for London-based roles)</p></li><li><p style="min-height:1.5em">Comprehensive medical insurance for our France-based team</p></li><li><p style="min-height:1.5em">All the hardware, tools, and services you need</p></li><li><p style="min-height:1.5em">Meaningful equity package</p></li><li><p style="min-height:1.5em">Covered subscriptions for AI agents and IDEs</p></li><li><p style="min-height:1.5em">Team off-sites twice a year: we've recently been to the Alps and to Saint-Tropez</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>How we hire</strong></p><ol style="min-height:1.5em"><li><p style="min-height:1.5em">Introductory call with HR (25 min)</p></li><li><p style="min-height:1.5em">Take-home test task</p></li><li><p style="min-height:1.5em">Technical interview with Head of Applied Research (60 min)</p></li><li><p style="min-height:1.5em">Final conversation with our CEO (45 min)</p></li></ol><p style="min-height:1.5em">Please submit your application in English.</p><p>Find more <a href="https://www.arbeitnow.fr/english-speaking-jobs">English Speaking Jobs in France</a> on Arbeitnow</a>