About this job
<p style="min-height:1.5em"><strong>TLDR: Multimodal ML Engineer to train and ship vision, audio, video, and speech models for an AI safety platform that operates at 100M+ API calls/month.</strong></p><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>About us</strong></p><p style="min-height:1.5em"><a target="_blank" rel="noopener noreferrer nofollow" href="https://whitecircle.ai/"><u>White Circle</u></a> is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do. We automatically test, enforce, and continuously improve these policies at scale.</p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others</p></li><li><p style="min-height:1.5em">We process over 100M+ API calls every month</p></li><li><p style="min-height:1.5em">We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model</p></li></ul><p style="min-height:1.5em">We’re a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built – you’re the one we need.</p><p style="min-height:1.5em"><strong>You will</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Train and fine-tune large-scale multimodal models (vision-language, audio, speech) from scratch and from pretrained checkpoints</p></li><li><p style="min-height:1.5em">Extend models across modalities: image understanding, video temporal modeling, long-context processing, and streaming audio</p></li><li><p style="min-height:1.5em">Design and run experiments: architecture changes, data mixes, training recipes</p></li><li><p style="min-height:1.5em">Build and maintain multimodal data pipelines — from raw images, video, and audio recordings to training-ready datasets, including synthetic data generation</p></li><li><p style="min-height:1.5em">Train and optimize MoE architectures for efficient multimodal inference</p></li><li><p style="min-height:1.5em">Build alignment pipelines: SFT, DPO, GRPO, reward modeling — across modalities, not just text</p></li><li><p style="min-height:1.5em">Optimize models for production: quantization, distillation, batching, streaming and low-latency serving</p></li><li><p style="min-height:1.5em">Deploy models end-to-end: from research checkpoint to production serving</p></li><li><p style="min-height:1.5em">Define evaluation metrics and benchmarks that actually matter for the product: visual QA, spatial reasoning, video comprehension, speech and audio understanding</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>You’ll fit right in if you</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">3+ years training large-scale deep learning models in multimodal domains (vision-language, audio, speech, or acoustic)</p></li><li><p style="min-height:1.5em">Strong PyTorch skills with hands-on distributed training experience (DeepSpeed, FSDP, or similar)</p></li><li><p style="min-height:1.5em">Deep experience with multimodal architectures — you understand how vision/audio encoders, projectors, and LLMs fit together (LLaVA, Qwen-VL, InternVL, Audio Flamingo, Omni Qwen, Audio Qwen, Whisper, HuBERT, Conformer, or similar)</p></li><li><p style="min-height:1.5em">Hands-on with RLHF/alignment for multimodal: GRPO, DPO, reward modeling — not just for text</p></li><li><p style="min-height:1.5em">Experience with video and/or audio sequence modeling: temporal modeling, long-context processing, efficient attention, streaming inference</p></li><li><p style="min-height:1.5em">Track record of shipping models to production: you've hit latency targets and optimized inference, not just reported benchmark scores</p></li><li><p style="min-height:1.5em">Comfortable with large-scale multimodal dataset curation: image-text pairs, video-instruction data, audio preprocessing, augmentation, synthetic data generation</p></li><li><p style="min-height:1.5em">Familiar with MoE architectures and their tradeoffs for multimodal workloads</p></li><li><p style="min-height:1.5em">Strong engineering fundamentals: clean code, version control, testing, documentation</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>A big plus:</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Understanding of audio signal processing fundamentals (spectrograms, mel features, noise reduction)</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>Why White Circle</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Paid time off in line with your local regulations, no matter where you work from</p></li><li><p style="min-height:1.5em">Work from Paris (hybrid) with a relocation package available, or work from London (note: we are unable to provide relocation support for London-based roles)</p></li><li><p style="min-height:1.5em">Comprehensive medical insurance for our France-based team (please note that we are in the process of setting up our UK office and therefore cannot offer medical insurance for London-based roles yet)</p></li><li><p style="min-height:1.5em">Meaningful equity package</p></li><li><p style="min-height:1.5em">All the hardware, tools, and services you need</p></li><li><p style="min-height:1.5em">Covered subscriptions for AI agents and IDEs</p></li><li><p style="min-height:1.5em">Team off-sites twice a year: we’ve recently been to the Alps and to Saint-Tropez</p></li></ul><div style="min-height:1.2em;margin-top:0;margin-bottom:0"> </div><p style="min-height:1.5em"><strong>How we hire</strong></p><ol style="min-height:1.5em"><li><p style="min-height:1.5em">Introductory call with HR (25 min)</p></li><li><p style="min-height:1.5em">Take-home test task</p></li><li><p style="min-height:1.5em">Technical interview with Head of Applied Research (60 min)</p></li><li><p style="min-height:1.5em">Final conversation with our CEO (45 min)</p></li></ol><p>Find <a href="https://www.arbeitnow.fr">Jobs in France</a> on Arbeitnow</a>