About this job
<p style="min-height:1.5em">We're looking for a <strong>Multimodal ML Engineer</strong> to join <em>White Circle</em>, an AI Safety company building the policy enforcement and optimization layer for AI systems. Backed by $11M from senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, and DeepMind, White Circle processes 100M+ API calls monthly and runs its own LLMs in production.</p><p style="min-height:1.5em"><strong>You will</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Train and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints.</p></li><li><p style="min-height:1.5em">Design experiments, build multimodal data pipelines, and train MoE architectures.</p></li><li><p style="min-height:1.5em">Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end.</p></li><li><p style="min-height:1.5em">Define evaluation metrics that actually matter for the product.</p></li></ul><p style="min-height:1.5em"><strong>Requirements</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">3+ years training large-scale multimodal models.</p></li><li><p style="min-height:1.5em">Strong PyTorch and distributed training experience (DeepSpeed, FSDP).</p></li><li><p style="min-height:1.5em">Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.</p></li><li><p style="min-height:1.5em">Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).</p></li><li><p style="min-height:1.5em">Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.</p></li><li><p style="min-height:1.5em">Relocation to Paris or London (hybrid) required.</p></li></ul><p style="min-height:1.5em"><strong>Bonus</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Audio signal processing fundamentals – spectrograms, mel features, noise reduction.</p></li><li><p style="min-height:1.5em">MoE architecture experience.</p></li></ul><p style="min-height:1.5em"><strong>We offer</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">$100k–$250k/year salary + equity; higher figures can be negotiated.</p></li><li><p style="min-height:1.5em">Official employment, visa and relocation help.</p></li></ul><hr><p>Compensation: $100K – $250K • Higher figures and equity are negotiable</p><ul><li> • $100K – $250K • Higher figures and equity are negotiable</li></ul><p>Find more <a href="https://www.arbeitnow.fr/english-speaking-jobs">English Speaking Jobs in France</a> on Arbeitnow</a>