About this job
<p><span style="color: #121317"><span style="">At TechBiz Global, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an </span></span><strong><span style="color: #121317">Senior AI DevOps / LLMOps</span></strong><span style="color: #616161"> </span><span style="color: #121317">specialist<span style=""> to join one of our </span></span><strong><span style="color: #121317"><span style="">clients</span></span></strong><span style="color: #121317"><span style="">' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you.<br /></span></span></p><p><strong><span style="color: #121317">Key Responsibilities</span></strong></p><ol><li><p><span style="color: #121317">Automation of Build-to-Production</span></p></li></ol><p><span style="color: #121317">- Design and implement robust CI/CD pipelines tailored for AI, covering model weights,</span></p><p><span style="color: #121317">dataset versioning, and application code.</span></p><p><span style="color: #121317">- Develop specialized workflows for PromptOps, ensuring that system prompts are</span></p><p><span style="color: #121317">version-controlled, tested for regressions, and deployed with the same rigor as traditional</span></p><p><span style="color: #121317">code.</span></p><p><span style="color: #121317">-Automate the deployment of Agentic workflows, managing the complexities of stateful</span></p><p><span style="color: #121317">AI interactions and multi-agent handoffs.</span></p><p><span style="color: #121317">2. AI Infrastructure as Code (IaC)</span></p><p><span style="color: #121317">- Provision and manage high-performance compute environments (GPU clusters, TPU</span></p><p><span style="color: #121317">pods) using Terraform, Pulumi, or Ansible.</span></p><p><span style="color: #121317">- Define and enforce Policy-as-Code for AI endpoints to ensure compliance with security,</span></p><p><span style="color: #121317">cost-usage limits, and data residency requirements.</span></p><p><span style="color: #121317">- Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless</span></p><p><span style="color: #121317">parity between On-Premises development and Cloud production.</span></p><p><span style="color: #121317">3. Safe Experimentation & Controlled Releases</span></p><p><span style="color: #121317">- Architect Progressive Delivery strategies for AI, including Canary releases, Blue-Green</span></p><p><span style="color: #121317">deployments, and Shadowing (where new models run in parallel with production to</span></p><p><span style="color: #121317">compare outputs).</span></p><p><span style="color: #121317">- Build “Evaluation-in-the-Loop” gates within the pipeline to automatically test for bias,</span></p><p><span style="color: #121317">hallucination, and performance degradation before a release.</span></p><p><span style="color: #121317">- Implement A/B testing frameworks specifically designed for LLM outputs and agentic</span></p><p><span style="color: #121317">behavior.</span></p><p><span style="color: #121317">4. Monitoring &amp; Observability</span></p><p><span style="color: #121317">- Establish deep observability into Inference Endpoints, tracking metrics like tokens-per-</span></p><p><span style="color: #121317">second, latency, and drift in model accuracy.</span></p><p><span style="color: #121317">-Integrate feedback loops that capture production “edge cases” to feed back into the</span></p><p><span style="color: #121317">training and fine-tuning pipelines.</span></p><br /><br /><p><strong><span style="color: #121317">Must-Have Technical Skills:</span></strong></p><p><span style="color: #121317">-Orchestration: Advanced Kubernetes (K8s) skills, specifically with KubeFlow, Ray, or</span></p><p><span style="color: #121317">NVIDIA Triton.</span></p><p><span style="color: #121317">-CI/CD &amp; IaC: Expertise in GitHub Actions/GitLab CI, and Terraform or Pulumi.</span></p><p><span style="color: #121317">- AI Tooling: Experience with Weights &amp; Biases, MLflow, LangSmith, or Arize</span></p><p><span style="color: #121317">Phoenix.</span></p><p><span style="color: #121317">-Hardware: Understanding of GPU virtualization, CUDA drivers, and on-premises</span></p><p><span style="color: #121317">hardware management.<br />-Security: Familiarity with Open Policy Agent (OPA) and secret management (Vault).<br /></span></p><p><strong><span style="color: #121317">Experience:</span></strong></p><p><span style="color: #121317">- 10+ years in DevOps, SRE, or Cloud Engineering.</span></p><p><span style="color: #121317">- 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs</span></p><p><span style="color: #121317">from notebook to production.</span></p><p><span style="color: #121317">-Proven experience managing Hybrid Cloud environments (e.g., AWS/Azure + Private</span></p><p><span style="color: #121317">Data Center).</span></p><p>Find <a href="https://www.arbeitnow.com/">Jobs in Germany</a> on Arbeitnow</a>