About this job
<p><span style="color: #333333">At our company, it’s all about #OneTeam! Join gridscale and help shape the future of the cloud together with OVH.</span></p><p><span style="color: #333333">As a leading tech company, we’ve been working for over two decades to reduce our environmental footprint - with innovative solutions and an open cloud designed to be sustainable from the ground up: #SustainableByDesign.</span></p><p style=";"></p><h4><strong>Our Tech Stack 🚀</strong></h4><p>OpenStack · Kubernetes · KVM · Linux · Bare-metal </p><p>· Ansible · Terraform · FluxCD/ ArgoCD · Git · Go · Python </p><p>· Claude Code/ Cursor/ agentic coding tooling<br /></p><h3><strong>Your Role💻</strong></h3><p>You'll help build, operate, and industrialize OVHcloud's on-premise cloud platform (OPCP). You'll join a small, senior team that owns the OpenStack-based infrastructure and the Kubernetes / GitOps stack our customer-facing cloud runs on and that treats AI-assisted engineering as a first-class part of how we work.</p><p>The platform is actively in build mode, so joining now means real influence on the architecture, the automation strategy, and how we adopt AI in platform engineering. As a Senior, you shape the focus of your role around your strengths and interests: there's a clear backbone of automation, compute-lifecycle, and platform work, plus an explicit AI-substrate workstream. You're at home in a security-oriented, highly automated (GitOps) environment, keep an overview in ambiguous situations, and make well-founded decisions on that basis.</p><p style=";"></p><h3><strong><span style="color: #000000">Your Tasks</span></strong></h3><ul><li><p>Design and build OpenStack-based on-prem infrastructure that deploys itself autonomously - discovering available hardware and bringing up a functional datacenter in minutes.</p></li><li><p>Develop Infrastructure as Code with Ansible and Terraform - typically spec-first with LLM assistance, then human-validated; push this further via custom agent / sub-agent setups, agentic test generation, and prompt-engineered review loops.</p></li><li><p>Drive the ongoing development of our Kubernetes stack and GitOps workflows (FluxCD / ArgoCD).</p></li><li><p>Own the full lifecycle of our compute infrastructure - from bare-metal (firmware, provisioning, hardware health) through hypervisors to virtual compute nodes - and build the automation that keeps capacity healthy and rolls out updates without disturbing tenant workloads.</p></li><li><p>Build and extend the AI substrate that compounds our output: Markdown knowledge bases as retrieval substrate, agentic prototypes for incident triage and capacity planning, and deeper integration of agentic coding tools into daily work.</p></li><li><p>Contribute to the self-healing direction, turning today's manual runbooks into tomorrow's reasoning agents. Auto-remediation isn't a separate team here - it's how platform work is meant to land.</p></li><li><p>Design and implement test suites aligned with functional and technical specs (non-regression, performance, security).</p></li><li><p>Document and package the solution so users can deploy and operate it without friction, and keep improving the platform based on telemetry and user feedback.</p></li><li><p>Act as a technical reference and mentor across automation, platform engineering, and AI-tooling topics.<br /></p></li></ul><h3><strong>What we offer you💼</strong></h3><ul><li><p>A platform that is genuinely in build mode - your architectural decisions stick.</p></li><li><p>A senior team where seniority means autonomy, not just a title.</p></li><li><p>AI-augmented engineering as a first-class workflow -Claude Code and comparable agentic tooling, Markdown-KB-as-substrate, and room to push the practice further. Modern tooling that compounds your work instead of just sitting next to it.</p></li><li><p>Exceptional team spirit across all departments and national borders - we live #OneTeam</p></li><li><p>Exciting work in a highly innovative, international environment with cutting-edge technologies</p></li><li><p>32 vacation days, increasing with length of service</p></li><li><p>Flexible working hours, home-office options, and a secure permanent position with market- and performance-based compensation</p></li><li><p>Employer-funded pension plan and an attractive insurance package</p></li><li><p>OVHcloud covers 50% of public transportation costs</p></li><li><p>Up to €400 per year toward sports activities (gym membership, classes, etc.)</p></li><li><p>Attractive discounts at numerous shops and companies through Corporate Benefits</p></li><li><p>A contribution toward leasing your cargo bike</p></li><li><p>Regular company events and free cold and hot beverages</p></li></ul><ul><li><p>Several years of hands-on experience running production infrastructure (SRE, Platform, or DevOps).</p></li><li><p>Solid OpenStack experience - deployed, operated, and debugged it in production.</p></li><li><p>End-to-end compute infrastructure management, from bare-metal lifecycle through hypervisor and virtual compute node operations (migration, host evacuation, graceful drains, capacity rebalancing). The skill matters more than the specific tooling - what counts is having done it at scale and automated it.</p></li><li><p>Strong with Infrastructure as Code (Ansible, Terraform) and GitOps (FluxCD or ArgoCD), plus solid Linux administration including on bare-metal.</p></li><li><p>Active, daily practice of AI-assisted engineering, with opinions formed from real use. You can describe a workflow where an LLM saved you half a day, and one where you should have skipped it. Theoretical interest doesn't count.</p></li><li><p>Fluent English, written and spoken - our team is distributed, and this is the working language.</p><p style=";"></p></li><li><p><strong>Nice to Have</strong></p><ul><li><p>Production experience with Kubernetes and the cloud-native ecosystem.</p></li><li><p>Production-quality Go and/or Python.</p></li><li><p>Deeper agentic tooling craft (Claude Code, Cursor, Aider): custom agent / sub-agent setups, hooks, prompt engineering, your own workflows or skills and managing a Markdown-first knowledge base as substrate for AI workflows.</p></li><li><p>Advanced compute-node tuning (CPU pinning, NUMA, hugepages, SR-IOV / PCI passthrough) and basic network debugging (VLANs, BGP).</p></li><li><p>Observability tooling (Prometheus, Loki, Grafana, etc.) and auto-remediation / self-healing systems (StackStorm, Event-Driven Ansible, or similar).</p></li><li><p>Experience in security-critical environments and with edge or multi-site deployments.</p><p style=";"></p></li></ul></li><li><p><strong>Soft Skills</strong></p><ul><li><p>A continuous-improvement mindset and ownership for what you build.</p></li><li><p>You see AI tooling as a structural shift in how engineering gets done - not a trend, not a threat and want to shape how the team adopts it.</p></li><li><p>You enjoy sharing knowledge, learning from peers, and can synthesize ideas clearly.</p></li></ul></li></ul><p>Find <a href="https://www.arbeitnow.com/">Jobs in Germany</a> on Arbeitnow</a>