About this job
<p>For our growing team, we are looking for an experienced Sovereign Cloud / Site Reliability Engineer (SRE) to support the secure, reliable, and continuous operation of modern cloud-native and Kubernetes platforms.</p>
<p>You will work with an experienced DevOps/SRE team on highly secure, business-critical platforms, taking responsibility for platform operations, automation, monitoring, incident management, security, and continuous improvement.</p>
<p>The role focuses particularly on Kubernetes, CI/CD, Infrastructure as Code, observability, logging, and operational excellence within a highly available Unified Observability Platform in a 24x7 operational environment.</p>
<p><strong>Mandatory Requirements – Please Read Before Applying</strong></p>
<p><strong>Citizenship</strong></p>
<p>You must hold valid citizenship in a country that is a full member of both the <strong>European Union (EU) and NATO</strong>.</p>
<p>If you hold multiple citizenships, <strong>all citizenships must be from countries that are members of both the EU and NATO.</strong></p>
<p><strong>Employment in Germany</strong></p>
<p>You must:</p>
<ul>
<li>
<p>Be employed directly by a legal entity registered in Germany.</p>
</li>
<li>
<p>Hold a German employment contract.</p>
</li>
<li>
<p>Be subject exclusively to German labor law.</p>
</li>
<li>
<p>Comply with all applicable German tax, social-security, and employment regulations.</p>
</li>
<li>
<p>Reside in Germany.</p>
</li>
</ul>
<p>**Employment through non-German entities, including foreign subcontractors or affiliates, does not meet this requirement.<br>
**</p>
<p><strong>Security Clearance – Ü2</strong></p>
<p>You must hold a <strong>valid and verifiable Ü2 security clearance</strong> in accordance with the German Security Clearance Act (<em>Sicherheitsüberprüfungsgesetz – SÜG</em>) and applicable preventive personnel sabotage-protection requirements.</p>
<h2>Tasks</h2>
<p>Kubernetes & Platform Operations</p>
<ul>
<li>Operate and maintain Kubernetes clusters with Gardener, including workloads, deployments, Helm charts and platform components.</li>
<li>Troubleshoot availability, performance and deployment issues.</li>
<li>Ensure secure, scalable, resilient and highly available platforms.</li>
</ul>
<p>CI/CD & Automation</p>
<ul>
<li>Operate and optimize Jenkins and ArgoCD pipelines for automated deployments.</li>
<li>Implement Infrastructure as Code (IaC) and Git-based deployment workflows.</li>
<li>Develop automation and operational tools using Python, Go and/or Bash.</li>
<li>Automate provisioning, health and compliance checks, alerting and reporting.</li>
</ul>
<p>Monitoring & Observability</p>
<ul>
<li>Manage Prometheus, Thanos and OpenTelemetry environments, including scrape jobs, alert rules and PromQL.</li>
<li>Develop and maintain Grafana dashboards.</li>
<li>Continuously improve monitoring and alerting capabilities.</li>
</ul>
<p>Logging & Log Management</p>
<ul>
<li>Operate and optimize Elasticsearch/OpenSearch, Logstash and Kibana.</li>
<li>Monitor log ingestion, storage, performance and reliability.</li>
<li>Support centralized troubleshooting, anomaly detection, security and compliance.</li>
</ul>
<p>Integration & Operations</p>
<ul>
<li>Integrate observability and logging platforms with ServiceNow, PagerDuty and other enterprise tools via secure APIs.</li>
<li>Handle operational requests.</li>
<li>Participate in Scrum, DevOps and service-improvement activities.</li>
<li>Collaborate with internal teams, SAP, suppliers and stakeholders.</li>
</ul>
<p>Incident & Problem Management</p>
<ul>
<li>Participate in a 24/7 on-call and shift rotation, including weekends and public holidays.</li>
<li>Respond to platform, monitoring, logging and deployment incidents.</li>
<li>Perform Root Cause Analysis (RCA).</li>
<li>Support Major Incident Management (MIM).</li>
<li>Implement sustainable corrective actions.</li>
<li>Continuously improve platform stability, resilience and operational processes.</li>
</ul>
<h2>Requirements</h2>
<p>Key Technologies</p>
<p>Kubernetes | Gardener | Helm | Jenkins | ArgoCD | Git | IaC | Python | Go | Bash | Prometheus | PromQL | Thanos | OpenTelemetry | Grafana | Elasticsearch | OpenSearch | Logstash | Kibana | ServiceNow | PagerDuty | REST APIs</p>
<p>Technical Skills & Experience</p>
<ul>
<li>Proven experience as a <strong>Sovereign Cloud Engineer and/or Site Reliability Engineer (SRE)</strong>.</li>
<li>Strong experience operating and maintaining <strong>Kubernetes clusters</strong>, ideally with Gardener.</li>
<li>Hands-on experience with <strong>Kubernetes workloads, deployments, Helm charts and platform components</strong>.</li>
<li>Experience troubleshooting <strong>availability, performance and deployment issues</strong>.</li>
<li>Experience with <strong>Jenkins and ArgoCD</strong> for automated CI/CD deployments.</li>
<li>Solid understanding of <strong>Infrastructure as Code (IaC)</strong> and Git-based deployment workflows.</li>
<li>Programming or scripting experience with <strong>Python, Go and/or Bash</strong>.</li>
<li>Experience with <strong>automation of provisioning, health checks, compliance checks, alerting and reporting</strong>.</li>
<li>Hands-on experience with <strong>Prometheus, PromQL, Thanos and OpenTelemetry</strong>.</li>
<li>Experience developing and maintaining <strong>Grafana dashboards</strong> and monitoring/alerting solutions.</li>
<li>Experience with <strong>Elasticsearch/OpenSearch, Logstash and Kibana</strong>.</li>
<li>Understanding of <strong>log ingestion, storage, performance and reliability</strong>.</li>
<li>Experience supporting <strong>centralized troubleshooting, anomaly detection, security and compliance</strong>.</li>
<li>Experience integrating platforms with <strong>ServiceNow, PagerDuty and other enterprise tools via secure REST APIs</strong>.</li>
<li>Experience with <strong>incident management, Root Cause Analysis (RCA) and Major Incident Management (MIM)</strong>.</li>
<li>Experience working in <strong>24/7 operational environments</strong>, including on-call and shift rotations.</li>
<li>Strong understanding of <strong>platform security, scalability, resilience, reliability and high availability</strong>.</li>
<li>Experience working in <strong>Scrum and DevOps environments</strong>.</li>
<li>Strong collaboration skills and experience working with <strong>internal teams, SAP, suppliers and other stakeholders</strong>.</li>
</ul>
<p>Languages</p>
<ul>
<li><strong>English:</strong> Required</li>
<li><strong>German:</strong> A plus</li>
</ul>
<h2>Benefits</h2>
<p>What You Can Expect</p>
<ul>
<li>A technically challenging role within a <strong>highly secure and business-critical cloud environment</strong>.</li>
<li>The opportunity to work with modern <strong>cloud-native, Kubernetes and observability technologies</strong>.</li>
<li>An <strong>international working environment</strong> with teams and stakeholders across different locations.</li>
<li>The opportunity to contribute to <strong>automation, platform reliability and continuous improvement</strong>.</li>
<li>Remote / hybrid working possibilities within Germany.</li>
</ul>
<p>We Look Forward to Hearing from You</p>
<p>Are you ready to bring your expertise to a challenging cloud engineering environment?</p>
<p>We look forward to receiving your application and getting to know you.</p>
<p>Find more <a href="https://www.arbeitnow.com/english-speaking-jobs">English Speaking Jobs in Germany</a> on Arbeitnow</a>