Skip to main content
Back to jobs

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

Fact Finder

Berlin 9/8/2026 Experienced
Apply on source ↗ Save

<br><strong>Introduction</strong><p><p><strong>At a glance </strong></p><ul><li><span><span>Location &amp; </span><span>work </span><span>model</span><span>: </span></span><span><span>Berlin, hybrid</span></span></li><li><span><span>Tech stack: </span></span><span><span>Kubernetes on our own servers, Harvester (</span><span>KubeVirt</span><span>), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph</span></span></li><li><span><span>Team: </span></span><span><span>A growing SRE team – you report to our CTPO for now and to the Team Lead SRE we're hiring next; two system administrators in Pforzheim run the physical hardware</span></span></li><li><span><span>Process: </span></span><span><span>Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team</span></span></li><li><span><span>Languages: </span></span><span><span>Fluent English required; German is a plus, not a must</span></span></li></ul><p><strong>Why this role is special </strong></p>Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next. <br><br>SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time. <br><br><p><strong>Your first 90 days </strong></p>You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.</p><br><strong>Your mission</strong><p><ul><li>Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions</li><li>Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by case</li><li>Eliminate toil through automation and GitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks</li><li>Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, cluster auto scaler) today's architecture makes hard</li><li>Plan capacity, performance  and cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations</li></ul></p><br><strong>Your profile</strong><p><strong>Must-haves: <br></strong><ul><li><span><span>Kubernetes in production – built, not just used</span></span><span><span>: you've set up and maintained clusters on your own servers (e.g. </span><span>kubeadm</span><span>, RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experience </span><span>isn't </span><span>enough for this role</span></span></li><li><span><span>Lived SRE practice</span></span><span><span>: SLOs, error budgets, incident management, </span><span>on-call</span></span></li><li><span><span>Hands-on experience with </span><span>GitOps</span><span>or comparable infrastructure/deployment automation </span></span><span><span>– experience with Argo CD or Flux is a strong plus</span></span></li><li><span><span>Solid observability skills </span></span><span><span>– metrics, logs, traces, alerting that people trust</span></span></li><li><span><span>A strong automation instinct </span></span><span><span>– you'd rather fix a problem's cause than repeat its workaround</span></span></li><li><span><span>A collaborative, enabling mindset </span></span><span><span>– you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, and don't fall in love with your own solution</span></span></li></ul><strong>Nice-to-haves (genuinely optional – we'll teach you the rest): <br></strong><ul><li><span><span>Harvester, </span><span>KubeVirt</span><span>, vSphere/</span><span>ESXi</span><span>, OpenStack or similar virtualization/HCI platforms</span></span></li><li><span><span>Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)</span></span></li><li><span><span>Auto-scaling (HPA, VPA, KEDA, cluster </span><span>auto scaler</span><span>) and capacity/cost planning</span></span></li><li><span><span>Experience building Kubernetes operators/CRDs</span></span></li><li><span><span>German language skills</span></span></li></ul>Certifications (CKA, CKS) are welcome but no substitute for hands-on experience – in the tech interview we'll ask about what you've actually built and operated. <br><br><strong>You don't tick every box – or your title was never “SRE”?</strong> Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you – production experience and engineering mindset matter more to us than titles or buzzwords.</p><br><strong>THE JOY OF WORKING WITH US</strong><p><ul><li>Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.</li><li>Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.</li><li>AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.</li><li>Ownership &amp; growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.</li><li>Flexible work: Hybrid work model three office days per week with a focus on outcomes.</li><li>Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.</li></ul></p><br><strong>Job Location</strong><p><span style="color:rgb(66,66,66);font-family:Verdana, Geneva, sans-serif, Inter, Arial, Helvetica, sans-serif;font-size:14px;font-style:normal;font-weight:400;text-transform:none;background-color:rgb(255,255,255);display:inline;">Berlin, Munich, Pforzheim or Stockholm (all Hybrid)</span></p><p>Find <a href="https://www.arbeitnow.com">Jobs in Germany</a> on Arbeitnow</a>

join

Related on GOH