Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

Fact FinderBerlinArbeitnow٨‏/٩‏/٢٠٢٦
إعلان
IntroductionAt a glance  Location & work model: Berlin, hybridTech stack: Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/CephTeam: A growing SRE team – you report to our CTPO for now and to the Team Lead SRE we're hiring next; two system administrators in Pforzheim run the physical hardwareProcess: Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the teamLanguages: Fluent English required; German is a plus, not a mustWhy this role is special  Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next.  SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.  Your first 90 days  You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next. Your missionDefine and own SLOs, SLIs and error budgets; drive data-informed reliability decisionsLead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by caseEliminate toil through automation and GitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacksHelp build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, cluster auto scaler) today's architecture makes hardPlan capacity, performance  and cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations Your profileMust-haves:  Kubernetes in production – built, not just used: you've set up and maintained clusters on your own servers (e.g. kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experience isn't enough for this roleLived SRE practice: SLOs, error budgets, incident management, on-callHands-on experience with GitOpsor comparable infrastructure/deployment automation – experience with Argo CD or Flux is a strong plusSolid observability skills – metrics, logs, traces, alerting that people trustA strong automation instinct – you'd rather fix a problem's cause than repeat its workaroundA collaborative, enabling mindset – you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, and don't fall in love with your own solutionNice-to-haves (genuinely optional – we'll teach you the rest):  Harvester, KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platformsContainer storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)Auto-scaling (HPA, VPA, KEDA, cluster auto scaler) and capacity/cost planningExperience building Kubernetes operators/CRDsGerman language skillsCertifications (CKA, CKS) are welcome but no substitute for hands-on experience – in the tech interview we'll ask about what you've actually built and operate
إعلان