FactFinder

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

Berlin, Germany full-time Mid Salary not listed
full-time Mid level Technology & IT Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

Introduction

At a glance

Location &workmodel:Berlin, hybrid

Tech stack:Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph

Team:A growing SRE team – you report to our CTPO for now and to the Team Lead SREwe'rehiring next; two system administrators in Pforzheim run the physical hardware

Process:Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team

Languages:Fluent Englishrequired; German is a plus, nota must

Why this role is special

Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next.

SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.

Your first 90 days

You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.

Your mission

Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions

Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by case

Eliminatetoil through automation andGitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks

Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, clusterautoscaler) today's architecture makes hard

Plan capacity,performanceand cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations

Your profile

Must-haves:

Kubernetes in production – built, not just used:you'veset up andmaintainedclusters on your own servers (e.g.kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experienceisn'tenough for this role

Lived SRE practice: SLOs, error budgets, incident management,on-call

Hands-on experience withGitOpsor comparable infrastructure/deployment automation– experience with Argo CD or Flux is a strong plus

Solid observability skills– metrics, logs, traces, alerting that people trust

A strong automation instinct–you'drather fix a problem's cause than repeat its workaround

A collaborative, enabling mindset– you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, anddon'tfall in love with your own solution

Nice-to-haves (genuinely optional – we'll teach you the rest):

Harvester,KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platforms

Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)

Auto-scaling (HPA, VPA, KEDA, clusterautoscaler) and capacity/cost planning

Experience building Kubernetes operators/CRDs

German language skills
Certifications (CKA, CKS) are welcome but no substitute for hands-on experience – in the tech interview we'll ask about what you've actually built and operated.

You don't tick every box – or your title was never “SRE”? Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you – production experience and engineering mindset matter more to us than titles or buzzwords.

THE JOY OF WORKING WITH US

Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.

Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.

AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.

Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.

Flexible work: Hybrid work model three office days per week with a focus on outcomes.

Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.

Job Location

Berlin, Munich, Pforzheim or Stockholm (all Hybrid)

Find Jobs in Germany on Arbeitnow

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Mid
Technology & IT Operations SQL Senior Site Reliability Mid level

Likely questions

  1. Tell us about work you have done that is close to the Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d) role.
  2. How would you approach your first 30 days at FactFinder?
  3. Which of Operations, SQL and Senior have you used recently, and what did it help you achieve?
  4. Describe a time you solved a problem without waiting to be told exactly what to do.
  5. How do you handle busy days, changing priorities, or pressure at work?

Prepare before the call

  • A recent example that proves your experience with Operations, SQL and Senior.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d) role because I can bring practical experience in Operations, SQL and Senior, learn the team quickly, and contribute to the outcomes FactFinder needs from this hire.

Related jobs.

More roles from this company or category.