FactFinder

Senior Site Reliability Engineer (all genders)

Remote, United States remote Entry Salary not listed
remote Technology & IT Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

Introduction

FACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. Both products are moving toward a modern, hybrid platform based on Kubernetes and Harvester – with the option to scale fully into the cloud in the mid-term. As a Senior Site Reliability Engineer (SRE), you make sure our systems stay fast, available, and scalable throughout this transformation. You work closely with the Hosting team and experienced engineers, and actively shape our journey toward a modern SaaS company.

Your mission

You define and own SLOs, SLIs, and error budgets across both products and make data-driven decisions on reliability and performance.

You drive incident response: fast detection, clear communication, blameless postmortems, and meaningful follow-through.

You consistently reduce manual work through automation and GitOps (e.g. Argo CD / Flux) and build out self-healing and self-service capabilities.

You support the development of an NG Search Operator (custom Kubernetes operator / CRDs) and the rollout of auto-scaling (HPA, VPA, KEDA, cluster autoscaler).

You evolve our observability – metrics, logs, traces, alerting, and runbooks that actually help on call.

You plan capacity and cost across on-premise (Frankfurt, Stockholm) and cloud – including burst scenarios into the public cloud.

You leverage AI tools to noticeably accelerate diagnosis, alerting, and operational workflows.

Your profile

Experience as an SRE, infrastructure, or production engineer in a SaaS or platform environment – or a strong software/operations background with a clear drive to grow into an SRE role.

Solid understanding of SLOs, error budgets, incident management, and observability.

Hands-on experience with Kubernetes and interest in cluster lifecycle, upgrades, and operator patterns.

Experience with or strong interest in Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).

Familiarity with GitOps (Argo CD / Flux), container storage (Longhorn, Ceph), and Kubernetes networking (load balancing, ingress).

Knowledge of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and capacity planning on-prem and in the cloud.

Understanding of networking in production-grade datacenters (incl. VLAN).

A strong automation instinct and a mindset to structurally eliminate toil.

Practical experience using AI tools in day-to-day operations.

Fluent English; German is a plus.

THE JOY OF WORKING WITH US

Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.

Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.

AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.

Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.

Flexible work: Hybrid work model with a focus on outcomes.

Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.

Job Location

Berlin, Munich, Pforzheim or Stockholm (Hybrid)

Find Jobs in Germany on Arbeitnow

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Role
Technology & IT Operations Senior Site Reliability Engineer remote

Likely questions

  1. Tell us about work you have done that is close to the Senior Site Reliability Engineer (all genders) role.
  2. How would you approach your first 30 days at FactFinder?
  3. Which of Operations, Senior and Site have you used recently, and what did it help you achieve?
  4. Describe a time you solved a problem without waiting to be told exactly what to do.
  5. How do you stay organised and communicate clearly when working remotely?

Prepare before the call

  • A recent example that proves your experience with Operations, Senior and Site.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the Senior Site Reliability Engineer (all genders) role because I can bring practical experience in Operations, Senior and Site, learn the team quickly, and contribute to the outcomes FactFinder needs from this hire.

Related jobs.

More roles from this company or category.