Team Lead - Site Reliability Engineering (all genders)

BerlinCompetitiveOnsite0 applicants

About this role

Introduction

FACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. We are actively modernizing our hosting toward Kubernetes on Harvester – as an on-prem hybrid with the option to scale fully into the cloud in the mid-term. As Team Lead Site Reliability Engineering (all genders), you own the reliability, scalability, and cost of our hosting environments, drive this transformation end-to-end, and lead the team that delivers it.

Your mission

You own the operational health of our hosting across on-premise (Frankfurt, Stockholm) and cloud – availability, performance, and incident management.

You actively drive the modernization toward Kubernetes on Harvester: cluster topology, storage (Longhorn), networking (VLAN, load balancing, ingress), backup, and disaster recovery.

You build a production-grade k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps (Argo CD / Flux), observability, and policy guardrails.

You shape the NG Search Operator (custom Kubernetes operator) and solve auto-scaling (HPA, VPA, KEDA, cluster autoscaler) for the current architecture.

You concretely define our on-prem hybrid model: which workloads run where, how we burst into the cloud, how we keep latency and cost under control – while keeping the architecture portable enough for a future cloud-only move.

You own capacity planning and hosting cost and turn cost into a deliberate, managed lever.

You lead and develop our currently 4-person Hosting team, own performance and technical direction, and set the standards and ownership culture.

You make AI a core part of our operations: diagnosis, automation, monitoring, and insight.

Your profile

Strong background in infrastructure or platform engineering across on-premise and cloud.

Hands-on depth with Kuberne

Responsibilities

  • You own the operational health of our hosting across on-premise (Frankfurt, Stockholm) and cloud – availability, performance, and incident management.
  • You actively drive the modernization toward Kubernetes on Harvester: cluster topology, storage (Longhorn), networking (VLAN, load balancing, ingress), backup, and disaster recovery.
  • You build a production-grade k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps (Argo CD / Flux), observability, and policy guardrails.
  • You shape the NG Search Operator (custom Kubernetes operator) and solve auto-scaling (HPA, VPA, KEDA, cluster autoscaler) for the current architecture.
  • You concretely define our on-prem hybrid model: which workloads run where, how we burst into the cloud, how we keep latency and cost under control – while keeping the architecture portable enough for a future cloud-only move.

Requirements

  • You own capacity planning and hosting cost and turn cost into a deliberate, managed lever.
  • You lead and develop our currently 4-person Hosting team, own performance and technical direction, and set the standards and ownership culture.
  • You make AI a core part of our operations: diagnosis, automation, monitoring, and insight.
  • Strong background in infrastructure or platform engineering across on-premise and cloud.
  • Hands-on depth with Kuberne

EU Requirements

Job Details

Posted3 September 2026
Closes3 October 2026
Work ModeOnsite

Contact

Similar Jobs

Finding similar jobs...