Team Lead - Site Reliability Engineering (all genders)
Fact Finder · Berlin, Germany
On-siteabout 10 hours agoApply →Introduction FACT-Finder entwickelt Product-Discovery-Technologie für den eCommerce und ist mit den Produkten Next Generation und Infinity bei führenden Online-Shops in Europa im Einsatz. Aktuell modernisieren wir unser Hosting konsequent in Richtung Kubernetes auf Harvester – als on-prem Hybrid mit der Option, mittelfristig vollständig in die Cloud zu skalieren. Als Team Lead Site Reliability Engineering (all genders) verantwortest du Reliability, Skalierbarkeit und Kosten unserer Hosting-Umgebungen, treibst diese Transformation end-to-end voran und führst das Team, das sie umsetzt. Deine Aufgaben Du verantwortest die operative Gesundheit unseres Hostings über On-Premise (Frankfurt, Stockholm) und Cloud hinweg – Verfügbarkeit, Performance, Incident Management. Du treibst die Modernisierung in Richtung Kubernetes auf Harvester aktiv voran: Cluster-Topologie, Storage (Longhorn), Networking (VLAN, Load Balancing, Ingress), Backup und Disaster Recovery. Du baust eine produktionsreife k8s-Plattform auf: Lifecycle, Upgrades, RBAC, Secrets, GitOps (Argo CD / Flux), Observability und Policy-Guardrails. Du gestaltest den NG Search Operator (Custom Kubernetes Operator) und löst Auto-Scaling (HPA, VPA, KEDA, Cluster Autoscaler) für die aktuelle Architektur. Du definierst unser On-Prem-Hybrid-Modell konkret: welche Workloads laufen wo, wie burst’en wir in die Cloud, wie halten wir Latenz und Kosten im Griff – und hältst die Architektur portabel genug für einen späteren Cloud-Only-Schritt. Du verantwortest Kapazitätsplanung und Hosting-Kosten und machst Kosten zu einem gezielt steuerbaren Hebel. Du führst und entwickelst unser derzeit 4-köpfiges Hosting-Team fachlich und disziplinarisch, verantwortest Performance und prägst die technische Standards- und Ownership-Kultur. Du machst AI zum festen Bestandteil unserer Operations: Diagnose, Automatisierung, Monitoring und Insight. Dein Profil Fundierter Hintergrund in Infrastructure oder Platform Engineering über On-Premise und Cloud hinweg. Hands-on Tiefe mit Kubernetes in Produktion: Cluster-Lifecycle, Upgrades, Networking, Storage, RBAC, Observability, GitOps-Delivery. Nachweislich starke Führungserfahrung, exzellente Kommunikation und Stakeholder-Management. Idealerweise praktische Erfahrung mit Harvester oder vergleichbaren HCI-/Virtualisierungsplattformen (KubeVirt, vSphere/ESXi, OpenStack). Erfahrung mit einer echten Migration von Bare Metal / klassischen VMs auf eine k8s-basierte Plattform – inklusive stateful Workloads, Storage-Migration, Cutover und Rollback. Sicherer Umgang mit Kubernetes Operators (Custom Controllers / CRDs), idealerweise für stateful Systeme wie Search, Datenbanken oder Streaming. Solides Verständnis von Auto-Scaling-Primitiven (HPA, VPA, Cluster Autoscaler, KEDA) und deren Zusammenspiel mit Kapazitätsplanung. Erfahrung mit On-Prem-Hybrid-Architekturen und der Verantwortung für Reliability, Kapazität und Kosten produktiver Systeme. Hands-on Fluency im Einsatz von AI-Tools im operativ
Senior Site Reliability Engineer (all genders)
Fact Finder · Berlin, Germany
On-siteabout 10 hours agoApply →Introduction FACT-Finder entwickelt Product-Discovery-Technologie für den eCommerce und ist mit den Produkten Next Generation und Infinity bei führenden Online-Shops in Europa im Einsatz. Beide Produkte bewegen sich aktuell in Richtung einer modernen, hybriden Plattform auf Basis von Kubernetes und Harvester – mit der Option, mittelfristig vollständig in die Cloud zu skalieren. Als Senior Site Reliability Engineer (SRE) sorgst du dafür, dass unsere Systeme über diese Transformation hinweg schnell, verfügbar und skalierbar bleiben. Du arbeitest mit dem Hosting-Team und erfahrenen Engineers zusammen und gestaltest den Weg zu einer modernen SaaS-Company aktiv mit. Deine Aufgaben Du definierst und verantwortest SLOs, SLIs und Error Budgets über beide Produkte hinweg und triffst datenbasierte Entscheidungen zu Reliability und Performance. Du treibst Incident Response voran: schnelle Detektion, klare Kommunikation, blameless Postmortems und nachhaltige Follow-ups. Du reduzierst manuelle Arbeit konsequent durch Automatisierung und GitOps (z. B. Argo CD / Flux) und baust self-healing sowie self-service Fähigkeiten aus. Du unterstützt den Aufbau eines NG Search Operators (Custom Kubernetes Operator / CRDs) und die Einführung von Auto-Scaling (HPA, VPA, KEDA, Cluster Autoscaler). Du entwickelst unsere Observability weiter – Metrics, Logs, Traces, Alerting und Runbooks, die on-call wirklich helfen. Du planst Kapazität und Kosten über On-Premise (Frankfurt, Stockholm) und Cloud hinweg – inklusive Burst-Szenarien in die Public Cloud. Du nutzt AI-Tools, um Diagnose, Alerting und operative Workflows spürbar zu Dein Profil Erfahrung als SRE, Infrastructure oder Production Engineer in einem SaaS- oder Plattform-Umfeld – oder ein starker Software-/Operations-Hintergrund mit klarem Willen, in die SRE-Rolle hineinzuwachsen. Solides Verständnis von SLOs, Error Budgets, Incident Management und Observability. Hands-on Erfahrung mit Kubernetes und Interesse an Cluster-Lifecycle, Upgrades und Operator-Pattern. Erfahrung oder starkes Interesse an Harvester bzw. vergleichbaren HCI-/Virtualisierungsplattformen (KubeVirt, vSphere/ESXi, OpenStack). Vertrautheit mit GitOps (Argo CD / Flux), Container Storage (Longhorn, Ceph) und Kubernetes Networking (Load Balancing, Ingress). Kenntnisse zu Auto-Scaling-Primitiven (HPA, VPA, Cluster Autoscaler, KEDA) und Kapazitätsplanung on-prem und in der Cloud. Verständnis für Netzwerke in produktionsnahen Rechenzentren (u. a. VLAN). Ausgeprägter Automatisierungsinstinkt und eine Haltung, Toil strukturell zu eliminieren. Praktische Erfahrung im Einsatz von AI-Tools im operativen Betrieb. Sehr gute Englischkenntnisse; Deutsch von Vorteil. THE JOY OF WORKING WITH US Impact from day one: Deine Arbeit wirkt direkt auf die Umsätze führender eCommerce-Marken in Europa. Moderner Tech-Stack: Kubernetes, Harvester, GitOps, Auto-Scaling und ein spannender Weg in Richtung Cloud – mit Raum, Dinge neu und richtig zu bauen. AI-first Mindset: Wir nutzen A
Senior Site Reliability Engineer (x/f/m)
Doctolib · Berlin, Germany
On-siteabout 10 hours agoApply →<h2><strong>Your Impact</strong></h2> <p>We are looking for a <strong>Senior Site Reliability Engineer</strong> to join our SRE team dedicated to platform reliability within Platform Engineering.</p> <p>Your mission will be to ensure Doctolib's platform remains reliable, scalable, and resilient at a European scale across infrastructure, observability, and cross-cutting reliability initiatives. You will work within a team driving reliability standards across 170+ applications, contributing directly to supporting 520,000 health professionals and 90 million patients in their daily healthcare journey.</p> <p>Working in the tech team at Doctolib means taking ownership of critical systems, driving reliability improvements end-to-end, and partnering closely with product and engineering teams to enable fast, safe delivery.</p> <h2><strong>What you'll do</strong></h2> <p>Your responsibilities include but are not limited to:</p> <ul> <li>Build and maintain infrastructure automation and infrastructure-as-code at scale, with a strong focus on consistency, reliability, and developer experience across 170+ applications</li> <li>Identify and lead large-scale cross-cutting reliability initiatives, including improvements to incident detection, response, and postmortem analysis capabilities</li> <li>Design, build, and improve infrastructure components that support reliability, scalability, and observability across the platform</li> <li>Define and drive SLOs, error budgets, and alerting standards across multiple product teams</li> <li>Take part in the on-call rotation, and actively contribute to improving our on-call experience by reducing noise and ensuring actionable telemetry</li> <li>Partner with software engineering teams to embed reliability practices early in the development lifecycle</li> </ul> <h2><strong>Who you are</strong></h2> <p>Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.</p> <p><strong>You'll be a great fit if you:</strong></p> <ul> <li>Have solid hands-on experience (5y+) in a Site Reliability Engineering role within a large-scale, multi-team production environment</li> <li>Have proven experience with cloud platforms such as AWS, GCP, or Azure</li> <li>Have strong experience with containerization and orchestration technologies, Kubernetes is a must, its deployment and scaling strategies ecosystem</li> <li>Have experience working closely with software engineering teams to co-design reliable systems</li> <li>Have hands-on experience with infrastructure as code, particularly Terraform, applied in real production
Senior Site Reliability Engineer (all genders)
FactFinder · Berlin, Germany
On-siteabout 16 hours agoApply →Introduction FACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. Both products are moving toward a modern, hybrid platform based on Kubernetes and Harvester – with the option to scale fully into the cloud in the mid-term. As a Senior Site Reliability Engineer (SRE), you make sure our systems stay fast, available, and scalable throughout this transformation. You work closely with the Hosting team and experienced engineers, and actively shape our journey toward a modern SaaS company. Your mission You define and own SLOs, SLIs, and error budgets across both products and make data-driven decisions on reliability and performance. You drive incident response: fast detection, clear communication, blameless postmortems, and meaningful follow-through. You consistently reduce manual work through automation and GitOps (e.g. Argo CD / Flux) and build out self-healing and self-service capabilities. You support the development of an NG Search Operator (custom Kubernetes operator / CRDs) and the rollout of auto-scaling (HPA, VPA, KEDA, cluster autoscaler). You evolve our observability – metrics, logs, traces, alerting, and runbooks that actually help on call. You plan capacity and cost across on-premise (Frankfurt, Stockholm) and cloud – including burst scenarios into the public cloud. You leverage AI tools to noticeably accelerate diagnosis, alerting, and operational workflows. Your profile Experience as an SRE, infrastructure, or production engineer in a SaaS or platform environment – or a strong software/operations background with a clear drive to grow into an SRE role. Solid understanding of SLOs, error budgets, incident management, and observability. Hands-on experience with Kubernetes and interest in cluster lifecycle, upgrades, and operator patterns. Experience with or strong interest in Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack). Familiarity with GitOps (Argo CD / Flux), container storage (Longhorn, Ceph), and Kubernetes networking (load balancing, ingress). Knowledge of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and capacity planning on-prem and in the cloud. Understanding of networking in production-grade datacenters (incl. VLAN). A strong automation instinct and a mindset to structurally eliminate toil. Practical experience using AI tools in day-to-day operations. Fluent English; German is a plus. THE JOY OF WORKING WITH US Impact from day one : Your work directly influences the revenue of leading eCommerce brands across Europe. Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right. AI-first mindset : We use AI as a real part of our daily work, not as a buzzword. Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role. Flexible work:
Senior Site Reliability Engineer (all genders)
FactFinder · Berlin, Germany
On-siteabout 16 hours agoApply →Introduction FACT-Finder entwickelt Product-Discovery-Technologie für den eCommerce und ist mit den Produkten Next Generation und Infinity bei führenden Online-Shops in Europa im Einsatz. Beide Produkte bewegen sich aktuell in Richtung einer modernen, hybriden Plattform auf Basis von Kubernetes und Harvester – mit der Option, mittelfristig vollständig in die Cloud zu skalieren. Als Senior Site Reliability Engineer (SRE) sorgst du dafür, dass unsere Systeme über diese Transformation hinweg schnell, verfügbar und skalierbar bleiben. Du arbeitest mit dem Hosting-Team und erfahrenen Engineers zusammen und gestaltest den Weg zu einer modernen SaaS-Company aktiv mit. Deine Aufgaben Du definierst und verantwortest SLOs, SLIs und Error Budgets über beide Produkte hinweg und triffst datenbasierte Entscheidungen zu Reliability und Performance. Du treibst Incident Response voran: schnelle Detektion, klare Kommunikation, blameless Postmortems und nachhaltige Follow-ups. Du reduzierst manuelle Arbeit konsequent durch Automatisierung und GitOps (z. B. Argo CD / Flux) und baust self-healing sowie self-service Fähigkeiten aus. Du unterstützt den Aufbau eines NG Search Operators (Custom Kubernetes Operator / CRDs) und die Einführung von Auto-Scaling (HPA, VPA, KEDA, Cluster Autoscaler). Du entwickelst unsere Observability weiter – Metrics, Logs, Traces, Alerting und Runbooks, die on-call wirklich helfen. Du planst Kapazität und Kosten über On-Premise (Frankfurt, Stockholm) und Cloud hinweg – inklusive Burst-Szenarien in die Public Cloud. Du nutzt AI-Tools, um Diagnose, Alerting und operative Workflows spürbar zu Dein Profil Erfahrung als SRE, Infrastructure oder Production Engineer in einem SaaS- oder Plattform-Umfeld – oder ein starker Software-/Operations-Hintergrund mit klarem Willen, in die SRE-Rolle hineinzuwachsen. Solides Verständnis von SLOs, Error Budgets, Incident Management und Observability. Hands-on Erfahrung mit Kubernetes und Interesse an Cluster-Lifecycle, Upgrades und Operator-Pattern. Erfahrung oder starkes Interesse an Harvester bzw. vergleichbaren HCI-/Virtualisierungsplattformen (KubeVirt, vSphere/ESXi, OpenStack). Vertrautheit mit GitOps (Argo CD / Flux), Container Storage (Longhorn, Ceph) und Kubernetes Networking (Load Balancing, Ingress). Kenntnisse zu Auto-Scaling-Primitiven (HPA, VPA, Cluster Autoscaler, KEDA) und Kapazitätsplanung on-prem und in der Cloud. Verständnis für Netzwerke in produktionsnahen Rechenzentren (u. a. VLAN). Ausgeprägter Automatisierungsinstinkt und eine Haltung, Toil strukturell zu eliminieren. Praktische Erfahrung im Einsatz von AI-Tools im operativen Betrieb. Sehr gute Englischkenntnisse; Deutsch von Vorteil. THE JOY OF WORKING WITH US Impact from day one: Deine Arbeit wirkt direkt auf die Umsätze führender eCommerce-Marken in Europa. Moderner Tech-Stack: Kubernetes, Harvester, GitOps, Auto-Scaling und ein spannender Weg in Richtung Cloud – mit Raum, Dinge neu und richtig zu bauen. AI-first Mindset: Wir nutzen A
Principal Site Reliability Engineer
commercetools · Berlin, Germany (Hybrid)
Hybrid7 days agoApply →Site Reliability Engineer
zattoo · Berlin, Germany
On-site13 days agoApply →Senior Site Reliability Engineer (f/m/d)
Forto Logistics SE & Co. KG · Berlin, Germany
On-siteabout 1 month agoApply →Director, Site Reliability Engineering (x/f/m)
doctolib · Berlin, Germany
On-siteabout 1 month agoApply →Senior Site Reliability Engineer
Forto · Berlin, Germany
On-siteabout 1 month agoApply →Senior Site Reliability Engineer at Forto. Apply via Ashby.
Senior Site Reliability Engineer Kubernetes Platform (m/w/d)
SysEleven GmbH · Berlin, Germany
On-siteabout 2 months agoApply →Site Reliability Engineer (SRE)
Airapps · Berlin Metropolitain Area, Germany
On-siteabout 2 months agoApply →Site Reliability Engineer (SRE) at Airapps. Apply via Ashby.
Senior Site Reliability Engineer (SRE)
nebius · Berlin, Germany; Remote - Europe
Remoteabout 2 months agoApply →Senior Site Reliability Engineer - Observability (x/f/m)
doctolib · Berlin, Germany
On-siteabout 2 months agoApply →