SRE Lead: Kubernetes, Hybrid Cloud, AI-Driven Ops

FACT-Finder

Germany (OH)

Hybrid

USD 103,603 - 172,672

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Learning budget
Hybrid work model

Job summary

FACT-Finder seeks a Team Lead Site Reliability Engineering to own reliability and cost of hosting across on-prem and cloud. You will drive the transformation toward Kubernetes on Harvester, shaping the NG Search Operator and implementing scalable, AI-assisted operations.

You will lead a 4-person Hosting team, define topology, upgrades, and policy guardrails while ensuring latency and availability, with a focus on cost efficiency and cloud-ready portability.

Qualifications

  • Strong background in infrastructure or platform engineering across on-premise and cloud.
  • Hands-on depth with Kubernetes in production: cluster lifecycle, upgrades, networking, storage, RBAC, observability, GitOps delivery.
  • Proven people leadership experience, excellent communication and stakeholder management skills.
  • Ideally practical experience with Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).
  • Experience leading a real migration from bare metal / classic VMs to a k8s-based platform – including stateful workloads, storage migration, cutover, and rollback.
  • Comfort designing or operating Kubernetes operators (custom controllers / CRDs), ideally for stateful systems like search, databases, or streaming.

Responsibilities

  • You own the operational health of our hosting across on-premise and cloud – availability, performance, and incident management.
  • You actively drive the modernization toward Kubernetes on Harvester: cluster topology, storage, networking, backup, and disaster recovery.
  • You build a production-grade k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps, observability, and policy guardrails.
  • You shape the NG Search Operator and solve auto-scaling for the current architecture.
  • You concretely define our on-prem hybrid model: workloads, cloud bursting, latency, and cost control.
  • You own capacity planning and hosting cost and turn cost into a deliberate, managed lever.
  • You lead and develop our 4-person Hosting team, own performance and technical direction, and set ownership culture.
  • You make AI a core part of our operations: diagnosis, automation, monitoring, and insight.

Skills

Kubernetes production
Team leadership
Cloud & on-prem
GitOps
Performance tuning
Networking
RBAC
Storage management
Automation

Tools

Kubernetes
Harvester
OpenStack
vSphere/ESXi
KubeVirt

Job description

FACT-Finder seeks a Team Lead Site Reliability Engineering to own reliability and cost of hosting across on-prem and cloud. You will drive the transformation toward Kubernetes on Harvester, shaping the NG Search Operator and implementing scalable, AI-assisted operations.

You will lead a 4-person Hosting team, define topology, upgrades, and policy guardrails while ensuring latency and availability, with a focus on cost efficiency and cloud-ready portability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Kubernetes, GitOps & Hybrid Cloud
Senior SRE: Kubernetes, GitOps & Hybrid Cloud

FACT-Finder • Germany (OH)

Hybrid
USD 103,000 - 139,000
Hybrid work model
Open feedback culture
Reliability discipline
Hybrid Lead SRE: Scale, Automate & Stabilize Cloud Services
Hybrid Lead SRE: Scale, Automate & Stabilize Cloud Services

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior SRE: Kubernetes, CI/CD & Cloud Reliability Leader
Senior SRE: Kubernetes, CI/CD & Cloud Reliability Leader

techchaintalent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior SRE - Hybrid, Platform Reliability Lead
Senior SRE - Hybrid, Platform Reliability Lead

TransUnion LLC • Reston (VA)

Hybrid
USD 112,000 - 188,000
Day-one medical, dental, vision
Company-paid basic life/AD&D
12 weeks paid parental leave
+2
Lead SRE: Cloud Infra, Kubernetes & CI/CD
Lead SRE: Cloud Infra, Kubernetes & CI/CD

HTC Global Services • Orlando (FL)

On-site
USD 150,000 - 210,000
Health Insurance
401(k) matching
Paid Time Off
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
OpenSearch SRE: Cloud Reliability & Ops (Contract)
OpenSearch SRE: Cloud Reliability & Ops (Contract)

Information Consulting Services • Herndon (VA)

On-site
USD 120,000 - 190,000
Senior SRE Lead — Hybrid Cloud Reliability
Senior SRE Lead — Hybrid Cloud Reliability

Cognizant • Bridgewater (MA)

Hybrid
USD 63,000 - 100,000
Medical/Dental/Vision/Life Insurance
401(k) plan and contributions
Paid time off
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE Lead – Cloud, Kubernetes & Security (Hybrid)
Senior SRE Lead – Cloud, Kubernetes & Security (Hybrid)

Socket.dev • San Jose (CA)

Hybrid
USD 124,000 - 271,000