Senior Site Reliability Engineer - Managed Kubernetes

Lambda

Deutschland

Vor Ort

EUR 138.000 - 190.000

Vollzeit

Vor 13 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Lambda is seeking a Kubernetes-focused engineer to operate and scale bare-metal clusters across thousands of nodes. You will handle incidents, degradation, and resizing, while assisting customers with workloads, storage, and authentication.

You will join a team that collaborates with HPC Ops and Datacenters to ensure reliability. The role requires presence in designated office locations four days a week with a set work-from-home day.

Aufgaben

  • Operate and maintain bare-metal Kubernetes clusters, scaling up to thousands of nodes
  • Handle cluster degradation, recovery, resizing, and incident response using fleet management tools
  • Participate in a well-managed on-call rotation for critical incidents
  • Assist customers with Kubernetes questions, workload integration, storage, and authentication
  • Work closely with our HPC Ops and Datacenters teams to ensure reliability and performance

Jobbeschreibung

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU. If you'd like to build the world's best AI cloud, join us. Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday. Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.

What You’ll Do
  • Operate and maintain bare-metal Kubernetes clusters, scaling up to thousands of nodes
  • Handle cluster degradation, recovery, resizing, and incident response using fleet management tools
  • Participate in a well-managed on-call rotation for critical incidents
  • Assist customers with Kubernetes questions, workload integration, storage, and authentication
  • Work closely with our HPC Ops and Datacent...
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer - Core Cloud Platform
Senior Site Reliability Engineer - Core Cloud Platform

Lambda • Deutschland

Hybrid
EUR 90.000 - 150.000
Senior Software Engineer - Managed Kubernetes
Senior Software Engineer - Managed Kubernetes

Lambda • Deutschland

Hybrid
EUR 156.000 - 225.000
Senior Site Reliability Engineer - SDN
Senior Site Reliability Engineer - SDN

Lambda • Deutschland

Hybrid
EUR 122.000 - 158.000
AI Operations Engineer - IT/Internal Infrastructure
AI Operations Engineer - IT/Internal Infrastructure

Lambda • Deutschland

Hybrid
EUR 90.000 - 120.000
Security Engineer
Security Engineer

Lambda • Deutschland

Hybrid
EUR 60.000 - 95.000
Data Center Asset Management Specialist
Data Center Asset Management Specialist

Lambda • Deutschland

Hybrid
EUR 56.000 - 78.000
Technical Program Manager - Data Center Delivery
Technical Program Manager - Data Center Delivery

Lambda • Deutschland

Hybrid
EUR 122.000 - 176.000
Manager, Investor Relations
Manager, Investor Relations

Lambda • Deutschland

Hybrid
EUR 130.000 - 200.000
Senior Manager, Revenue Technical Accounting
Senior Manager, Revenue Technical Accounting

Lambda • Deutschland

Remote
EUR 65.000 - 85.000
Internal Audit Lead - IT Systems and Controls
Internal Audit Lead - IT Systems and Controls

Lambda • Deutschland

Hybrid
EUR 104.000 - 157.000