Senior Platform Engineer (Core Infrastructure)

Lambda

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health, dental vision
401k match
5 sick days and 12 paid holidays
Flexible PTO
Parental, medical, caregiver leave

Job summary

Lambda is seeking a senior Platform/SRE engineer to architect and run Kubernetes at scale across AWS and bare-metal data centers. You will own cluster reliability, performance, security, and automation, partnering with product teams to deliver scalable cloud-native services and robust CI/CD pipelines.

Ideal candidates have 5+ years in platform or SRE roles, deep Kubernetes knowledge, Go or Python scripting ability, and strong networking/security fundamentals.

Qualifications

  • 5+ years in Platform, Infrastructure, or SRE roles, incl. running Kubernetes in production at scale.
  • Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting).
  • Strong coding skills in Go or Python for automation and tooling.
  • Experience with multi-cluster, multi-cloud, or hybrid environments.
  • Solid grounding in networking, service meshes, and container runtimes.
  • Practical security experience: network policies, secrets management, and image scanning.

Responsibilities

  • Architect, deploy, and operate Kubernetes clusters across AWS and Lambda’s datacenters.
  • Build and maintain automation for cluster lifecycle management — provisioning, upgrades, and scaling.
  • Own reliability, performance, and security of Kubernetes workloads in production.
  • Implement observability, logging, and alerting for clusters and critical workloads.
  • Partner with product teams to design scalable, cloud-native services and CI/CD pipelines.
  • Set standards for resource management, networking, and RBAC across the platform.
  • Lead incident response, root-cause analysis, and post-mortems for platform issues.
  • Mentor engineers and raise the bar for platform engineering across the org.

Skills

Kubernetes
Observability
Terraform
Pulumi
Go
Python
GitOps
Multi-cluster
Cost optimization
Security best practices
Networking
RBAC

Tools

Helm
Kustomize
Argo Workflows
Temporal
Cadence

Job description

  • Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance
  • Architect, deploy, and operate Kubernetes clusters across AWS and Lambda’s bare-metal datacenters
  • Build and maintain automation for cluster lifecycle management — provisioning, upgrades, and scaling
  • Own the reliability, performance, and security of Kubernetes workloads in production
  • Implement observability, logging, and alerting for clusters and critical workloads
  • Partner with product teams to design scalable, cloud-native services and CI/CD pipelines
  • Set the standards for resource management, networking, and RBAC across the platform
  • Lead incident response, root-cause analysis, and post-mortems for platform issues
  • Mentor engineers and raise the bar for platform engineering across the org
Benefits
  • Health, dental vision
  • 401k match
  • 5 sick days and 12 paid holidays
  • Flexible PTO
  • Paid parental, medical, and caregiver leave
Hands-on with observability stacks (Prometheus, Grafana, OpenTelemetry)Solid grounding in networking, service meshes, and container runtimesProficient with infrastructure-as-code (Terraform, Pulumi, or equivalent)Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting)5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scaleStrong with Helm, Kustomize, or similar, and GitOps-based deliveryPractical security experience: network policies, secrets management, and image scanningStrong coding skills in Go or Python for automation and toolingKnowledge of GPU scheduling, HPC workloads, or ML/AI infrastructureExperience with multi-cluster, multi-cloud, or hybrid environmentsExperience with workflow orchestration / durable execution frameworks (Temporal, Cadence, or Argo Workflows)Exposure to cost optimization and capacity planning for large clustersContributions to CNCF or Kubernetes open-source projectsCKA/CKS certification
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer - Core Infrastructure
Senior Platform Engineer - Core Infrastructure

Lambda • Bellevue (WA)

Hybrid
USD 190,000 - 270,000
Health, dental, and vision coverage
401k with 2% company match
Flexible PTO
+3
Senior Platform Engineer - Cloud Infra & Kubernetes
Senior Platform Engineer - Cloud Infra & Kubernetes

Lambda Labs • United States

On-site
USD 180,000 - 260,000
Health insurance
401k plan
Flexible PTO
+2
Senior Platform Engineer - Core Infrastructure
Senior Platform Engineer - Core Infrastructure

Lambda Labs • United States

On-site
USD 180,000 - 260,000
Health insurance
401k plan
Flexible PTO
+2
Senior Platform Engineer — Cloud Kubernetes & Automation Leader
Senior Platform Engineer — Cloud Kubernetes & Automation Leader

Lambda • United States

Remote
USD 180,000 - 240,000
Health, dental vision
401k match
5 sick days and 12 paid holidays
+2
Senior Platform Engineer - Core Infrastructure
Senior Platform Engineer - Core Infrastructure

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, vision coverage
401k with company match
Wellness and commuter stipends
+1
Senior Platform Engineer – Kubernetes & Cloud Infra
Senior Platform Engineer – Kubernetes & Cloud Infra

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, vision coverage
401k with company match
Wellness and commuter stipends
+1
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

Socket.dev • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Health, dental, vision
4-day in-office work week
Wellness stipend
+1
Senior Site Reliability Engineer - Core Cloud Platform Software
Senior Site Reliability Engineer - Core Cloud Platform Software

Front Door Defense • San Jose (CA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k Plan with company match
Flexible paid time off
+1