Senior Site Reliability Engineer – Kubernetes & Cloud

Latitude AI

United States

On-site

USD 179,200 - 268,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation packages
Medical, dental, and vision insurance
Health savings account with employer-m
Employer-matched 401(k)
Paid parental leave
Unlimited vacation
15 paid holidays
Daily lunches and snacks

Job summary

Latitude AI is hiring a Site Reliability Engineer to build and run mission-critical systems for a Ford autonomy platform. You will implement monitoring, alerting, and automation to ensure health, reliability, and performance across the stack, collaborating with ingest, mapping, ML, and deployment teams.

Responsibilities include designing platform components, implementing Kubernetes controllers, and maintaining a high-availability production environment.

Qualifications

  • Bachelor's degree in a related field and 4+ years of relevant experience (or Master's/PhD with less experience).
  • Fundamental understanding of Linux OS internals, TCP/IP networking, and storage subsystems.
  • Hands-on development in Go or Python for production-ready software.
  • Experience scaling and securing services in AWS or GCP or cloud-native environments.
  • Experience with infrastructure-as-code to automate resource creation (Terraform, CloudFormation).
  • Experience authoring Kubernetes controllers in Go and running Kubernetes in production.
  • Experience with metrics (Prometheus), logging (Elasticsearch, Loki) and tracing (Jaeger, Tempo).
  • Ability to guide teams to scale services within budget and define SLOs.
  • Strong communication skills in a diverse, distributed team.

Responsibilities

  • Build monitoring to keep the platform healthy and its reliability measurable.
  • Create alerting and runbooks to speed up detection and remediation.
  • Debug complex multi-component issues and implement robust fixes.
  • Participate in on-call rotation and blameless postmortems for continuous improvement.
  • Design platform components enabling customers to work more easily and efficiently.
  • Develop Kubernetes controllers to automate operations.

Skills

Go or Python
Kubernetes
Cloud platforms (AWS/GCP)
Terraform/CloudFormation
Monitoring/Observability
SRE on-call culture
Communication

Education

Bachelor's degree in Computer Engineering, CS, EE, Robotics or related field
Master's degree (advantage)
PhD (advantage)

Tools

Terraform
CloudFormation
Elasticsearch
Prometheus
Jaeger/Tempo

Job description

Latitude AI is hiring a Site Reliability Engineer to build and run mission-critical systems for a Ford autonomy platform. You will implement monitoring, alerting, and automation to ensure health, reliability, and performance across the stack, collaborating with ingest, mapping, ML, and deployment teams.

Responsibilities include designing platform components, implementing Kubernetes controllers, and maintaining a high-availability production environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Scalable Cloud Reliability & Observability
Senior SRE: Scalable Cloud Reliability & Observability

Latitude AI • Palo Alto (CA)

On-site
USD 179,000 - 269,000
Health insurance
401(k) match
Unlimited vacation
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Latitude AI • United States

On-site
USD 179,000 - 269,000
Competitive compensation packages
Medical, dental, and vision insurance
Health savings account with employer-m
+5
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer - AI Cloud Platform
Senior Site Reliability Engineer - AI Cloud Platform

Lambda • United States

Hybrid
USD 160,000 - 220,000
Senior Enterprise Systems Engineer: Global Infra & Automation
Senior Enterprise Systems Engineer: Global Infra & Automation

Latitude AI • Pittsburgh

On-site
USD 110,000 - 170,000
Health insurance
401(k) plan with employer match
Paid parental leave
+2
Senior Site Reliability Engineer – AI Platform
Senior Site Reliability Engineer – AI Platform

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 130,000
Competitive salary
Flexible work environment
High-performance culture
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Latitude AI • Palo Alto (CA)

On-site
USD 179,000 - 269,000
Health insurance
401(k) match
Unlimited vacation
+2
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000