Senior Site Reliability Engineer

Latitude AI

United States

On-site

USD 179,200 - 268,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation packages
Medical, dental, and vision insurance
Health savings account with employer-m
Employer-matched 401(k)
Paid parental leave
Unlimited vacation
15 paid holidays
Daily lunches and snacks

Job summary

Latitude AI is hiring a Site Reliability Engineer to build and run mission-critical systems for a Ford autonomy platform. You will implement monitoring, alerting, and automation to ensure health, reliability, and performance across the stack, collaborating with ingest, mapping, ML, and deployment teams.

Responsibilities include designing platform components, implementing Kubernetes controllers, and maintaining a high-availability production environment.

Qualifications

  • Bachelor's degree in a related field and 4+ years of relevant experience (or Master's/PhD with less experience).
  • Fundamental understanding of Linux OS internals, TCP/IP networking, and storage subsystems.
  • Hands-on development in Go or Python for production-ready software.
  • Experience scaling and securing services in AWS or GCP or cloud-native environments.
  • Experience with infrastructure-as-code to automate resource creation (Terraform, CloudFormation).
  • Experience authoring Kubernetes controllers in Go and running Kubernetes in production.
  • Experience with metrics (Prometheus), logging (Elasticsearch, Loki) and tracing (Jaeger, Tempo).
  • Ability to guide teams to scale services within budget and define SLOs.
  • Strong communication skills in a diverse, distributed team.

Responsibilities

  • Build monitoring to keep the platform healthy and its reliability measurable.
  • Create alerting and runbooks to speed up detection and remediation.
  • Debug complex multi-component issues and implement robust fixes.
  • Participate in on-call rotation and blameless postmortems for continuous improvement.
  • Design platform components enabling customers to work more easily and efficiently.
  • Develop Kubernetes controllers to automate operations.

Skills

Go or Python
Kubernetes
Cloud platforms (AWS/GCP)
Terraform/CloudFormation
Monitoring/Observability
SRE on-call culture
Communication

Education

Bachelor's degree in Computer Engineering, CS, EE, Robotics or related field
Master's degree (advantage)
PhD (advantage)

Tools

Terraform
CloudFormation
Elasticsearch
Prometheus
Jaeger/Tempo

Job description

Latitude AI (lat.ai) is building the future of Ford’s autonomy roadmap to make travel safer, less stressful, and more enjoyable for everyone. Bringing this vision to scale, our fully in-house developed hands-free ADAS platform will debut on Ford’s all-new Universal Electric Vehicle in 2027.

When you join the Latitude team, you’ll work alongside leading experts across machine learning and robotics, cloud platforms, mapping, sensors and compute systems, test operations, systems and safety engineering – all dedicated to redefining the relationship between people and their vehicles for millions of customers.

As a Ford Motor Company subsidiary, we operate independently to develop automated driving technology at the speed of a technology startup. Latitude is headquartered in Pittsburgh with engineering centers in Dearborn, Mich., and Palo Alto, Calif.

Meet the team:

As a Site Reliability Engineer on the team, you will be responsible for helping to build and run these mission critical systems. Through the implementation of monitoring and automation, you will constantly ensure the health, reliability, scalability, and performance of the platforms.

The Site Reliability team interacts with engineering teams including ingest/data processing, mapping, labeling, triage, machine learning (detection, prediction, tracking), motion planning/control, offline simulation, and release/deployment teams to provide uniform service observability and incident response.

What you’ll do:
  • Build monitoring to ensure our platform is healthy and its reliability measurable
  • Build alerting and a set of runbooks to enable faster detection and remediation of platform issues
  • Debug complex issues that may combine multiple components of the stack and ensure proper fixes are implemented to prevent these issues from happening again
  • Participate in an on-call rotation and culture of continuous improvement through blameless postmortems
  • Design and implement components of the platform to enable features that make the work of our customers possible, simpler and more efficient
  • Build Kubernetes controllers to automate operations
What you'll need to succeed:
  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, Robotics or a related field and 4+ years of relevant experience (or Master's degree and 2+ years of relevant experience, or PhD)
  • Fundamental understanding of Linux operating system internals, TCP/IP networking, and storage subsystems
  • Hands on development in Go or Python to create robust software that can run reliably in production
  • Strong experience scaling and securing services in the cloud (AWS, GCP) or cloud native environments
  • Experience using infrastructure-as-code principles to automate the creation of infrastructure resources (e.g. Terraform, CloudFormation)
  • Experience authoring and maintaining Kubernetes Controllers in Go
  • Experience running Kubernetes and related core components in a large-scale, production environment
  • Experience with metrics (e.g. Prometheus), logging (e.g. Elasticsearch, Loki) and tracing (e.g. Jaeger, Tempo) systems
  • Understanding of engineering design limitations and ability to provide guidance to teams to scale their services to achieve desired performance within budget
  • A focus on increasing service reliability through defining and adhering to SLOs
  • Strong communication skills and the ability to work effectively in a diverse and distributed team
What we offer you:
  • Competitive compensation packages
  • High-quality individual and family medical, dental, and vision insurance
  • Health savings account with available employer match
  • Employer-matched 401(k) retirement plan with immediate vesting
  • Employer-paid group term life insurance and the option to elect voluntary life insurance
  • Paid parental leave
  • Paid medical leave
  • Unlimited vacation
  • 15 paid holidays
  • Daily lunches, snacks, and beverages available in all office locations
  • Pre-tax spending accounts for healthcare and dependent care expenses
  • Pre-tax commuter benefits
  • Monthly wellness stipend
  • Adoption/Surrogacy support program
  • Backup child and elder care program
  • Professional development reimbursement
  • Employee assistance program
  • Discounted programs that include legal services, identity theft protection, pet insurance, and more
  • Company and team bonding outlets: employee resource groups, quarterly team activity stipend, and wellness initiatives

Learn more about Latitude’s team, mission and career opportunities at lat.ai!

The expected base salary range for this full-time position in California is $179,200 - $268,800 USD. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Latitude employees are also eligible to participate in Latitude’s annual bonus programs, equity compensation, and generous Company benefits program, subject to eligibility requirements.

Candidates for positions with Latitude AI must be legally authorized to work in the United States on a permanent basis. Verification of employment eligibility will be required at the time of hire. Visa sponsorship is available for this position.

We are an Equal Opportunity Employer committed to a culturally diverse workforce. All qualified applicants will receive consideration for employment without regard to race, religion, color, age, sex, national origin, sexual orientation, gender identity, disability status or protected veteran status.

#LI-CG1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Latitude AI • Palo Alto (CA)

On-site
USD 179,000 - 269,000
Health insurance
401(k) match
Unlimited vacation
+2
Senior Software Engineer - Runtime Infrastructure Frameworks
Senior Software Engineer - Runtime Infrastructure Frameworks

Latitude AI • Detroit (MI)

On-site
USD 179,000 - 269,000
Competitive compensation
Health insurance
401(k) with employer match
+5
Senior Enterprise Systems Engineer
Senior Enterprise Systems Engineer

Latitude AI • Pittsburgh

On-site
USD 110,000 - 170,000
Health insurance
401(k) plan with employer match
Paid parental leave
+2
Senior Software Engineer - Deploy Infrastructure
Senior Software Engineer - Deploy Infrastructure

Latitude AI • Detroit (MI)

On-site
USD 179,000 - 269,000
Competitive compensation
Health insurance
401(k) retirement plan with vesting
+3
Senior Software Engineer - Runtime Infrastructure Frameworks
Senior Software Engineer - Runtime Infrastructure Frameworks

Latitude AI LLC • Northern (KY)

Hybrid
USD 179,000 - 269,000
Competitive compensation packages
Health insurance
401(k) retirement plan with company-m匹
Senior Staff Software Systems Engineer - Advanced Autonomy Systems Engineering Lead
Senior Staff Software Systems Engineer - Advanced Autonomy Systems Engineering Lead

Latitude AI • Pittsburgh

On-site
USD 253,000 - 380,000
High-quality medical insurance
401(k) retirement plan
Unlimited vacation
+1
Senior Software Engineer - Runtime Infrastructure Frameworks
Senior Software Engineer - Runtime Infrastructure Frameworks

Latitude • Pittsburgh, Palo Alto (CA), Detroit (MI)

On-site
USD 179,000 - 269,000
Medical, dental, vision insurance
401(k) with company match
Senior Software Engineer - Simulation Infrastructure
Senior Software Engineer - Simulation Infrastructure

Latitude AI • Pittsburgh

On-site
USD 179,200 - 268,800
Competitive compensation packages
Health insurance
Dental and vision insurance
+3
Senior / Staff Software Engineer- Unified Modeling (Machine Learning)
Senior / Staff Software Engineer- Unified Modeling (Machine Learning)

Latitude AI • Pittsburgh, Detroit (MI), Palo Alto (CA)

Hybrid
USD 179,000 - 269,000
Medical coverage
401(k) retirement plan with vesting
Paid parental leave
+1
Senior Software Engineer - ML Performance
Senior Software Engineer - ML Performance

Socket.dev • Palo Alto (CA), Pittsburgh

On-site
USD 179,000 - 269,000
Medical insurance
401(k) retirement plan
Paid parental leave
+2