Site Reliability Engineer — Observability Platform

Devopsroles

United States

Remote

USD 85,000 - 193,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision coverage
Parental leave and back-up childcare
Adoption and surrogacy support
Vehicle discount program
Tuition assistance
Paid time off and holidays

Job summary

Ford is hiring an experienced Site Reliability Engineer to architect, extend, and scale our global observability platform. You will design scalable pipelines, define SLIs/SLOs, and build infrastructure-as-code templates to standardize instrumentation across hybrid environments.

Join a team focused on reliability, performance, and continuous innovation. The role requires strong distributed systems design, automation, and production operations expertise, with hands-on work in cloud platforms

Qualifications

  • Bachelor’s degree in CS or equivalent.
  • At least 3+ years in an SRE role.
  • 5+ years programming in Python/Go/Java/Scala/C/C++.
  • 3+ years IaC templates in Terraform/ToFu.
  • 3+ years with APM/monitoring tools (Dynatrace, New Relic, ELK, Splunk, Prometheus, DataDog).
  • 3+ years with Java/J2EE, Spring Boot, NoSQL/SQL databases, cloud (GCP/AWS/Azure), and Docker/Kubernetes.
  • Experience with REST APIs and microservices.
  • CI/CD automation and SDLC practices.
  • Strong observability and MTTR/MTTD focus.
  • Understanding of networking basics (TCP/IP).

Responsibilities

  • Design scalable observability pipelines across metrics, logging, tracing, and alerting.
  • Define SLIs/SLOs and error budgets to maximize availability.
  • Build IaC templates to standardize instrumentation.
  • Develop automation to improve resilience and scalability of apps.
  • Perform safe destructive testing to reveal vulnerabilities.
  • Create tooling to improve reliability and time-to-market.
  • Reduce toil via automation to focus on engineering and innovation.
  • Collaborate with development teams on scalable systems.
  • Identify stability risks and mitigation plans.
  • Monitor key metrics like errors, latency, capacity, and resource utilization.
  • Analyze performance of new and production systems to drive improvements.
  • Troubleshoot distributed systems and lead root-cause analyses.
  • Participate in incident response and postmortems.
  • Mentor team members and share knowledge.
  • Evaluate AI/ML capabilities for anomaly detection and insights.
  • Embed observability best practices into system design and deployment workflows.

Education

Bachelor’s Degree in Computer Science or equivalent

Tools

Python
Go
Java/Scala
C
C++
Terraform
ToFu
Dynatrace
New Relic
ELK
Splunk
Prometheus
Sensu
Nagios
Kafka
DataDog
J2EE
NoSQL/SQL Databases
Spring Boot
GCP
AWS
Azure
Docker
Kubernetes
RESTful APIs
CI/CD

Job description

Ford is hiring an experienced Site Reliability Engineer to architect, extend, and scale our global observability platform. You will design scalable pipelines, define SLIs/SLOs, and build infrastructure-as-code templates to standardize instrumentation across hybrid environments.

Join a team focused on reliability, performance, and continuous innovation. The role requires strong distributed systems design, automation, and production operations expertise, with hands-on work in cloud platforms

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE: Observability Platform (Remote)
SRE: Observability Platform (Remote)

Ford • Northern (KY)

Remote
USD 85,000 - 193,000
Medical, dental, vision coverage
Paid time off and holidays
Vehicle discount program
+2
Site Reliability Engineer - Observability Platform
Site Reliability Engineer - Observability Platform

Devopsroles • United States

Remote
USD 85,000 - 193,000
Medical, dental, vision coverage
Parental leave and back-up childcare
Adoption and surrogacy support
+3
Senior Site Reliability Engineer - Observability & Scale
Senior Site Reliability Engineer - Observability & Scale

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Site Reliability Engineer - Observability Platform
Site Reliability Engineer - Observability Platform

Ford • Northern (KY)

Remote
USD 85,000 - 193,000
Medical, dental, vision coverage
Paid time off and holidays
Vehicle discount program
+2
Platform SRE Engineer: Observability & Automation
Platform SRE Engineer: Observability & Automation

Fathom • United States

Remote
USD 120,000 - 180,000
Site Reliability Engineer – Observability & Automation Lead
Site Reliability Engineer – Observability & Automation Lead

National Oilwell Varco • Houston (TX)

On-site
USD 140,000 - 180,000
Site Reliability Engineer — Global Scale & Observability
Site Reliability Engineer — Global Scale & Observability

Human Ventures, LLC. • United States

On-site
USD 90,000 - 100,000
SRE Platform Engineer - Go, Cloud & Monitoring (Remote)
SRE Platform Engineer - Go, Cloud & Monitoring (Remote)

Artha • Salem (OR), Northern (KY)

Hybrid
USD 85,000 - 193,000
Medical coverage
Parental leave
Back-up child care
+1
Site Reliability Engineer: Distributed Systems Observability
Site Reliability Engineer: Distributed Systems Observability

AppLab Systems, Inc • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer — Platform Observability & Autonomy
Site Reliability Engineer — Platform Observability & Autonomy

MaintainX • San Francisco (CA), Northern (KY)

Hybrid
USD 130,000 - 180,000