Site Reliability Engineer - Observability Platform

Devopsroles

United States

Remote

USD 85,000 - 193,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision coverage
Parental leave and back-up childcare
Adoption and surrogacy support
Vehicle discount program
Tuition assistance
Paid time off and holidays

Job summary

Ford is hiring an experienced Site Reliability Engineer to architect, extend, and scale our global observability platform. You will design scalable pipelines, define SLIs/SLOs, and build infrastructure-as-code templates to standardize instrumentation across hybrid environments.

Join a team focused on reliability, performance, and continuous innovation. The role requires strong distributed systems design, automation, and production operations expertise, with hands-on work in cloud platforms

Qualifications

  • Bachelor’s degree in CS or equivalent.
  • At least 3+ years in an SRE role.
  • 5+ years programming in Python/Go/Java/Scala/C/C++.
  • 3+ years IaC templates in Terraform/ToFu.
  • 3+ years with APM/monitoring tools (Dynatrace, New Relic, ELK, Splunk, Prometheus, DataDog).
  • 3+ years with Java/J2EE, Spring Boot, NoSQL/SQL databases, cloud (GCP/AWS/Azure), and Docker/Kubernetes.
  • Experience with REST APIs and microservices.
  • CI/CD automation and SDLC practices.
  • Strong observability and MTTR/MTTD focus.
  • Understanding of networking basics (TCP/IP).

Responsibilities

  • Design scalable observability pipelines across metrics, logging, tracing, and alerting.
  • Define SLIs/SLOs and error budgets to maximize availability.
  • Build IaC templates to standardize instrumentation.
  • Develop automation to improve resilience and scalability of apps.
  • Perform safe destructive testing to reveal vulnerabilities.
  • Create tooling to improve reliability and time-to-market.
  • Reduce toil via automation to focus on engineering and innovation.
  • Collaborate with development teams on scalable systems.
  • Identify stability risks and mitigation plans.
  • Monitor key metrics like errors, latency, capacity, and resource utilization.
  • Analyze performance of new and production systems to drive improvements.
  • Troubleshoot distributed systems and lead root-cause analyses.
  • Participate in incident response and postmortems.
  • Mentor team members and share knowledge.
  • Evaluate AI/ML capabilities for anomaly detection and insights.
  • Embed observability best practices into system design and deployment workflows.

Education

Bachelor’s Degree in Computer Science or equivalent

Tools

Python
Go
Java/Scala
C
C++
Terraform
ToFu
Dynatrace
New Relic
ELK
Splunk
Prometheus
Sensu
Nagios
Kafka
DataDog
J2EE
NoSQL/SQL Databases
Spring Boot
GCP
AWS
Azure
Docker
Kubernetes
RESTful APIs
CI/CD

Job description

We made history and now we work to transform the future – for our customers, our communities and our families. You'll see your work on the road every day, helping people move freely and pursue their dreams. At Ford, you can build more than vehicles. Come build what matters.

Enterprise Technology plays a critical part in shaping the future of mobility. If you’re looking for the chance to leverage advanced technology to redefine the transportation landscape, enhance the customer experience and improve people’s lives, this is the opportunity for you. Join us and challenge your IT expertise and analytical skills to help create vehicles that are as smart as you are.

The Observability Platform Team designs, builds, and operates the monitoring and observability infrastructure that underpins visibility into application performance across hybrid environments—on-prem and cloud. Our platform integrates AI-driven analytics with intuitive dashboards to give engineering teams the telemetry, metrics, logs, and traces they need to detect issues faster, reduce MTTR, and drive continuous performance optimization.

We're hiring an experienced Site Reliability Engineer to architect, extend, and scale our global observability platform. This role sits at the intersection of software and systems engineering, requiring strong skills in distributed systems design, infrastructure automation, and production operations to ensure high availability, scalability, and maintainability of our monitoring stack.

  • Design and implement scalable observability pipelines spanning metrics, logging, tracing, and alerting
  • Define and operationalize Service Level Indicators (SLIs) and Service Level Objectives (SLOs), and establish error budgets to effectively drive maximum availability and uptime
  • Build reusable infrastructure-as-code templates and frameworks to standardize observability instrumentation and onboarding
  • Architect, design, and develop automation to improve the resilience, recoverability, availability, and scalability of supported applications
  • Leverage experience to safely perform destructive testing to seek and discover vulnerabilities
  • Develop tooling to improve reliability, quality, and time-to-market for software solutions
  • Identify and reduce or eliminate toil via automation to maximize time spent on engineering and innovation
  • Collaborate with development teams to design, build, and operate scalable and resilient software systems using cloud-native principles
  • Proactively identify stability risks and work with engineering leadership to establish appropriate mitigation plans
  • Regularly review key technical metrics such as transaction errors, logging, response times, caching strategies, conversion/bounce rates, capacity, and resource utilization
  • Conduct performance analysis and optimization of new and in-production systems, measuring and optimizing performance to get ahead of customer needs and drive continuous innovation
  • Solve complex architecture, design, and business problems by simplifying processes, optimizing systems, and removing bottlenecks
  • Recognize, validate, and evangelize emerging technologies and architectures that align with business objectives
  • Troubleshoot complex, distributed production systems and drive root-cause analysis for platform-level incidents
  • Participate in incident response, support, recovery, and postmortem analysis
  • Provide technical guidance and mentorship to other team members
  • Continuously evaluate and integrate AI/ML capabilities to enhance anomaly detection, alerting precision, and performance insights
  • Collaborate cross-functionally with engineering teams to embed observability best practices into system design and deployment workflows
  • Bachelor’s Degree in Computer Science or equivalent experience
  • 3+ years of experience in an SRE role
  • 5+ years of programming experience with one or more of: Python, Go, Java/Scala, C, or C++
  • 3+ years of experience building reusable infrastructure-as-code templates & frameworks in Terraform or ToFu
  • 3+ years of experience with APM and monitoring tools such as Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog
  • 3+ years of experience with J2EE, NoSQL/SQL datastores, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes in developing multi-tier applications
  • Experience with RESTful APIs and microservices platforms
  • Working knowledge of the TCP/IP stack, internet routing, and load balancing
  • Strong proficiency with Google Cloud Platform and its library of services
  • Experience with automated, test-driven development in CI/CD pipelines
  • Thorough understanding of software development and agile methodologies
  • Understanding of, and ability to implement, effective observability strategies to improve MTTD/MTTR (Mean Time to Detect/Resolve)

As an established global company, we offer the benefit of choice. You can choose what your Ford future will look like: will your story span the globe, or keep you close to home? Will your career be a deep dive into what you love, or a series of new teams and new skills? Will you be a leader, a changemaker, a technical expert, a culture builder…or all of the above? No matter what you choose, we offer a work life that works for you, including:

  • Immediate medical, dental, vision and prescription drug coverage
  • Flexible family care days, paid parental leave, new parent ramp-up programs, subsidized back-up childcare and more
  • Family building benefits including adoption and surrogacy expense reimbursement, fertility treatments, and more
  • Vehicle discount program for employees and family members and management leases
  • Tuition assistance
  • Established and active employee resource groups
  • Paid time off for individual and team community service
  • A generous schedule of paid holidays, including the week between Christmas and New Year’s Day
  • Paid time off and the option to purchase additional vacation time.

This position is a range of salary grades 6-8 and ranges from $85,400-$192,900. Final determination of salary grade will be based on candidate's skills and experience, and base salary will be set within the applicable range according to job scope, responsibility and competitive market value. Internal applicants: moving into this role may result in an adjustment to your current compensation based on the posted pay range for this role, taking into consideration your qualifications and other relevant factors.

Visa sponsorship is not available for this position.

Candidates for positions with Ford Motor Company must be legally authorized to work in the United States. Verification of employment eligibility will be required at the time of hire.

We are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, age, sex, national origin, sexual orientation, gender identity, disability status or protected veteran status. In the United States, if you need a reasonable accommodation for the online application process due to a disability, please call 1-888-336-0660.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Observability Platform
Site Reliability Engineer - Observability Platform

Ford • Northern (KY)

Remote
USD 85,000 - 193,000
Medical, dental, vision coverage
Paid time off and holidays
Vehicle discount program
+2
Senior Software Engineer
Senior Software Engineer

Lincoln Motor Company • United States

On-site
USD 133,000 - 251,000
Medical, dental, vision coverage
Family care days
Paid parental leave
+3
W Site Reliability Engineer - Deviantart Wix Salem, Oregon, US
W Site Reliability Engineer - Deviantart Wix Salem, Oregon, US

Artha • Salem (OR), Northern (KY)

Hybrid
USD 85,000 - 193,000
Medical coverage
Parental leave
Back-up child care
+1
Senior Software Engineer
Senior Software Engineer

Ford • Dearborn (MI)

Hybrid
USD 100,000 - 193,000
Medical, dental, vision coverage
Paid time off
Tuition assistance
+2
Software Engineer, Distributed Systems & Networking
Software Engineer, Distributed Systems & Networking

Lincoln Motor Company • Palo Alto (CA)

Hybrid
USD 139,000 - 233,000
Medical, dental, vision coverage
Vehicle discount program
Tuition assistance
+1
Full Stack Software Engineer - Ford Pro
Full Stack Software Engineer - Ford Pro

Lincoln Motor Company • Dearborn (MI)

On-site
USD 85,000 - 193,000
Medical coverage
Dental coverage
Vision coverage
+6
Sr Kubernetes Platform Engineer
Sr Kubernetes Platform Engineer

Ford • United States

Remote
USD 100,000 - 193,000
Medical/dental/vision coverage
Flexible remote work policy
Vehicle discount program
+3
Technology Specialist
Technology Specialist

Ford Motor Company • Dearborn (MI)

Hybrid
USD 74,000 - 145,000
Hybrid work model
Competitive compensation
Tuition assistance
Cloud Systems Developer
Cloud Systems Developer

Lincoln Motor Company • United States

Remote
USD 85,000 - 193,000
Medical coverage
Parental leave
Back-up child care
+1
Software Engineer, Distributed Systems & Networking
Software Engineer, Distributed Systems & Networking

Ford • Palo Alto (CA)

On-site
USD 139,000 - 233,000
Medical, dental, vision
Parental leave
Vehicle discount program
+1