Senior Systems Engineer

Cox Automotive Inc.

Atlanta (GA)

On-site

USD 92,300 - 153,900

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cox Automotive Inc. seeks a Senior Systems Engineer for Observability & Resilience to architect scalable patterns that help engineering teams observe, detect, and respond faster.

The role emphasizes AI-driven detection, automation, and enterprise tool adoption across CloudWatch, New Relic, SolarWinds, and Splunk. You will build spec-driven operational standards, own relationships with engineering teams, and develop automated responses to reduce time to notice and time to resolution, evolving our

Qualifications

  • Bachelor's degree in a related discipline.
  • 4+ years of related experience.
  • Hands-on with enterprise tooling such as CloudWatch, New Relic, SolarWinds, Splunk, PagerDuty, and ServiceNow ITOM.

Responsibilities

  • Deliver AI-driven observability capabilities and automation.
  • Build and champion patterns for observability adoption across engineering teams.
  • Establish spec-driven standards for operations and resilience.

Skills

Observability patterns
Resilience engineering
AI/automation
CloudWatch
New Relic
PagerDuty

Education

Bachelor's degree

Tools

CloudWatch
New Relic
SolarWinds
Splunk
PagerDuty
ServiceNow ITOM

Job description

Senior Systems Engineer - Observability & Resilience
About The Team

We're refining our Observability and Resilience team that architects and scales intelligent observability patterns across Cox Automotive; enables automated resilience. This isn't traditional monitoring-we create, curate and govern patterns that make it efficient for our engineering organizations to be observable and resilient. To that end we write code that surfaces problems before they become outages, build AI-driven detection and response systems, and enable rapid resolution. This is a ground-floor opportunity to transform the team from implementers of monitoring into technology leaders, true experts in their trade with deep relationship with our engineering organizations.

About The Position

We are seeking highly skilled system engineers for the observability and resilience team. Individuals should be well versed in a variety of technical infrastructure and engineering capabilities and keenly focused on understanding and improving observability and resilience. We are focused on AI, automation and determined to continuously drive efficiencies in our incident response. This individual should be focused on establishing spec driven standards for operations generally and for observability specifically. You will form strong relationship with our engineering teams and build, curate and champion effective patterns for observability adoption. We currently work with Cloudwatch, New Relic, Solarwinds, Splunk and rely heavily on Pager Duty and Service Now ITOM for ingestion. We are not seeking users of these tools. We are seeking thought leaders, owners, developers, enablers - emphasis on back end, adoption and user experience. It is our vision to transform the team from implementers of monitoring to technology leaders; experts in their trade. We will be experts in our tools and drive innovation and adoption. This role will deliver new capabilities, using AI to develop new tools and new ways to seamlessly adopt the enterprise toolset. You will champion new tools and methods. You will develop and own relationships with engineering teams. You will ensure that our key assets are fully observable and you will help develop automated responses to common issues driving down the time to notice an issue, and crushing the time between recognition and resolution using innovative capabilities. You will be in rapidly changing, uncertain conditions. AI is changing everything. You must thrive in a challenging and ever evolving environment. We are in the greatest time ever to be a serious, battle-hardened operator. For the first time, we can specify how the operations harness is to be used while software is being developed. If you are excited about that and seeking a challenge creating, enabling and championing world class observability and response, bring it!!!

We are seeking a skilled and detail-oriented IT Monitoring Engineer to join our team. In this role, you will leverage tools such as ServiceNow ITOM, SolarWinds, and New Relic to ensure seamless monitoring and automation of our IT infrastructure. You will be responsible for administering monitoring tools, writing synthetic testing scripts, automating alert responses, and documenting monitoring solutions. This is a critical role that requires strong technical expertise, communication skills, and the ability to collaborate with cross-functional teams.

Who You Are
  • Deliver new capabilities, using AI to develop agentic flows and new ways to seamlessly adopt the enterprise toolset
  • Champion new tools and methods across the organization
  • Develop and own relationships with engineering teams
  • Ensure that our key assets are fully observable
  • Build automated responses to common issues, driving down the time to notice an issue and crushing the time between recognition and resolution
  • Establish spec-driven standards for operations and observability
  • Reimagine how we detect, triage, and respond to incidents at scale
  • Experiment with new approaches to observability, monitoring, and alerting
  • Define what modern observability and response engineering looks like for our organization
About You
  • Bachelor's degree in a related discipline and 4 years' experience in a related field. The right candidate could also have a different combination, such as a master's degree and 2 years' experience; a Ph.D. and up to 1 year of experience; or 16 years' experience in a related field
  • Focus on Service Now ITOM, Discovery, CSDN as the hub of observability and reactivity.
  • Professional experience optimizing the integration and flow of monitoring and ITIL systems.
  • Hands-on with enterprise tooling such as CloudWatch, New Relic, SolarWinds, Splunk, PagerDuty, and ServiceNow ITOM-as a builder and integrator
  • Professional experience with writing synthetic tests in Python, Ruby or JavaScript for Playwright Puppeteer or Selenium.
  • Distributed systems expertise and understanding of failure modes
  • Deep observability experience-instrumentation, metrics, logs, traces, and alerting at scale
  • Experience building internal platforms, developer tools, or automation that scales
  • Git/version control and CI/CD pipeline experience
  • Infrastructure as code and API design experience
  • Track record eliminating toil through intelligent automation
  • Production ownership experience (on-call, incident response, observability)
  • Systems thinking mindset-understanding how components interact at scale Eager to dig into problems and bring proposed solutions to group discussion
  • Open to feedback and able to creatively adapt multiple ideas into solutions
  • Strong technical writing including high and low-level diagramming techniques
  • Analytical skills and careful attention to detail
  • Availability for rotational on-call duties outside of standard business hours may be required
Why This Role Is Different
You're a key player transforming a team:

You will create and develop key relationships with our engineering teams. You will drive a roadmap to enable and govern solid patterns for observability and resilience.

AI at the Forefront

Work with cutting-edge LLM technology to solve real production and observability problems.

Spec-Driven Operations

For the first time, we can specify how the operations harness is to be used while software is being developed. Help define that standard.

Leadership Exposure

Grow into technical acumen, work with leadership across all levels, and shape observability and reliability strategy.

USD 92,300.00 - 153,900.00 per year

Compensation

Compensation includes a base salary in the range of $92,300.00 - $153,900.00. The base salary may vary within the anticipated base pay range based on factors such as the ultimate location of the position and the selected candidate's knowledge, skills, and abilities. Position may be eligible for additional compensation that may include an incentive program.

Benefits

The Company offers eligible employees the flexibility to take as much vacation with pay as they deem consistent with their duties, the company's needs, and its obligations; seven paid holidays throughout the calendar year; and up to 160 hours of paid wellness annually for their own wellness or that of family members. Employees are also eligible for additional paid time off in the form of bereavement leave, time off to vote, jury duty leave, volunteer time off, military leave, and parental leave.

EOE, including disability/vets

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Engineer
Senior Systems Engineer

Cox • Atlanta (GA)

On-site
USD 92,000 - 154,000
Vacation with pay
Seven paid holidays
Paid wellness up to 160 hours
Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)

O'Neil Digital Solutions, LLC • Los Angeles (CA)

On-site
USD 115,000 - 125,000
10% annual bonus target
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)

Data Analysis Incorporated • Los Angeles (CA)

On-site
USD 115,000 - 125,000
10% yearly bonus
Dynamic workplace environment
Sr Observability Engineer
Sr Observability Engineer

IT Associates • Irvine (CA)

Hybrid
USD 150,000 - 210,000
Principal Consultant - Intelligent Operations
Principal Consultant - Intelligent Operations

Medium • Chicago (IL)

On-site
USD 235,000 - 265,000
Medical, Dental, and Vision Insurance
401(k)
Paid company holidays
+2
Sr. Forward Deployed Engineer
Sr. Forward Deployed Engineer

Socket.dev • San Francisco (CA)

On-site
USD 130,000 - 175,000
Senior Observability Engineer — AI-Driven Reliability
Senior Observability Engineer — AI-Driven Reliability

IT Associates • Irvine (CA)

On-site
Principal Consultant - Intelligent Operations
Principal Consultant - Intelligent Operations

AHEAD • United States

On-site
USD 180,000 - 260,000
Medical, Dental, and Vision Insurance
401(k)
Paid company holidays
+2
Director, Professional Services
Director, Professional Services

LogicMonitor • Austin (TX)

On-site
USD 163,000 - 240,000
Health insurance
401K with company match
Unlimited vacation policy
+1