Sr Software Engineer - Reliability Engineering

Cox Automotive

New York (NY)

Hybrid

USD 122,000 - 203,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Hybrid work option
Incentive program

Job summary

Cox Automotive in the USA is seeking a Sr Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems—treating resilience and debuggability as first-class concerns.

You'll own projects end-to-end: architecture → code → deployment → production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring; work on a team

Qualifications

  • 5+ years software engineering, platform engineering, or infrastructure engineering experience.
  • Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
  • AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
  • Terraform or equivalent infrastructure-as-code experience.
  • Docker and container orchestration (Kubernetes or similar).
  • Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
  • System design thinking: can architect scalable systems and reason about trade-offs.
  • Availability for rotational on-call duties outside of standard business hours may be required.
  • Bachelor's degree in a related discipline and 4 years' experience; or master's/PhD variants acceptable.

Responsibilities

  • Design and implement SRE best practices and system health management.
  • Build observability into systems with logging, metrics, distributed tracing, and alerts.
  • Drive AWS cost optimization through architectural improvements and resource utilization.
  • Develop production-grade software and architecture with emphasis on reliability and maintainability.
  • Participate in on-call rotations; debug and resolve production incidents; document postmortems.

Skills

Python
Go
Java
System design
On-call readiness

Education

Bachelor's degree in related discipline

Tools

Terraform
Docker
Kubernetes
AWS

Job description

Company

Cox Automotive - USA

Job Family Group

Engineering / Product Development

Job Profile

Sr Software Engineer

Management Level

Individual Contributor

Flexible Work Option

Hybrid - Ability to work remotely part of the week

Travel %

Yes, 5% of the time

Work Shift

Day

Compensation

Compensation includes a base salary in the range of $121,800.00 - $203,000.00. The base salary may vary within the anticipated base pay range based on factors such as the ultimate location of the position and the selected candidate's knowledge, skills, and abilities. Position may be eligible for additional compensation that may include an incentive program.

Job Description

We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems-treating resilience and debuggability as first‑class concerns. You'll own projects end‑to‑end: from architecture → code → deployment → production. You'll split time between infrastructure-as‑code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily.

What You’ll Do:
SRE Best Practices & System Health Management
  • Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
  • Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
  • Own system health: proactively maintain application performance, minimize downtime, ensure consistent user‑experience.
  • Evolve team’s SRE standards and practices.
Application Monitoring & Observability
  • Build observability into systems: logging, metrics, distributed tracing, alert design.
  • Improve monitoring frameworks; enable faster incident detection and resolution.
  • Design dashboards and alerts that help teams understand system behavior.
  • Partner with application teams on service instrumentation.
AWS Cost Optimization
  • Drive significant reductions in cloud spend through architectural improvements and resource utilization.
  • Review infrastructure for efficiency; identify and eliminate waste.
  • Balance cost, performance, and reliability in design decisions.
Software Development & Architecture
  • Build production systems, APIs, internal tools, and automation with clean, well-tested code.
  • Design for maintainability, operational simplicity, and reliability.
  • Participate in code review and technical design discussions.
  • Mentor junior engineers on code quality and architectural thinking.
Operations & Incident Response
  • Participate in on‑call rotations; debug and resolve production incidents.
  • Conduct postmortem analysis; drive systemic improvements.
  • Develop operational procedures and runbooks.
Qualifications:
  • 5+ years software engineering, platform engineering, or infrastructure engineering experience.
  • Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
  • AWS hands‑on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
  • Terraform or equivalent infrastructure-as-code experience.
  • Docker and container orchestration (Kubernetes or similar).
  • Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
  • System design thinking: can architect scalable systems and reason about trade-offs.
  • Availability for rotational on‑call duties outside of standard business hours may be required.
  • Bachelor's degree in a related discipline and 4 years' experience in a related field. The right candidate could also have a different combination, such as a master's degree and 2 years' experience; a Ph.D. and up to 1 year of experience; or 16 years' experience in a related field.
Highly Valued:
  • Experience with observability tools (New Relic, Splunk, Prometheus).
  • Incident response experience; familiar with postmortem practices.
  • Interest in or hands‑on experience with SRE concepts (SLOs, resilience, failure modes).
  • Cost optimization mindset; has identified and eliminated cloud waste.
  • Windows and Linux system troubleshooting and performance analysis.
  • Experience with CI/CD pipelines and deployment automation.
Drug Testing

To be employed in this role, you'll need to clear a pre-employment drug test. Cox Automotive does not currently administer a pre-employment drug test for marijuana for this position. However, we are a drug‑free workplace, so the possession, use or being under the influence of drugs illegal under federal or state law during work hours, on company property and/or in company vehicles is prohibited.

Benefits

The Company offers eligible employees the flexibility to take as much vacation with pay as they deem consistent with their duties, the company's needs, and its obligations; seven paid holidays throughout the calendar year; and up to 160 hours of paid wellness annually for their own wellness or that of family membe

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox • Village of North Hills (NY)

On-site
USD 121,000 - 203,000
Paid vacation
7 holidays per year
Up to 160 hours wellness time
+6
Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox Automotive Inc. • Village of North Hills (NY)

On-site
USD 122,000 - 203,000
Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox Enterprises • Village of North Hills (NY)

On-site
USD 150,000 - 185,000
LEAD SITE RELIABILITY ENGINEER
LEAD SITE RELIABILITY ENGINEER

CAI Cox Automotive Corp Svcs., LLC • Austin (TX)

Hybrid
USD 120,000 - 150,000
Health care insurance (medical, dental, vision)
Retirement planning (401(k))
Paid days off
Senior Systems Engineer
Senior Systems Engineer

CAI Cox Automotive Corp Svcs., LLC • Atlanta (GA)

On-site
USD 92,300 - 153,900
Health insurance
401(k) retirement plan
Paid time off
Software Engineer II - 20269
Software Engineer II - 20269

Cox Automotive Inc. • Atlanta (GA)

On-site
USD 89,000 - 134,000
Sr Software Engineer
Sr Software Engineer

ApplyMint • Atlanta (GA)

Hybrid
USD 102,000 - 169,000
Flexible work option
Comprehensive benefits
Software Engineer II - 20269
Software Engineer II - 20269

CAI Cox Automotive Corp Svcs., LLC • Atlanta (GA)

Hybrid
USD 89,000 - 134,000
Hybrid work option
Incentive program
Health insurance
+1
Senior SRE Engineer: Reliability, Observability & Cloud
Senior SRE Engineer: Reliability, Observability & Cloud

Cox Automotive Inc. • Village of North Hills (NY)

On-site
USD 122,000 - 203,000
Principal Software Engineer
Principal Software Engineer

ApplyMint • Atlanta (GA)

Hybrid
USD 163,000 - 272,000
Hybrid work schedule
Health insurance
Paid time off