Sr Software Engineer - Reliability Engineering

Cox Enterprises

Village of North Hills (NY)

On-site

USD 150,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cox Enterprises is seeking a Sr. Software Engineer - Reliability Engineer to design and build resilient systems with strong observability across the stack.

You will own projects end-to-end, from architecture to production, and work across infrastructure-as-code, incident response, and system improvements. You will write production-grade code in a multi-language stack, manage AWS resources, use Terraform, and mentor teammates while shaping reliability practices for a platform serving millions of

Qualifications

  • 5+ years software engineering, platform engineering, or infrastructure engineering experience.
  • Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
  • AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
  • Terraform or equivalent infrastructure-as-code experience.
  • Docker and container orchestration (Kubernetes or similar).
  • Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
  • System design thinking: architect scalable systems and reason about trade-offs.
  • Availability for rotational on-call duties outside of standard business hours may be required.
  • Bachelor’s degree in a related discipline and 4 years’ experience; alternatives allowed.

Responsibilities

  • Design and build infrastructure, observability tooling, and operational systems.
  • Own projects end-to-end: from architecture → code → deployment → production.
  • Collaborate across teams to ensure reliability and performance.
  • Mentor junior engineers on code quality and architectural thinking.
  • Participate in on-call rotations; debug and resolve production incidents.
  • Develop runbooks and postmortems to drive systemic improvements.

Skills

Python
Go
Java
System design thinking
Linux debugging
On-call experience

Education

Bachelor's degree in a related field

Tools

Terraform
AWS
Docker
Kubernetes
New Relic
Splunk
Prometheus

Job description

We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems—treating resilience and debuggability as first-class concerns. You'll own projects end-to-end: from architecture → code → deployment → production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily.

What You'll Do:

SRE Best Practices & System Health Management

  • Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
  • Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
  • Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.
  • Evolve team's SRE standards and practices.

Application Monitoring & Observability

  • Build observability into systems: logging, metrics, distributed tracing, alert design.
  • Improve monitoring frameworks; enable faster incident detection and resolution.
  • Design dashboards and alerts that help teams understand system behavior.
  • Partner with application teams on service instrumentation.

AWS Cost Optimization

  • Drive significant reductions in cloud spend through architectural improvements and resource utilization.
  • Review infrastructure for efficiency; identify and eliminate waste.
  • Balance cost, performance, and reliability in design decisions.

Software Development & Architecture

  • Build production systems, APIs, internal tools, and automation with clean, well-tested code.
  • Design for maintainability, operational simplicity, and reliability.
  • Participate in code review and technical design discussions.
  • Mentor junior engineers on code quality and architectural thinking.

Operations & Incident Response

  • Participate in on-call rotations; debug and resolve production incidents.
  • Conduct postmortem analysis; drive systemic improvements.
  • Develop operational procedures and runbooks.
Qualifications:
  • 5+ years software engineering, platform engineering, or infrastructure engineering experience.
  • Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
  • AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
  • Terraform or equivalent infrastructure-as-code experience.
  • Docker and container orchestration (Kubernetes or similar).
  • Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
  • System design thinking: can architect scalable systems and reason about trade-offs.
  • Availability for rotational on-call duties outside of standard business hours may be required.

  • Bachelor’s degree in a related discipline and 4 years’ experience in a related field. The right candidate could also have a different combination, such as a master’s degree and 2 years’ experience; a Ph.D. and up to 1 year of experience; or 16 years’ experience in a related field.

Highly Valued:

  • Experience with observability tools (New Relic, Splunk, Prometheus).
  • Incident response experience; familiar with postmortem practices.
  • Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).
  • Cost optimization mindset; has identified and eliminated cloud waste.
  • Windows and Linux system troubleshooting and performance analysis.
  • Experience with CI/CD pipelines and deployment automation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox Automotive Inc. • Village of North Hills (NY)

On-site
USD 122,000 - 203,000
Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox • Village of North Hills (NY)

On-site
USD 121,000 - 203,000
Paid vacation
7 holidays per year
Up to 160 hours wellness time
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

JPS Tech Solutions • Colorado

On-site
USD 160,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • Hapeville (GA)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000