Senior Reliability Engineer (Remote)

Kohl's

Menomonee Falls (WI)

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kohl's is seeking a Senior Reliability Engineer to ensure the resilience and availability of its systems and applications. In this role, you will collaborate with development teams, conduct risk assessments, and implement robust monitoring mechanisms.

Candidates should have a Bachelor's Degree and 4+ years of software development experience, with strong programming skills in languages such as Java or Python. Experience with cloud platforms and automation techniques is essential.

Qualifications

  • 4+ years of experience in software development.
  • Strong programming skills in one or more languages.
  • In-depth knowledge of systems architecture and network fundamentals.
  • Experience with multi-region application troubleshooting.

Responsibilities

  • Drive error budget and Service Level Objective adoption.
  • Establish consistent practices for operational excellence.
  • Identify repeated toil for automation opportunities.
  • Mentor and assist engineers on the team.

Skills

Java
Python
Go
Node.js
Cloud Monitoring
Automation

Education

Bachelor's Degree in MIS, Computer Science or related field

Tools

Docker
Kubernetes
AWS
Azure
GCP
CloudWatch
Grafana
Prometheus
Terraform

Job description

About the Role

As Senior Reliability Engineer, you will ensure the resilience and availability of Kohl’s systems and applications, collaborate closely with development teams, contribute to architectural designs, conduct risk assessments and design for failure, and implement robust monitoring and failover mechanisms.

What You’ll Do
  • Drive error budget and Service Level Objective (SLO) adoption across products
  • Drive incident response efforts, perform root cause analysis and implement preventative measures to enhance system reliability
  • Establish consistent practices that elevate Kohl’s operational excellence through automation and process improvements
  • Follow software lifecycle and drive reliability, observability, and efficiency across product teams within an assigned domain
  • Identify repeated toil and find opportunities for automation and risk reduction
  • On-call on a rotation to respond to production incidents and conduct blameless retrospectives and root‑cause analyses (RCAs) to drive a culture of continuous improvements
  • Proactively identify failures before they cause outages using chaos engineering techniques such as edge cases, failure modes and design review
  • Advise on capacity planning and provide continuous assessments on systems behavior and consumption
  • Work with product managers to identify and prioritize work for reliability best practices (i.e., leveraging SLIs/SLOs/Error Budgets)
  • Mentor and assist engineers on the team
  • Additional tasks may be assigned
Required
What Skills You Have
  • Bachelor’s Degree or equivalent in MIS, Computer Science or related field
  • 4+ years of experience in software development
  • Strong programming skills in one or more languages (Java, Python, Go or Node.js)
  • In-depth knowledge of systems architecture, operating system internals and network fundamentals
  • In-depth knowledge of application design patterns, event‑driven architecture, database schemas, and testing strategies
  • Experience with multi‑region application troubleshooting and performance tuning
  • Working experience with one cloud platform (GCP, AWS, or Azure)
  • Working experience with monitoring techniques and tools (e.g., CloudWatch, Grafana, Prometheus, OpenTelemetry, Tracing)
Preferred
  • In‑depth knowledge of containerization and container orchestration (e.g., Docker, Kubernetes, Rancher)
  • Experience with one or more configuration management systems (e.g., Chef, Ansible, Puppet)
  • Passion for and experience with AI and ML methodologies (MLOps)
  • Experience writing Infrastructure as code (e.g., Terraform, OpenTofu)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Reliability Engineer: Resilience & Automation
Senior Reliability Engineer: Resilience & Automation

Kohl's • Menomonee Falls (WI)

On-site
USD 90,000 - 120,000
Reliability Engineer
Reliability Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Senior Information Security Engineer (Remote)
Senior Information Security Engineer (Remote)

Kohl's • Menomonee Falls (WI)

On-site
USD 90,000 - 120,000
Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox Enterprises • Village of North Hills (NY)

On-site
USD 150,000 - 185,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Engineer 2 – Site Reliability Engineering
Senior Engineer 2 – Site Reliability Engineering

Jobtailor • Seattle (WA)

On-site
USD 180,000 - 240,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000