Senior Reliability Engineer

Jobtailor

Menomonee Falls (WI)

On-site

USD 120,000 - 160,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kohl’s seeks a skilled SRE/Software Reliability Engineer to enhance system resilience and availability. You will partner with development teams to design robust architectures, implement failover and monitoring strategies, and lead incident response with thorough root-cause analyses.

A strong background in Java/Python/Go, cloud platforms, and IaC is essential. You will drive error budgets and SLO adoption, reduce toil, and mentor engineers while collaborating across product teams to improve

Qualifications

  • Bachelor's degree or equivalent in MIS/CS or related field.
  • 4+ years of experience in software development.
  • Strong programming skills in Java, Python, Go, or Node.js.
  • Deep knowledge of systems architecture, OS internals, and networks.
  • Experience with multi-region troubleshooting and performance tuning.
  • Experience with at least one cloud platform (GCP, AWS, or Azure).
  • Experience with monitoring tools (CloudWatch, Grafana, Prometheus, OpenTelemetry).
  • Knowledge of containerization and orchestration (Docker, Kubernetes).
  • Experience with IaC (Terraform/OpenTofu).

Responsibilities

  • Ensure resilience and availability of Kohl's systems and applications.
  • Collaborate with development teams and contribute to architectural designs.
  • Lead incident response, root cause analysis, and preventative measures.
  • Drive reliability, observability, and automation across product teams.
  • Mentor engineers and advocate for best practices in SRE.
  • Participate in on-call rotations and blameless retrospectives.
  • Identify toil and opportunities for automation and risk reduction.

Skills

Java
Python
Go
Node.js
Systems Architecture
Application Design Patterns
Cloud Platforms
Docker
Kubernetes
Rancher
Terraform
OpenTelemetry
CloudWatch
Grafana
Prometheus
Multi-Region Troubleshooting
Performance Tuning

Education

Bachelor's Degree or equivalent in MIS, Computer Science or related field

Tools

Docker
Kubernetes
Rancher
Terraform
OpenTofu
Chef
Ansible
Puppet

Job description

Responsibilities
  • Ensure the resilience and availability of Kohl’s systems and applications
  • Collaborate closely with development teams
  • Contribute to architectural designs
  • Conduct risk assessments and design for failure
  • Implement robust monitoring and failover mechanisms
  • Drive error budget and Service Level Objective (SLO) adoption across products
  • Lead incident response, perform root cause analysis, and implement preventative measures
  • Establish operational excellence through automation and process improvements
  • Drive reliability, observability, and efficiency across product teams
  • Identify toil and opportunities for automation and risk reduction
  • Participate in an on‑call rotation for production incidents
  • Conduct blameless retrospectives and root-cause analyses
  • Proactively identify failures using chaos engineering, edge cases, failure modes, and design reviews
  • Advise on capacity planning and assess system behavior and consumption
  • Work with product managers to prioritize reliability best practices using SLIs, SLOs, and error budgets
  • Mentor and assist engineers on the team
  • Perform additional assigned tasks
Requirements
  • Bachelor's Degree or equivalent in MIS, Computer Science or related field
  • 4+ years of experience in software development
  • Strong programming skills in one or more languages: Java, Python, Go, or Node.js
  • In-depth knowledge of systems architecture, operating system internals, and network fundamentals
  • In-depth knowledge of application design patterns, event-driven architecture, database schemas, and testing strategies
  • Experience with multi-region application troubleshooting and performance tuning
  • Working experience with one cloud platform: GCP, AWS, or Azure
  • Working experience with monitoring techniques and tools such as CloudWatch, Grafana, Prometheus, OpenTelemetry, and tracing
  • Preferred: In-depth knowledge of containerization and container orchestration, such as Docker, Kubernetes, or Rancher
  • Preferred: Experience with configuration management systems, such as Chef, Ansible, or Puppet
  • Preferred: Passion for and experience with AI and ML methodologies (MLOps)
  • Preferred: Experience writing Infrastructure as Code, such as Terraform or OpenTofu
Core Competencies

Demonstrates expertise in systems architecture, application design patterns, and cloud platforms, with a strong focus on reliability, observability, and automation. Proficient in programming languages such as Java, Python, and Go, and experienced in implementing monitoring and failover mechanisms.

Highest-signal resume keywords
  • Systems Architecture
  • Java Programming
  • Cloud Platform Experience
  • Monitoring Techniques
  • Incident Response
Hard Skills
  • Java
  • Python
  • Go
  • Node.js
  • Systems Architecture
  • Application Design Patterns
  • Database Schemas
  • Infrastructure as Code
  • Multi-Region Application Troubleshooting
  • Performance Tuning
Soft Skills
  • Collaboration
  • Mentoring
  • Root Cause Analysis
  • Process Improvement
  • Blameless Retrospectives
Industry Keywords
  • Service Level Objective
  • Error Budget
  • Chaos Engineering
  • Capacity Planning
  • Automation
Tools & Technologies
  • GCP
  • AWS
  • Azure
  • CloudWatch
  • Grafana
  • Prometheus
  • OpenTelemetry
  • Docker
  • Kubernetes
  • Terraform
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Reliability Engineer (Remote)
Senior Reliability Engineer (Remote)

Kohl's • Menomonee Falls (WI)

On-site
USD 90,000 - 120,000
Senior Reliability Engineer (Remote)
Senior Reliability Engineer (Remote)

FashionUnited • Menomonee Falls (WI)

Remote
USD 120,000 - 150,000
Reliability Engineer (Remote)
Reliability Engineer (Remote)

Kohl's • Menomonee Falls (WI)

On-site
USD 95,000 - 135,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Reliability Observability Engineer, Level 2
Reliability Observability Engineer, Level 2

Jobtailor • Colorado

On-site
USD 120,000 - 180,000
Senior Reliability Engineer: Architect Resilience
Senior Reliability Engineer: Architect Resilience

Jobtailor • Menomonee Falls (WI)

On-site
USD 120,000 - 160,000
DevOps Engineer
DevOps Engineer

duPont REGISTRY • Miami (FL)

On-site
USD 120,000 - 180,000
Senior Reliability Engineer: Resilience & Automation
Senior Reliability Engineer: Resilience & Automation

Kohl's • Menomonee Falls (WI)

On-site
USD 90,000 - 120,000
Senior Software Engineer
Senior Software Engineer

Aditi Consulting • Urbandale (IA)

On-site
USD 120,000 - 150,000
Medical Insurance
Dental Insurance
Life Insurance
+14
Senior Forward Deployed Engineer
Senior Forward Deployed Engineer

Jobtailor • Washington

On-site
USD 140,000 - 185,000