Production Reliability Engineer — Windows & Kubernetes

WebMD

Newark (NJ)

On-site

USD 109,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Insurance
Paid Time Off
401(k) Retirement Plan with employer-m
Life and Disability Insurance
Employee Assistance Program (EAP)
Commuter/Transit Benefits

Job summary

WebMD is seeking a Production Engineer to join the SRE team to manage and evolve the backend stack, focusing on Windows-based infrastructure, Linux, and Kubernetes. You will automate with PowerShell and Python, own the Observability stack (Prometheus and ELK), and drive improvements across uptime and scalability.

The role requires 4+ years in IT/Infrastructure, strong troubleshooting of distributed systems, and solid communication.

Qualifications

  • Bachelor's degree in Computer Science or a related field.
  • 4+ years IT/Infrastructure experience, with 2+ years supporting high-traffic web apps.
  • Troubleshooting distributed systems, IIS and web applications.
  • Fundamental networking concepts and traffic management with F5 Big-IP.
  • Open, adaptive mindset toward Linux CLI and AI tool integration.
  • Excellent communication to convey technical concepts clearly.

Responsibilities

  • Manage Windows Server environments and Kubernetes clusters for high availability.
  • Automate tasks with PowerShell and Python to streamline workflows.
  • Own Observability stack (Prometheus, Grafana, ELK) for full-stack visibility.
  • Identify infra improvement opportunities and drive change with Ops teams.
  • Participate in on-call rotation to resolve production issues.

Skills

Windows Server OS
Linux CLI
On-call rotation experience
Strong communication
Adaptive mindset
Distributed systems troubleshooting
IIS/web apps troubleshooting

Education

Bachelor's degree in Computer Science or related field

Tools

PowerShell
Python
Go
Kubernetes
Prometheus
Grafana
ELK Stack
Puppet
Ansible
F5 Big-IP
Redis
Kafka
RabbitMQ
GitLab
Jenkins
Azure DevOps

Job description

WebMD is seeking a Production Engineer to join the SRE team to manage and evolve the backend stack, focusing on Windows-based infrastructure, Linux, and Kubernetes. You will automate with PowerShell and Python, own the Observability stack (Prometheus and ELK), and drive improvements across uptime and scalability.

The role requires 4+ years in IT/Infrastructure, strong troubleshooting of distributed systems, and solid communication.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Engineer & SRE - Windows, Kubernetes
Production Engineer & SRE - Windows, Kubernetes

WebMD LLC • Newark (NJ)

On-site
USD 109,000 - 120,000
Health Insurance
Paid Time Off
401(k) Retirement Plan with employer
+3
Production Engineer
Production Engineer

WebMD • Newark (NJ)

On-site
USD 109,000 - 120,000
Health Insurance
Paid Time Off
401(k) Retirement Plan with employer-m
+3
Production Engineer
Production Engineer

WebMD LLC • Newark (NJ)

On-site
USD 109,000 - 120,000
Health Insurance
Paid Time Off
401(k) Retirement Plan with employer
+3
Senior Production Engineer: Reliability Platform & SRE
Senior Production Engineer: Reliability Platform & SRE

Weights & Biases • Livingston (NJ)

On-site
USD 139,000 - 185,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+1
On-Prem Reliability Engineer
On-Prem Reliability Engineer

OpsMill • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer — Platform & Observability
Senior Site Reliability Engineer — Platform & Observability

Jobtailor • North Carolina

On-site
USD 180,000 - 240,000
Site Reliability Engineer — 99.9% Uptime, Kubernetes
Site Reliability Engineer — 99.9% Uptime, Kubernetes

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Reliability Engineer
Reliability Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Service Reliability Engineer
Service Reliability Engineer

NVIDIA Corporation • United States

On-site
USD 140,000 - 190,000