Senior Site Reliability Engineer

EPAM Systems

Poland

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

EPAM Systems seeks a Senior Site Reliability Engineer to own reliability and operability of production AI systems. You will bridge deployment and long-term operability while embedding cost, security, and quality discipline into every solution lifecycle.

You will own end-to-end deployments on Azure, including infrastructure as code, CI/CD pipelines, and environment management, ensuring environments are rebuildable from source.

Qualifications

  • 5+ years operating cloud production systems with on-call experience
  • Azure IaaS/PaaS expertise; IaC (Terraform/Bicep)
  • LLM-aware observability; Python/Bash automation
  • FinOps basics for AI workloads
  • Familiarity with incident triage, runbooks, and postmortems

Responsibilities

  • Own reliability and observability for production AI systems
  • Define SLOs and manage cost, latency, and quality dashboards
  • Lead incident management, runbooks, and blameless postmortems
  • Oversee security posture, patching, and access reviews
  • Collaborate with pods and architects on operability patterns

Tools

Terraform/Bicep
CI/CD tooling

Job description

Responsibilities


  • We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.

  • Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source

  • Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards

  • Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets

  • Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward

  • Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively

  • Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle

  • Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated

  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect



Requirements


  • 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away

  • Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling

  • Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation

  • Knowledge of FinOps basics for AI workloads

  • Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
AI Infrastructure / MLOps Engineer — NYC
AI Infrastructure / MLOps Engineer — NYC

LaStellar Group • New York (NY)

On-site
USD 140,000 - 180,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Dairy Farmers of America • Kansas City (KS)

On-site
USD 98,000 - 139,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Strategic Staffing Solutions • Detroit (MI)

Hybrid
USD 96,000 - 152,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Kansas Ag Connection • Kansas City (KS)

On-site
USD 140,000 - 190,000
Senior SRE - Azure
Senior SRE - Azure

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 100,000 - 140,000
Azure Cloud Engineer
Azure Cloud Engineer

Embrace Software Inc • United States

Remote
USD 120,000 - 180,000
Senior MLOps Engineer
Senior MLOps Engineer

Jobot • Atlanta (GA)

On-site
USD 150,000 - 175,000
Remote work 100%
Competitive salary + bonus + equity
Medical, dental, vision insurance
+3
Senior AI Delivery & Operations Engineer
Senior AI Delivery & Operations Engineer

Jobtailor • Illinois

On-site
USD 140,000 - 190,000