Senior Site Reliability Engineer

EPAM Systems

Warszawa

Hybrid

PLN 240,000 - 360,000

Full time

21 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Relocation opportunities
Work abroad opportunities
Health insurance
Employee stock purchase plan

Job summary

EPAM Systems seeks a Senior Site Reliability Engineer to own reliability and observability for production AI systems in a hybrid model. You will own deployment end-to-end on Azure, implement infrastructure as code, and ensure environments are rebuildable from source.

You will build LLM-aware observability with traces on model calls, cost and latency dashboards, and define SLOs. You will manage on-call rotations, incident responses, and security postures across the lifecycle.

Qualifications

  • 5+ years operating cloud production systems.
  • Experience in Azure IaaS/PaaS operations.
  • Proficiency in observability stacks with LLM tracing and Python/Bash automation.
  • Knowledge of FinOps basics for AI workloads.
  • Familiarity with incident triage, runbook drafting and automation.

Responsibilities

  • Own deployment end-to-end on Azure with IaC, CI/CD, and environment management.
  • Build LLM-aware observability with traces, eval sampling, drift detection, and dashboards.
  • Define and defend SLOs for availability, latency and quality.
  • Run incident management with on-call models and blameless postmortems.
  • Monitor token economics, budgets, and capacity for AI workloads; ensure security posture.

Skills

Cloud operations
SLOs and on-call
Python scripting
FinOps for AI
Incident management
Observability
Automation
Cost optimization

Tools

Azure
Terraform
Bicep
CI/CD tooling
Kubernetes

Job description

We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.

Responsibilities
  • Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source
  • Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards
  • Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets
  • Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward
  • Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively
  • Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution’s entire lifecycle
  • Shape operability requirements before handover, ensuring they land in the pod’s definition of done, and run hypercare jointly with sign-off on what will be operated
  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect
Requirements
  • 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away
  • Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling
  • Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation
  • Knowledge of FinOps basics for AI workloads
  • Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code
We offer
  • We gather like-minded people:
    • Top tech minds driving innovation in AI, cloud and digital platform modernization
    • Supportive team and agile, startup-like culture
    • Hybrid by design mode and opportunity to work remotely within Poland
    • Chance to work abroad for up to 60 days annually
    • Business-driven relocation opportunities
  • We provide growth opportunities:
    • Career development programs
    • Thought leadership, mentoring, soft skills and well‑being programs
    • Certification (Anthropic, Gemini, GCP, Azure, AWS)
    • English classes
  • We cover it all:
    • Stable pay
    • Participation in the Employee Stock Purchase Plan with a 15% discount
    • Benefits package (health insurance, multisport, shopping vouchers)
    • Referral bonuses up to $2,000
    • Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
    • Corporate, social and well‑being events
  • Please, note:
    • Benefits listed above are available to employees only
    • We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
    • We will reach out to selected candidates exclusively

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI‑Native enterprises, driving measurable value from innovation and digital investments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Azure Cloud Engineer
Lead Azure Cloud Engineer

EPAM Systems • Poland

Hybrid
PLN 200,000 - 320,000
Hybrid by design
Remote within Poland
Relocation opportunities
+3
Lead Azure AI Security Engineer
Lead Azure AI Security Engineer

EPAM Systems • Poland

Hybrid
PLN 240,000 - 320,000
Hybrid by design remote within Poland
Relocation opportunities
Health insurance
Senior AI-Native Full-Stack Software Engineer
Senior AI-Native Full-Stack Software Engineer

EPAM Systems • Warszawa

Hybrid
PLN 300,000 - 480,000
Health insurance
Multisport
Shopping vouchers
+2
Senior AI & Platform Integration Engineer
Senior AI & Platform Integration Engineer

EPAM Systems • Poland

Hybrid
PLN 130,000 - 170,000
Hybrid work with remote Poland
International assignments up to 60+
Health insurance
+3
Senior Azure Cloud Engineer
Senior Azure Cloud Engineer

EPAM Systems • Poland

Hybrid
PLN 220,000 - 320,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+3
Site Reliability Engineer
Site Reliability Engineer

EPAM Systems • Poland

Hybrid
PLN 180,000 - 300,000
Health insurance
Multisport
Relocation support
+2
Lead AI Security Engineer
Lead AI Security Engineer

EPAM Systems • Poland

Hybrid
PLN 260,000 - 420,000
Hybrid by design
Remote within Poland
Relocation opportunities
+2
Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems • Wrocław

Hybrid
PLN 180,000 - 280,000
Health insurance
Hybrid work model
Relocation opportunities
+3
Senior DevOps Engineer (Azure AI / GenAI Solutions)
Senior DevOps Engineer (Azure AI / GenAI Solutions)

Talanto • Województwo pomorskie

Hybrid
PLN 210,000 - 320,000
Hybrid work model
Remote work within Poland
Work abroad up to 60 days annually
+2
Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems • Kraków

Hybrid
PLN 180,000 - 240,000
Hybrid / remote within Poland
Possibility to work abroad up to 60 as
Relocation opportunities
+2