Senior Site Reliability Engineer

EPAM Systems, Inc.

Warszawa

Hybrid

PLN 170,000 - 320,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Multisport
Shopping vouchers
Referral bonuses

Job summary

EPAM Systems, Inc. is seeking a Senior Site Reliability Engineer to own reliability of production AI systems, bridging deployment and operability on Azure with cost, security, and quality baked in.

You will lead end-to-end deployments, build LLM-aware observability, define SLOs, manage incidents, drive FinOps for AI workloads, and ensure security posture across the lifecycle while collaborating with pods and architectural teams.

Qualifications

  • 5+ years of experience operating cloud production systems.
  • Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling.
  • Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation.
  • Knowledge of FinOps basics for AI workloads.
  • Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code.

Responsibilities

  • Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source
  • Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards
  • Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets
  • Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward
  • Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively
  • Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle
  • Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated
  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect

Skills

Azure IaaS/PaaS
Terraform/Bicep
CI/CD tooling
LLM tracing
Container orchestration
Python scripting
Bash scripting
FinOps basics

Job description

We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.ResponsibilitiesOwn deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from sourceBuild LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboardsDefine and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgetsRun incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterwardManage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactivelyKeep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycleShape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operatedFeed operational patterns, failure modes, and cost learnings back to the pods and the ArchitectRequirements5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work awayExpertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD toolingProficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automationKnowledge of FinOps basics for AI workloadsFamiliarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation codeWe offerWe gather like-minded people:Top tech minds driving innovation in AI, cloud and digital platform modernizationSupportive team and agile, startup-like cultureHybrid by design mode and opportunity to work remotely within PolandChance to work abroad for up to 60 days annuallyBusiness-driven relocation opportunitiesWe provide growth opportunities:Career development programsThought leadership, mentoring, soft skills and well-being programsCertification (Anthropic, Gemini, GCP, Azure, AWS)English classesWe cover it all:Stable payParticipation in the Employee Stock Purchase Plan with a 15% discountBenefits package (health insurance, multisport, shopping vouchers)Referral bonuses up to $2,000Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and moreCorporate, social and well-being eventsPlease, note:Benefits listed above are available to employees onlyWe are open for working with Contractors. Terms of B2B cooperation agreements are agreed individuallyWe will reach out to selected candidates exclusivelyEPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems, Inc. • Wrocław

Hybrid
PLN 180,000 - 300,000
Health insurance
Multisport
Shopping vouchers
+2
Azure Systems Architect
Azure Systems Architect

EPAM Systems, Inc. • Poland

On-site
PLN 200,000 - 350,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+4
Lead AI Security Engineer
Lead AI Security Engineer

EPAM Systems, Inc. • Polska

Hybrid
PLN 300,000 - 420,000
Health insurance
Employee stock plan
Relocation opportunities
+1
Senior DevOps Engineer (Azure AI / GenAI Solutions)
Senior DevOps Engineer (Azure AI / GenAI Solutions)

EPAM Systems, Inc. • Polska

Hybrid
PLN 180,000 - 260,000
Health insurance
Multisport
Shopping vouchers
+1
Senior C# Azure Engineer
Senior C# Azure Engineer

EPAM Systems, Inc. • Wrocław

On-site
PLN 240,000 - 400,000
Health insurance
Multisport
Shopping vouchers
Senior .NET Engineer with AI
Senior .NET Engineer with AI

EPAM Systems, Inc. • Łódź

Hybrid
PLN 180,000 - 260,000
Health insurance
Employee Stock Purchase Plan
Certification programs
+2
Senior .NET Engineer with AI
Senior .NET Engineer with AI

EPAM Systems, Inc. • Poznań

Hybrid
PLN 180,000 - 320,000
Health insurance
English classes
Employee stock purchase plan
+2
Senior DevOps Engineer with Azure
Senior DevOps Engineer with Azure

EPAM Systems, Inc. • Wrocław

Hybrid
PLN 180,000 - 240,000
Health insurance
Multisport
Shopping vouchers
+5
Senior .NET Engineer with AI
Senior .NET Engineer with AI

EPAM Systems, Inc. • Warszawa

Hybrid
PLN 180,000 - 260,000
Health insurance
English classes
Relocation opportunities
+1
Lead Azure AI Security Engineer
Lead Azure AI Security Engineer

EPAM Systems • Łódź

On-site
PLN 250,000 - 400,000
Hybrid by design
Remote within Poland
Relocation opportunities
+2