Chief Site Reliability Engineer

EPAM Systems

Argentina

On-site

ARS 1,800,000 - 3,000,000

Full time

3 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off
Upskilling programs

Job summary

EPAM Systems seeks a Chief Site Reliability Engineer to drive reliability and DevOps maturity for mission-critical platforms. You will lead resilient cloud and Kubernetes operations, improve CI/CD and release processes, and guide rapid incident response.

The role demands hands-on leadership, top-level SRE delivery, and collaboration with global engineering teams. The ideal candidate has extensive SRE and cloud experience, strong Python automation skills, and proven ability to scale

Qualifications

  • 7+ years of site reliability engineering experience
  • 7+ years of cloud platform experience (AWS/Azure)
  • Proven leadership to set direction and raise standards
  • Enterprise-scale release management and CI/CD automation experience
  • Advanced Python programming for tooling and automation
  • Hands-on Kubernetes experience with clusters and workloads
  • DevSecOps platform experience with GitLab preferred
  • Strong analytical and problem-solving skills
  • Upper-Intermediate English proficiency (B2)

Responsibilities

  • Lead reliability strategy for critical infrastructure
  • Design resilient cloud architectures on AWS and Azure
  • Improve CI/CD pipelines, source control, and release workflows
  • Automate infrastructure and operations with Python
  • Harden platform foundations across networking, compute, security, IAM
  • Operate and evolve Kubernetes usage patterns
  • Coordinate on-call response and incident resolution
  • Investigate root causes and drive corrective actions
  • Collaborate with teams to deliver high-leverage solutions
  • Define and track reliability metrics and controls

Skills

Python programming
Leadership
Analytical thinking
English (B2)

Tools

AWS
Azure
GitLab
Kubernetes

Job description

We are seeking a Chief Site Reliability Engineer to strengthen reliability and DevOps maturity across mission‑critical platforms in a fast‑changing environment. You will drive resilient cloud and Kubernetes operations, improve CI/CD and release processes, and lead rapid incident response.

Responsibilities
  • Lead reliability strategy for critical infrastructure to enable rapid business change
  • Design resilient cloud architectures and operational patterns across AWS and Azure
  • Build and improve CI/CD pipelines, source control practices, and release workflows
  • Automate infrastructure and operational tasks using Python to reduce toil and risk
  • Harden platform foundations across networking, compute, security, IAM, and configuration automation
  • Operate and evolve Kubernetes usage patterns to improve stability and delivery speed
  • Coordinate on‑call response and resolve business‑critical incidents under time pressure
  • Investigate systemic issues, perform root‑cause analysis, and drive corrective actions
  • Partner with engineering teams to deliver high‑leverage solutions over quick fixes
  • Define and track reliability metrics and operational controls to measure maturity
Requirements
  • Extensive site reliability engineering experience (7+ years) supporting critical infrastructure
  • Strong cloud platform experience (7+ years) with leading providers, including Amazon Web Services and Microsoft Azure
  • Proven leadership skills to set direction, influence stakeholders, and raise engineering standards
  • Enterprise‑scale release management experience delivering reliable software delivery processes
  • Deep CI/CD expertise across pipelines, source control, and infrastructure automation
  • Advanced Python programming skills to build tooling and automation
  • Solid Kubernetes experience as a developer working with clusters and workloads
  • Hands‑on DevSecOps platform experience with GitLab preferred
  • Strong analytical skills for complex problem‑solving and strategic decision‑making
  • Upper‑Intermediate English proficiency (B2, Upper‑Intermediate)
Nice to have
  • Amazon Web Services certification or equivalent hands‑on expertise
  • Microsoft Azure certification or equivalent hands‑on expertise
  • AI Architecture experience applied to reliability and operational decision‑making
  • AI Solution Engineering experience supporting platform automation and operations
  • Gen AI Solutions Development exposure for operational tooling or incident workflows
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award‑winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Argentina

On-site
ARS 1,400,000 - 2,000,000
Healthcare benefits
Paid time off
Upskilling programs
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • Argentina

On-site
ARS 136,430,000 - 181,906,000
International projects with top brands
Global teams of diverse peers
Healthcare benefits
+6
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems • Argentina

On-site
ARS 2,400,000 - 4,200,000
Healthcare benefits
Unlimited LinkedIn Learning access
Global career opportunities
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Argentina

On-site
ARS 135,892,000 - 196,289,000
Flexible work format
Education reimbursement
Mentorship program
+2
Chief SRE: Lead Cloud Reliability & DevOps
Chief SRE: Lead Cloud Reliability & DevOps

EPAM Systems • Argentina

On-site
ARS 1,800,000 - 3,000,000
Healthcare benefits
Paid time off
Upskilling programs
Senior Cloud Engineer (AWS)
Senior Cloud Engineer (AWS)

EPAM Systems • Argentina

On-site
ARS 2,000,000 - 4,500,000
Healthcare benefits
Unlimited LinkedIn Learning access
Paid time off and sick leave
+3
Senior Sre - Ai Reliability & Platform Engineer
Senior Sre - Ai Reliability & Platform Engineer

T2S - Group International . Your Partner In Executive Search • Córdoba

Remote
ARS 183,450,000 - 259,887,000
Professional growth
Competitive compensation
A selection of exciting projects
+1
Senior Devsecops Engineer — Aws Cloud Platform & Sre
Senior Devsecops Engineer — Aws Cloud Platform & Sre

Talan • Partido de Quilmes

On-site
ARS 212,695,000 - 319,042,000
Professional growth
Competitive compensation: USD-based
Exciting projects
+1
Tech Lead (.Net)
Tech Lead (.Net)

Skydropx - Frenet • Buenos Aires

Hybrid
ARS 213,233,000 - 289,387,000
Professional growth
USD-based pay
Exciting projects
+1
Lead Infrastructure Engineer (Automation & Orchestration)
Lead Infrastructure Engineer (Automation & Orchestration)

EPAM Systems • Argentina

Hybrid
ARS 136,430,000 - 227,383,000
Healthcare benefits
Professional development
Global career opportunities
+2