Lead Site Reliability Engineer

EPAM Systems

Brasil

Presencial

BRL 250 000 - 450 000

Tempo integral

Há 3 dias
Torna-te num dos primeiros candidatos

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

EPAM Systems in Brazil seeks a Lead Site Reliability Engineer to partner with backend developers, product managers, and fellow SREs to anticipate risks and drive incident response and postmortems. You will architect and sustain monitoring, dashboards, and SLOs/SLIs with a focus on debit card processes to ensure reliability and resilience.

The role requires extensive experience with distributed systems, Kubernetes and cloud or on-prem environments, and strong automation skills.

Qualificações

  • At least 5 years of experience as a Site Reliability Engineer, DevOps Engineer, or similar role.

Responsabilidades

  • Architect and sustain monitoring, alerting, and observability systems for debit card services.

Conhecimentos

Go
Python
Java

Ferramentas

Kubernetes
Docker
Grafana
Prometheus
Datadog

Descrição da oferta de emprego

We are seeking a Lead Site Reliability Engineer to become part of our team. In this position, you'll partner closely with backend developers, product managers, and fellow SREs to spot potential risks before they escape into full-blown incidents — and when problems do surface, you'll take charge of driving the response and subsequent analysis.

Responsibilities
  • Architect and sustain monitoring, alerting, and observability systems, covering metrics, logs, and traces, to support debit card services
  • Develop and manage dashboards that display availability, latency, error rates, and other critical reliability indicators for both engineering teams and leadership
  • Establish and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets tied to essential debit card processes, such as authorization, settlement, card issuance, and dispute handling
  • Continuously oversee system health and stability, catching signs of degradation early to prevent negative effects on customers
  • Take part in on-call schedules, driving or contributing to incident resolution, root cause investigation, and blameless retrospectives
  • Team up with product and engineering groups to assess architecture from the standpoint of reliability, scalability, and resilience to failure
  • Streamline repetitive operational tasks through automation and custom tooling to cut down on manual toil
  • Perform capacity forecasting and load testing to confirm systems can handle growing transaction demands
  • Enhance the safety of releases by leveraging canary deployments, automated rollbacks, and progressive rollout techniques
  • Help develop and uphold reliability standards, runbooks, and operational documentation throughout the team
Requirements
  • At least 5 years of relevant experience working as a Site Reliability Engineer, DevOps Engineer, or Production/Infrastructure Engineer
  • A minimum of one year of experience guiding and overseeing teams
  • Practical experience using observability and monitoring platforms such as Grafana, Prometheus, Datadog, or comparable internal metrics systems
  • Solid grasp of SLOs, SLIs, error budgets, and broader reliability engineering concepts
  • Command of at least one programming language typically used for automation and tooling, such as Go, Python, or Java
  • Background working with distributed systems and awareness of common failure patterns in high-volume, low-latency settings
  • Working familiarity with container orchestration tools and infrastructure, including Kubernetes and Docker, plus experience with cloud or on-premises environments at scale
  • Exposure to incident management workflows, covering on-call duties, root cause diagnosis, and postmortem reviews
  • Solid scripting and automation capabilities aimed at minimizing operational toil, using languages such as Bash or Python
  • Practical understanding of CI/CD pipelines and secure deployment techniques, including canary releases, blue-green deployments, and rollback approaches
  • Strong communication abilities, capable of turning system performance data into understandable insights for technical and non-technical audiences alike
  • Excellent English communication skills (B2 level or higher)
Nice to have
  • Background working within payments, fintech, or other tightly regulated, transaction-sensitive industries
  • Understanding of PCI-DSS or similar financial industry compliance and security standards
  • Exposure to chaos engineering or fault-injection methodologies for resilience testing
  • Experience with database reliability topics, such as query optimization, replication, and failover, within transactional systems
  • Background developing or supporting internal tools and platforms designed for large-scale observability
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems, Inc. • São Paulo

Presencial
BRL 180 000 - 320 000
Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certificates
Senior Reliability and Platform Engineer
Senior Reliability and Platform Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 300 000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Híbrido
BRL 180 000 - 260 000
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Senior SRE
Senior SRE

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000
Senior DevOps Analyst – SRE, Kubernetes, Cloud
Senior DevOps Analyst – SRE, Kubernetes, Cloud

Jobtailor • São Paulo

Presencial
BRL 90 000 - 130 000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • São Paulo

Teletrabalho
BRL 120 000 - 180 000
Home office setup
Competitive compensation
Career growth plans
+4