Lead Site Reliability Engineer

EPAM Systems, Inc.

São Paulo

Presencial

BRL 180 000 - 320 000

Tempo integral

Há 3 dias
Torna-te num dos primeiros candidatos

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Vantagens oferecidas por esta oferta de emprego

Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certificates

Resumo da oferta

EPAM Systems, Inc. is seeking a Lead Site Reliability Engineer to collaborate with backend developers, product managers, and fellow SREs to detect risks before incidents occur and drive effective responses when issues arise.

You will architect and maintain monitoring, dashboards, and SLOs/SLIs to ensure service reliability for debit card processes, while mentoring teammates and reducing toil through automation.

Qualificações

  • At least 5 years of relevant experience as an SRE/DevOps/Production/Infrastructure engineer.
  • Minimum of 1 year of experience guiding and overseeing teams.
  • Practical experience with observability platforms (Grafana/Prometheus/Datadog) and container orchestration.

Responsabilidades

  • Architect and sustain monitoring, alerting, and observability systems for debit card services.
  • Develop dashboards showing availability, latency, and error rates for engineering teams and leadership.
  • Establish and monitor SLOs, SLIs, and error budgets for core debit card processes.
  • Oversee system health, drive incident response, and conduct blameless postmortems.
  • Collaborate with product/engineering on reliability, scalability, and resilience of architectures.
  • Automate repetitive operational tasks to reduce toil.
  • Perform capacity forecasting and load testing for growing transaction demands.
  • Support safe releases via canary deployments, automated rollbacks, and progressive rollout techniques.
  • Develop and uphold reliability standards, runbooks, and operational documentation.

Conhecimentos

Observability & monitoring concepts
Incident management & postmortems
Automation & scripting
Communication skills
Team leadership / mentoring

Formação académica

Bachelor's degree in Computer Science or related field

Ferramentas

Grafana
Prometheus
Datadog
Kubernetes
Docker
CI/CD tooling

Descrição da oferta de emprego

We are seeking a Lead Site Reliability Engineer to become part of our team. In this position, you'll partner closely with backend developers, product managers, and fellow SREs to spot potential risks before they escalate into full-blown incidents — and when problems do surface, you'll take charge of driving the response and subsequent analysis.ResponsibilitiesArchitect and sustain monitoring, alerting, and observability systems, covering metrics, logs, and traces, to support debit card servicesDevelop and manage dashboards that display availability, latency, error rates, and other critical reliability indicators for both engineering teams and leadershipEstablish and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets tied to essential debit card processes, such as authorization, settlement, card issuance, and dispute handlingContinuously oversee system health and stability, catching signs of degradation early to prevent negative effects on customersTake part in on-call schedules, driving or contributing to incident resolution, root cause investigation, and blameless retrospectivesTeam up with product and engineering groups to assess architecture from the standpoint of reliability, scalability, and resilience to failureStreamline repetitive operational tasks through automation and custom tooling to cut down on manual toilPerform capacity forecasting and load testing to confirm systems can handle growing transaction demandsEnhance the safety of releases by leveraging canary deployments, automated rollbacks, and progressive rollout techniquesHelp develop and uphold reliability standards, runbooks, and operational documentation throughout the teamRequirementsAt least 5 years of relevant experience working as a Site Reliability Engineer, DevOps Engineer, or Production/Infrastructure EngineerA minimum of one year of experience guiding and overseeing teamsPractical experience using observability and monitoring platforms such as Grafana, Prometheus, Datadog, or comparable internal metrics systemsSolid grasp of SLOs, SLIs, error budgets, and broader reliability engineering conceptsCommand of at least one programming language typically used for automation and tooling, such as Go, Python, or JavaBackground working with distributed systems and awareness of common failure patterns in high-volume, low-latency settingsWorking familiarity with container orchestration tools and infrastructure, including Kubernetes and Docker, plus experience with cloud or on-premises environments at scaleExposure to incident management workflows, covering on-call duties, root cause diagnosis, and postmortem reviewsSolid scripting and automation capabilities aimed at minimizing operational toil, using languages such as Bash or PythonPractical understanding of CI/CD pipelines and secure deployment techniques, including canary releases, blue-green deployments, and rollback approachesStrong communication abilities, capable of turning system performance data into understandable insights for technical and non-technical audiences alikeExcellent English communication skills (B2 level or higher)Nice to haveBackground working within payments, fintech, or other tightly regulated, transaction-sensitive industriesUnderstanding of PCI-DSS or similar financial industry compliance and security standardsExposure to chaos engineering or fault-injection methodologies for resilience testingExperience with database reliability topics, such as query optimization, replication, and failover, within transactional systemsBackground developing or supporting internal tools and platforms designed for large-scale observabilityWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Brasil

Presencial
BRL 250 000 - 450 000
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Lead PHP Software Engineer
Lead PHP Software Engineer

EPAM Systems • Brasil

Presencial
BRL 260 000 - 420 000
International projects with top brands
Global teams of diverse peers
Employee financial programs
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • São Paulo

Teletrabalho
BRL 120 000 - 180 000
Home office setup
Competitive compensation
Career growth plans
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AgileEngine, LLC • Brasil

Presencial
BRL 240 000 - 360 000
Sr. Backend Engineer - LATAM
Sr. Backend Engineer - LATAM

Insight Global • Torrinha

Presencial
BRL 200 000 - 320 000
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • São Bernardo do Campo

Híbrido
BRL 250 000 - 360 000
Professional growth
Competitive compensation
Exciting projects
+1
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Híbrido
BRL 180 000 - 260 000