We are seeking a Lead Site Reliability Engineer to become part of our team. In this position, you'll partner closely with backend developers, product managers, and fellow SREs to spot potential risks before they escalate into full-blown incidents — and when problems do surface, you'll take charge of driving the response and subsequent analysis.ResponsibilitiesArchitect and sustain monitoring, alerting, and observability systems, covering metrics, logs, and traces, to support debit card servicesDevelop and manage dashboards that display availability, latency, error rates, and other critical reliability indicators for both engineering teams and leadershipEstablish and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets tied to essential debit card processes, such as authorization, settlement, card issuance, and dispute handlingContinuously oversee system health and stability, catching signs of degradation early to prevent negative effects on customersTake part in on-call schedules, driving or contributing to incident resolution, root cause investigation, and blameless retrospectivesTeam up with product and engineering groups to assess architecture from the standpoint of reliability, scalability, and resilience to failureStreamline repetitive operational tasks through automation and custom tooling to cut down on manual toilPerform capacity forecasting and load testing to confirm systems can handle growing transaction demandsEnhance the safety of releases by leveraging canary deployments, automated rollbacks, and progressive rollout techniquesHelp develop and uphold reliability standards, runbooks, and operational documentation throughout the teamRequirementsAt least 5 years of relevant experience working as a Site Reliability Engineer, DevOps Engineer, or Production/Infrastructure EngineerA minimum of one year of experience guiding and overseeing teamsPractical experience using observability and monitoring platforms such as Grafana, Prometheus, Datadog, or comparable internal metrics systemsSolid grasp of SLOs, SLIs, error budgets, and broader reliability engineering conceptsCommand of at least one programming language typically used for automation and tooling, such as Go, Python, or JavaBackground working with distributed systems and awareness of common failure patterns in high-volume, low-latency settingsWorking familiarity with container orchestration tools and infrastructure, including Kubernetes and Docker, plus experience with cloud or on-premises environments at scaleExposure to incident management workflows, covering on-call duties, root cause diagnosis, and postmortem reviewsSolid scripting and automation capabilities aimed at minimizing operational toil, using languages such as Bash or PythonPractical understanding of CI/CD pipelines and secure deployment techniques, including canary releases, blue-green deployments, and rollback approachesStrong communication abilities, capable of turning system performance data into understandable insights for technical and non-technical audiences alikeExcellent English communication skills (B2 level or higher)Nice to haveBackground working within payments, fintech, or other tightly regulated, transaction-sensitive industriesUnderstanding of PCI-DSS or similar financial industry compliance and security standardsExposure to chaos engineering or fault-injection methodologies for resilience testingExperience with database reliability topics, such as query optimization, replication, and failover, within transactional systemsBackground developing or supporting internal tools and platforms designed for large-scale observabilityWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.