EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.We are seeking a Senior SRE Engineer to join our team on a full-time, in-office basis, responsible for ensuring the reliability, scalability, and performance of critical cloud infrastructure and services.ResponsibilitiesDesign and maintain high-availability and disaster recovery strategies for critical workloadsBuild and manage cloud infrastructure using Infrastructure-as-Code toolsImplement and optimize CI/CD pipelines for automated deploymentsMonitor system performance and reliability through observability platformsDevelop scripts and automation tools to streamline operationsEnsure robust networking, security, and identity/access management practicesLead cross-team collaboration to resolve complex reliability challengesDrive continuous improvement initiatives across infrastructure and platform automationRequirements4-8 years of overall experience in IT4+ years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure rolesExpertise in AWS services such as EC2, S3, RDS, IAM, VPC, and LambdaKnowledge of Infrastructure-as-Code using Terraform, AWS CDK, or CloudFormationBackground in CI/CD tools such as Jenkins, GitHub Actions, or GitLab CIProficiency in containerization and orchestration technologies including Docker, Kubernetes, and ECS/EKSCompetency in monitoring and observability tools such as Datadog, New Relic, Prometheus, Grafana, ELK, and CloudWatchSkills in scripting or programming languages such as Python, Bash, or GoUnderstanding of networking, security, and identity/access management in cloud environmentsExcellent communication, problem-solving, and leadership skills with the ability to influence across teamsNice to haveAWS or other Cloud Certification such as Solutions Architect or DevOps EngineerFamiliarity with AIOps, Serverless Architectures, and event-driven systemsUnderstanding of FinOps practices and cost optimization frameworks, along with SaaS monitoring tools such as Sumo Logic and PagerDutyExposure to Atlassian tools including Jira, Confluence, and Bitbucket, plus experience with SQL/NoSQL databasesShowcase of leading cross-functional reliability initiatives or platform-wide automation projectsWe offerOpportunity to work on technical challenges that may impact across geographiesVast opportunities for self-development: online university, knowledge sharing opportunities globally, learning opportunities through external certificationsOpportunity to share your ideas on international platformsSponsored Tech Talks & HackathonsUnlimited access to LinkedIn learning solutionsPossibility to relocate to any EPAM office for short and long-term projectsFocused individual developmentBenefit package:Health benefitsRetirement benefitsPaid time offFlexible benefitsForums to explore beyond work passion (CSR, photography, painting, sports, etc.)