Senior Site Reliability Engineer (SRE)

EPAM Systems

España

On-site

PHP 5,084,000 - 7,262,000

Full time

24 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Private health insurance
EPAM Employee Stock Purchase Plan
100% paid sick leave
Referral Program
Professional certification
Language courses

Job summary

EPAM Systems is seeking a Senior Site Reliability Engineer (SRE) for remote work in Spain. You will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable systems in the financial domain.

You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting SLOs. The role offers influence on system design for reliability and performance in a global delivery context, leveraging cloud

Qualifications

  • Bachelor's degree in Computer Science, Engineering or related field.
  • Experience with cloud environments (AWS/GCP/Azure).
  • Knowledge of SRE principles (SLOs/SLIs, postmortems, automation).
  • Proficiency in Python or similar scripting languages.

Responsibilities

  • Define and maintain SLOs/SLIs and error budgets for critical services.
  • Collaborate with cross-functional teams to embed reliability into design.
  • Automate operational tasks to reduce toil and improve performance.
  • Troubleshoot incidents and maintain robust monitoring/observability.
  • Plan capacity for high availability and resiliency.
  • Contribute to postmortems and continuous improvement.

Skills

Cloud platforms (AWS/GCP/Azure)
SRE principles
Python scripting
Monitoring/observability
Infrastructure as Code
CI/CD
Docker/Kubernetes

Education

Bachelor's degree

Tools

Terraform
Ansible
Jenkins / GitLab CI

Job description

We're looking for a Senior Site Reliability Engineer (SRE) to join our team in Spain in a remote working mode. In this role, you will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs).

This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences.

Responsibilities
  • Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services
  • Collaborate with cross-functional teams to embed reliability into application and infrastructure design
  • Automate operational tasks to reduce manual toil and improve service performance
  • Troubleshoot and resolve infrastructure and application incidents quickly and effectively
  • Implement robust monitoring and observability systems to detect and prevent outages
  • Plan capacity and scaling strategies to ensure high availability and resiliency
  • Contribute to incident postmortems and continuous improvement initiatives
  • Support the adoption of SRE best practices across all SDLC stages
Requirements
  • Bachelor’s degree in Computer Science, Engineering or related field
  • Proven experience working in cloud environments (AWS, GCP or Azure)
  • Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation)
  • Proficiency in Python or other scripting language for automation tasks
  • Strong understanding of monitoring tools and observability frameworks
  • Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab)
  • Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes
Nice to have
  • Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions
  • Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies
  • Background in DevOps practices and agile delivery frameworks
  • Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments
We offer
  • Private health insurance
  • EPAM Employees Stock Purchase Plan
  • 100% paid sick leave
  • Referral Program
  • Professional certification
  • Language courses

EPAM is a leading digital transformation services and product engineering company with 61,700+ EPAMers in 55+ countries and regions. Since 1993, our multidisciplinary teams have been helping make the future real for our clients and communities around the world. In 2018, we opened an office in Spain that quickly grew to over 1,450 EPAMers distributed between the offices in Málaga, Madrid and Cáceres as well as remotely across the country. Here you will collaborate with multinational teams, contribute to numerous innovative projects, and have an opportunity to learn and grow continuously.

Why Join EPAM
  • WORK AND LIFE BALANCE. Enjoy more of your personal time with flexible work options, 24 working days of annual leave and paid time off for numerous public holidays.
  • CONTINUOUS LEARNING CULTURE. Craft your personal Career Development Plan to align with your learning objectives. Take advantage of internal training, mentorship, sponsored certifications and LinkedIn courses.
  • CLEAR AND DIFFERENT CAREER PATHS. Grow in engineering or managerial direction to become a People Manager, in-depth technical specialist, Solution Architect, or Project/Delivery Manager.
  • STRONG PROFESSIONAL COMMUNITY. Join a global EPAM community of highly skilled experts and connect with them to solve challenges, exchange ideas, share expertise and make friends.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Security Architect
Security Architect

EPAM Systems • España

On-site
PHP 6,536,000 - 10,167,000
Private health insurance
EPAM Employees Stock Purchase Plan
100% paid sick leave
+3
ServiceNow Platform Architect
ServiceNow Platform Architect

EPAM Systems • España

On-site
PHP 8,708,000 - 13,062,000
Private health insurance
EPAM Employees Stock Purchase Plan
100% paid sick leave
+3
Senior SRE: Cloud Reliability & Automation (Remote, Spain)
Senior SRE: Cloud Reliability & Automation (Remote, Spain)

EPAM Systems • España

On-site
PHP 5,084,000 - 7,262,000
Private health insurance
EPAM Employee Stock Purchase Plan
100% paid sick leave
+3
Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Mexico

On-site
PHP 7,519,000 - 11,278,000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+1
Senior ServiceNow Developer
Senior ServiceNow Developer

EPAM Systems • España

Hybrid
PHP 3,994,000 - 5,447,000
Private health insurance
Employee stock plan
Paid sick leave
+3
SAP Datasphere (Databricks and SAC)
SAP Datasphere (Databricks and SAC)

EPAM Systems • España

On-site
PHP 4,720,000 - 6,536,000
Private health insurance
EPAM Employees Stock Purchase Plan
100% paid sick leave
+3
Senior SAP CAP Java Developer
Senior SAP CAP Java Developer

EPAM Systems • España

On-site
PHP 4,357,000 - 6,536,000
Private health insurance
EPAM Employees Stock Purchase Plan
100% paid sick leave
+3
Senior Solution Architect Salesforce (Native Spanish Speaker)
Senior Solution Architect Salesforce (Native Spanish Speaker)

EPAM Systems • España

On-site
PHP 6,536,000 - 9,441,000
Private health insurance
EPAM Employees Stock Purchase Plan
100% paid sick leave
+3
Partner Sales Manager - Salesforce (Spain & Portugal)
Partner Sales Manager - Salesforce (Spain & Portugal)

EPAM Systems • España

On-site
PHP 5,797,000 - 7,971,000
Private health insurance
EPAM Employees Stock Purchase Plan
100% paid sick leave
+3
SAP Master Data Consultant
SAP Master Data Consultant

EPAM Systems • España

On-site
PHP 5,084,000 - 7,988,000
Private health insurance
EPAM Employees Stock Purchase Plan
100% paid sick leave
+3