Senior Site Reliability Engineer

Publicis Groupe

India

On-site

INR 2,400,000 - 4,200,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Epsilon is seeking a Senior Site Reliability Engineer to design and operate highly scalable cloud platforms with a focus on reliability, observability, and security. The role partners with Engineering, DevOps, Platform, and Security to raise operational excellence and meet performance targets.

It requires deep expertise in Linux, Kubernetes, AWS/GCP/Azure, and IaC tooling, with hands-on leadership across incident response and automation.

Qualifications

  • 5+ years of proven experience in Site Reliability/DevOps or Platform Engineering.
  • Extensive Linux administration experience across multiple sites and environments.
  • Experience with virtualization and container orchestration (VMware, Docker, Kubernetes).
  • Experience with workflow automation (e.g., N8N).
  • Strong configuration management using Ansible, Puppet, or Chef.

Responsibilities

  • Design and implement reliability strategies for distributed systems on AWS/GCP.
  • Define and measure SLIs, SLOs, and reliability metrics.
  • Build observability solutions with monitoring, logging, tracing, and alerts.
  • Lead incident response, RCA, and postmortems to improve reliability.
  • Collaborate with engineering to improve performance, resiliency, and capacity.
  • Automate operations to reduce toil and improve efficiency.
  • Guide teams on reliability-focused architecture and capacity planning.
  • Maintain Linux administration, container orchestration, and patching.
  • Monitor health, performance, and capacity with observability tools.
  • Manage ELK stack lifecycle and Kibana dashboards for operational visibility.
  • Coordinate patch validation and deployment with application owners.
  • Ensure security/compliance standards and vulnerability remediation.

Skills

SRE
Kubernetes
Docker
Terraform
Ansible
Python
Observability
Linux admin
ELK Stack
Cloud platforms
Linux troubleshooting

Education

Bachelor’s degree in Computer Science or related field

Tools

VMware vSphere
KVM/Proxmox
N8N
JIRA/ServiceNow
Grafana
Kibana
ELK Stack
PowerShell
Terraform

Job description

Overview
About Business Unit:

At the core of all that Epsilon does is a team that sets the foundation of our IT infrastructure. The team drives innovation and efficiency through pioneering technology across Epsilon's platforms and business verticals. From being the first point of contact for infrastructure needs to final deployment, the team provides end-to-end solutions for our client-facing platforms. ETS supports all aspects of revenue-generating platforms for Epsilon and sets the architectural direction for our enterprise deployments. By adopting the newest technologies, such as Cloud, Automation, and Artificial Intelligence, the team is at the front of redefining our digital business and capturing new opportunities.

Overview:

Epsilon is seeking a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives.

This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design.

Job Description: - Senior Site Reliability Engineer
  • 5+ years of proven experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
  • 5+ years of proven experience supporting enterprise Linux environments (RHEL, CentOS, Alma) across multiple sites and centralized services.
  • 3+ years of proven experience implementing virtualization solutions using VMware vSphere, KVM, Proxmox or similar technologies with high availability features (vMotion, clustering, failover).
  • 3+ years of proven experience in configuring and setting up workflow automation (e.g., N8N).
  • Experience with configuration management and automation tools (e.g., Ansible, Puppet, Chef).
  • Advanced hands-on engineering with senior ownership of infrastructure design, automation, and service delivery.
  • Core Tools: Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python), Enterprise Monitoring & Observability Platforms,n8n
  • Strong experience with Linux system administration including user management, file systems, networking, and performance tuning.
  • Experience managing containers and orchestration (Docker, Kubernetes) is a plus.
  • Familiarity with monitoring and logging tools (Kibana, Grafana, ELK stack).
  • Demonstrated experience using ITIL-based ticketing systems (e.g., JIRA, ServiceNow).
  • Strong experience supporting production systems in AWS/GCP/Azure environments.
  • Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
  • Experience operating and fixing issues on Kubernetes platforms such as EKS and/or GKE.
  • Strong knowledge of observability tools such as, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Strong scripting and automation skills using Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python comparable languages.
  • Core Tools: Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python), Enterprise Monitoring & Observability Platforms,n8n
  • Solid understanding of networking, distributed systems, cloud security, and performance optimization.
  • Primary Expertise: Linux Server Engineering, Automation, DevOps, Monitoring, and Observability
  • Linux Platforms: Linux (RHEL, CentOS, AlmaLinux)
  • Technical Depth: Advanced hands-on engineering with senior ownership of infrastructure design, automation, and service delivery.
Responsibilities
  • Design and implement reliability strategies for distributed systems running across AWS and GCP.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
  • Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.
  • Automate operational processes and reduce toil through engineering solutions.
  • Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.
  • Experience with configuration management and automation tools (e.g., Ansible, Puppet, Chef).
  • Strong experience with Linux system administration including user management, file systems, networking, and performance tuning.
  • Experience managing containers and orchestration (Docker, Kubernetes) is a plus.
  • Familiarity with monitoring and logging tools (Nagios, Prometheus, Grafana, ELK stack).
  • Experience with databases such as MySQL, PostgreSQL, or equivalent tools.
  • Develop and enforce system standards, automation practices, and recommend improvements to enhance performance and reliability.
  • Install, configure, and maintain commercial and open-source applications on Linux operating systems (e.g., RHEL, CentOS, Ubuntu).
  • Manage and support BAU (Business As Usual) operational activities.
  • Handle ServiceNow incidents, requests, and change tickets.
  • Perform Linux operating system fixing and issue resolution.
  • Conduct root cause analysis (RCA) for incidents and service disruptions.
  • Ensure infrastructure stability, reliability, and service availability.
  • Drive automation initiatives to improve operational efficiency and reduce manual effort.
  • Collaborate with multi-functional teams to support enterprise infrastructure and platform services.
  • Monitor system health, performance, and capacity using observability and monitoring tools.
  • Implement OS patching activities across Linux servers and hybrid infrastructure environments.
  • Drive vulnerability remediation efforts to address security findings and compliance requirements.
  • Coordinate with application owners and partners for patch validation and deployment activities.
  • Ensure consistency to security, compliance, and operational standards.
  • Automate routine operational tasks using , Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell / Bash Scripting.
  • Monitor system health, performance, availability, and capacity using enterprise monitoring and observability platforms.
  • Support cloud and on-premises infrastructure across AWS, Azure, GCP, VMware, and enterprise data centre environments.
  • Drive continuous service improvements through automation, standardization, and proactive problem management
  • Manage the complete ELK (Elasticsearch, Logstash, Kibana) platform lifecycle, including deployment, upgrades, maintenance, capacity planning, and decommissioning.
  • Design, develop, and maintain Kibana dashboards, visualizations, alerts, and reporting solutions for infrastructure, security, and operational monitoring.
  • Ensure log ingestion, retention, performance optimization, data availability, and platform reliability across the ELK ecosystem.
  • Fix and resolve issues related to Elasticsearch clusters, Logstash pipelines, and Kibana dashboards.
  • Hands-on experience with Dell and HPE server hardware troubleshooting, health checks, firmware management, and diagnostics using iDRAC, iLO, and vendor management tools.
Qualifications
Education, Experience, and Licensing Requirements
  • Bachelor’s degree in computer science or related field (or equivalent experience).
  • Minimum of 5+ years of proven experience in Linux system administration or IT infrastructure.
  • Experience with databases such as MySQL, PostgreSQL, or equivalent tools.
  • Experience working across multiple operating systems with strong emphasis on Linux platforms.
  • Relevant certifications preferred:
    • Red Hat Certified System Administrator (RHCSA) or Engineer (RHCE)
    • Linux+ or equivalent
  • VMware or cloud certifications (AWS, Azure) are a plus
  • Python, Ansible, Terraform, Bitbucket, PowerShell, ELK Stack, Shell Scripting (Bash, Python comparable languages.
Set Yourself Apart With
  • Experience supporting large-scale cloud migration or modernization programs.
  • Expertise in incident management and production operations for high-availability systems.
  • Experience implementing chaos engineering or resilience testing practices.
  • RHEL/AWS/AZURE/GCP certifications.
  • Experience working in Agile, DevOps, or DevSecOps environments.
Additional Information

Epsilon is a global data, technology and services company that powers the marketing and advertising ecosystem. For decades, we’ve provided marketers from the world’s leading brands the data, technology and services they need to engage consumers with 1 View, 1 Vision and 1 Voice. 1 View of their universe of potential buyers. 1 Vision for engaging each individual. And 1 Voice to harmonize engagement across paid, owned and earned channels.

Epsilon’s comprehensive portfolio of capabilities across our suite of digital media, messaging and loyalty solutions bridge the divide between marketing and advertising technology. We process 400+ billion consumer actions each day using advanced AI and hold many patents of proprietary technology, including real-time modeling languages and consumer privacy advancements. Thanks to the work of every employee, Epsilon has been consistently recognized as industry-leading by Forrester, Adweek and the MRC. Epsilon is a global company with more than 9,000 employees around the world.

Our pillars aren't just words. They're how we show up every day.

  • People centricity: We focus on employee well-being in an environment where colleagues truly care about each other.
  • Collaboration: We work together, support one another, and collectively achieve goals.
  • Growth: There are endless opportunities for growth through learning, development and career advancement.
  • Innovation: We drive progress through cutting-edge solutions and forward-thinking approaches.
  • Flexibility: We’ve created a balance between work and personal life, and we encourage adaptability to solve problems creatively.

Our values guide us to create value for our clients, our people and consumers.

  • Act with integrity
  • Work together to win together
  • Innovate with purpose
  • Respect all voices
  • Empower with accountability

These pillars and values are our foundation—shaping our culture, guiding our decisions, and uniting us in common purpose.

Epsilon is an Equal Opportunity Employer.

Epsilon is committed to promoting diversity, inclusion, and equal employment opportunities by using reasonable efforts to attract, recruit, engage and retain qualified individuals of all ethnicities and backgrounds, including, but not limited to, women, people of color, LGBTQ individuals, people with disabilities and any other underrepresented groups, traits or characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Groupe Holdings B.V • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Information Systems Engineer
Senior Information Systems Engineer

Epsilon • Bengaluru

On-site
INR 600,000 - 900,000
Lead, Systems Administration
Lead, Systems Administration

Publicis Groupe • India

On-site
INR 1,500,000 - 2,000,000
Opportunities for growth and advancement
Collaborative work environment
Diversity and inclusion initiatives
Lead, Systems Administration
Lead, Systems Administration

Publicis Groupe Holdings B.V • Bengaluru

On-site
INR 1,500,000 - 2,200,000
Employee well-being programs
Growth opportunities
Flexible work-life balance
Senior Software Engineer
Senior Software Engineer

Publicis Group • India

On-site
INR 3,500,000 - 5,200,000
Senior Network/Systems Administration Analyst
Senior Network/Systems Administration Analyst

Publicis Groupe Holdings B.V • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Network/Systems Administration Analyst
Senior Network/Systems Administration Analyst

Epsilon • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Senior Network/Systems Administration Analyst
Senior Network/Systems Administration Analyst

Publicis Groupe • India

On-site
INR 2,500,000 - 4,200,000
Senior Network/Systems Administration Analyst
Senior Network/Systems Administration Analyst

Publicis Groupe ANZ • India

On-site
INR 1,800,000 - 3,000,000
Senior Big Data Administrator
Senior Big Data Administrator

Publicis Groupe • Bengaluru

Hybrid
INR 2,500,000 - 4,000,000