SRE / Production Engineering

Infosys

Bengaluru

On-site

INR 3,500,000 - 7,000,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Infosys is seeking a senior SRE/Production Engineer to lead reliability across critical services. You will drive Kubernetes, Terraform, CI/CD, and observability practices to ensure high availability and scalable systems.

You will own incident management, runbooks, SLO/SLI frameworks, and automation initiatives, mentoring teams to reduce toil and MTTR while improving release readiness across cloud environments.

Qualifications

  • 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
  • Strong hands-on experience in cloud operations, incident management, and production support for critical services.
  • Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
  • Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices.

Responsibilities

  • Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
  • Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
  • Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements.
  • Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
  • Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
  • Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes.
  • Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
  • Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
  • Standardize operational tooling and workflows to improve service health, change success rate, and MTTR.
  • Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
  • Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives.

Skills

SRE
Production engineering
Reliability engineering
Kubernetes
Terraform
CI/CD pipelines
Observability
Prometheus
Grafana
Incident management
Automation
Cloud operations

Education

BTECH/MTECH/MCA/MSC in CS/IT

Tools

Kubernetes
Prometheus
Grafana

Job description

SRE, Production engineering, Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)

Key Responsibilities: Reliability & Production Ownership

  • Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
  • Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
  • Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements. Incident Management & Operational Excellence
  • Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
  • Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
  • Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes. Cloud Operations & Automation
  • Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
  • Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
  • Standardize operational tooling and workflows to improve service health, change success rate, and MTTR. Leadership & Collaboration
  • Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
  • Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives. Minimum Qualifications:
  • BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).
  • 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
  • Strong hands‑on experience in cloud operations, incident management, and production support for critical services.
  • Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
  • Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices. Preferred Qualifications:
  • Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.
  • Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.
  • Expertise in release/change management practices that improve deployment safety and reduce production incidents.
  • Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.
  • Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer / Production Engineer
Site Reliability Engineer / Production Engineer

Infosys • Bengaluru

On-site
INR 5,000,000 - 7,000,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Bengaluru

On-site
INR 1,800,000 - 2,400,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Saika Technologies Inc. • Hyderabad, Bengaluru

Hybrid
INR 3,000,000 - 4,200,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Skillventory • Kamrup Metropolitan

On-site
INR 3,500,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

C1X • Chennai District

On-site
INR 1,800,000 - 3,200,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000