Site Reliability Engineer / Production Engineer

Infosys

Bengaluru

On-site

INR 5,000,000 - 7,000,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Infosys is seeking a Senior SRE/Production Engineer to lead reliability across services, focusing on Kubernetes, Terraform, and CI/CD pipelines. You will own SLOs/SLIs, incident response, and automation to reduce toil.

You will mentor engineers, drive reliability roadmaps, and collaborate with cross-functional teams to ensure high availability and scalable architectures in a cloud-first environment.

Qualifications

  • 12–14 years of experience in SRE or Reliability Engineering for large-scale systems.
  • Strong hands-on cloud operations, incident management, and production support.
  • Experience designing SLO/SLI frameworks and reliability KPIs.
  • Proven ability to drive automation and reduce toil.

Responsibilities

  • Lead production engineering practices to ensure availability and performance.
  • Define SLOs/SLIs and reliability roadmaps aligned to business priorities.
  • Own incident response processes and lead major incident bridges when needed.
  • Drive postmortems, root-cause analysis, and preventive actions.
  • Identify toil and implement automation to improve deployment safety and MTTR.
  • Mentor engineers and collaborate on reliability initiatives.

Education

BTECH / MTECH / MCA / MSC in Computer Science or related

Tools

Kubernetes
Terraform
CI/CD pipelines
Prometheus
Grafana

Job description

SRE / Production Engineering SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)

Preferred Qualifications
  • Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.
  • Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.
  • Expertise in release/change management practices that improve deployment safety and reduce production incidents.
  • Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.
  • Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.
Key Responsibilities
Reliability & Production Ownership
  • Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
  • Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
  • Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements.
Incident Management & Operational Excellence
  • Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
  • Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
  • Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes.
Cloud Operations & Automation
  • Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
  • Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
  • Standardize operational tooling and workflows to improve service health, change success rate, and MTTR.
Leadership & Collaboration
  • Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
  • Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives.
Minimum Qualifications
  • BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).
  • 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
  • Strong hands‑on experience in cloud operations, incident management, and production support for critical services.
  • Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
  • Demonstrated experience leading operational processes such as on‑call, postmortems, and production readiness practices.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE / Production Engineering
SRE / Production Engineering

Infosys • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Saika Technologies Inc. • Hyderabad, Bengaluru

Hybrid
INR 3,000,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Bengaluru

On-site
INR 1,800,000 - 2,400,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Skillventory • Kamrup Metropolitan

On-site
INR 3,500,000 - 7,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

C1X • Chennai District

On-site
INR 1,800,000 - 3,200,000
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000