MTS 2, Platform Reliability Engineer

The Networker

Bengaluru

On-site

INR 2,500,000 - 4,000,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

The Networker in Bengaluru, India, is seeking a Platform Reliability Engineer to own the reliability, operability, and evolution of our internal engineering platform. This hands-on role balances platform engineering, reliability, and intelligent automation to reduce toil and improve observability, enabling AI agents to operate safely at scale.

You will work with engineering teams to harden services, respond to incidents, and build automation that makes the platform increasingly self-managing,

Qualifications

  • 6+ years operating production platforms or large-scale distributed systems.
  • Proven incident management, on-call operations, and production debugging.
  • Strong programming skills in Java, Python, Go, or Shell.
  • Experience with observability tooling (monitoring, alerting, logging, tracing).
  • Experience building/maintaining CI/CD pipelines and release processes.
  • Experience with AI-driven automation or strong interest in this space.
  • Working knowledge of Linux-based production environments.
  • Strong communication and cross-team collaboration skills.

Responsibilities

  • Own reliability, availability, and performance of the internal platform and critical services.
  • Participate in on-call rotations; lead incident triage, debugging, root cause analysis, and post-mortems.
  • Build and operate platform automation and AI-powered workflows (including agent-based systems) to reduce manual operational effort.
  • Design and implement guardrails, validation pipelines, and safety mechanisms for automated and AI-generated changes to code and infrastructure.
  • Enable closed-loop automation systems (detect diagnose remediate validate) to improve system resilience.
  • Define and track SLIs and SLOs; use reliability data to guide engineering decisions.
  • Standardize build, deployment, and release workflows for safe, predictable delivery, including automation-friendly and AI-integrated pipelines.
  • Identify and remediate security vulnerabilities across systems and services, including risks introduced by automated changes.
  • Partner with development teams on service design, resilience, and operability, with an emphasis on automation-first and AI-compatible system design.

Skills

Java
Python
Go
Shell
Observability
CI/CD
Linux
Communication

Tools

CI/CD platforms
Monitoring tools
Logging/Tracing systems

Job description

Job Summary

Were looking for a Platform Reliability Engineer to own the reliability, operability, and evolution of our internal engineering platform. This is a hands‑on role at the intersection of platform engineering, reliability, and intelligent automation with a clear mandate: reduce toil, improve observability, and enable systems (AI agents) to safely operate at scale.

Youll work directly with engineering teams to harden services, respond to incidents, and build automation that makes the platform increasingly self‑managing over time. A key aspect of this role is designing and operating AI‑driven and agent‑based workflows, including the guardrails, validation systems, and observability needed to allow automated systems to safely generate and act on changes in production environments.

What Youll Do
  • Own reliability, availability, and performance of the internal platform and critical services
  • Participate in on‑call rotations; lead incident triage, debugging, root cause analysis, and post‑mortems
  • Build and operate platform automation and AI‑powered workflows (including agent‑based systems) to reduce manual operational effort
  • Design and implement guardrails, validation pipelines, and safety mechanisms for automated and AI‑generated changes to code and infrastructure
  • Enable closed‑loop automation systems (detect diagnose remediate validate) to improve system resilience
  • Define and track SLIs and SLOs; use reliability data to guide engineering decisions
  • Standardize build, deployment, and release workflows for safe, predictable delivery, including automation‑friendly and AI‑integrated pipelines
  • Identify and remediate security vulnerabilities across systems and services, including risks introduced by automated changes
  • Partner with development teams on service design, resilience, and operability, with an emphasis on automation‑first and AI‑compatible system design
Required Qualifications
  • 6+ years of experience operating production platforms or large‑scale distributed systems
  • Proven track record in incident management, on‑call operations, and production debugging
  • Strong programming skills in Java, Python, Go, Shell, or equivalent
  • Hands‑on experience with observability tooling (monitoring, alerting, logging, tracing)
  • Experience building or maintaining CI/CD pipelines and release processes
  • Familiarity with platform upgrades, dependency management, and system lifecycle operations
  • Experience building or integrating AI‑driven (agent‑based) automation frameworks, or strong interest in this space
  • Working knowledge of Linux‑based production environments
  • Strong communication and cross‑team collaboration skills
Nice to Have
  • Experience with SRE frameworks: SLOs, error budgets, reliability reviews
  • Experience with chaos engineering or resilience testing
  • Background in building self‑healing systems
  • History of driving platform standardization across large engineering organizations
What Success Looks Like at 6 Months
  • Platform reliability metrics are tracked, visible, and trending in the right direction
  • On‑call burden is measurably reduced through automation and better runbooks
  • At least one significant automation or autonomous remediation initiative shipped and adopted by engineering teams
  • Platform upgrades and rollouts are executed safely with documented processes
  • AI‑driven or automated changes are safely deployed with clear guardrails, observability, and rollback mechanism
Key Traits

Strong ownership mentality. Calm under pressure. Bias toward automation. Systems thinker who doesn’t just fix problems but builds systems that prevent, detect, and autonomously remediate issues over time.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MTS 2, Platform Reliability Engineer
MTS 2, Platform Reliability Engineer

eBay • Bengaluru

On-site
INR 4,200,000 - 6,400,000
Site Reliability Engineer
Site Reliability Engineer

Nexcess • India

On-site
INR 2,000,000 - 4,200,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

Sita • New Delhi

On-site
INR 2,600,000 - 4,200,000
Sr Software Engineer
Sr Software Engineer

Nvidia • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Senior Software Engineer - AI Platform Engineer
Senior Software Engineer - AI Platform Engineer

CloudBees • Chennai District

On-site
INR 3,500,000 - 5,500,000