Site Reliability Engineer

Nexcess

India

On-site

INR 2,000,000 - 4,200,000

Full time

12 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nexcess is seeking a Platform SRE to join the team responsible for the health, resiliency, and modernization of our hosting and cloud infrastructure. You will bridge systems engineering, SRE, and automation with AI-augmented tooling to scale reliability and drive modernization across production environments.

The role emphasizes technical debt reduction, automation of critical workflows, reliability metrics (SLIs/SLOs), and proactive incident response.

Qualifications

  • Advanced Linux systems administration and troubleshooting; kernel and firmware familiarity.
  • Strong Kubernetes, Docker, and container orchestration experience.
  • Hands-on automation and infrastructure-as-code (Terraform, Ansible, Puppet/Chef).
  • Experience with CI/CD pipelines and release automation.
  • Familiarity with AI SRE/AIOps tooling for proactive detection and remediation.
  • Experience defining and tracking SLIs/SLOs and observability tooling.

Responsibilities

  • Own lifecycle of patching across managed hosting fleets and application environments.
  • Design automation for provisioning, deployment, patching, remediation, and configuration management.
  • Develop automated release pipelines with human-in-the-loop approval gates; support CI/CD workflows.

Skills

Kubernetes
Advanced Linux
Docker
Python
Bash
Go
Terraform
Ansible
Puppet/Chef
AI SRE / AIOps tooling

Tools

Kubernetes
Docker
Terraform
Ansible
Puppet
Chef
Grafana
Datadog
Prometheus

Job description

The Platform SRE is a member of the Platform SRE team, the group responsible for the operational health, resiliency, and modernization of our hosting and cloud infrastructure. This role sits at the intersection of systems engineering, site reliability engineering, and automation : increasingly AI-augmented automation : combining hands-on technical debt remediation with large-scale automation and reliability practice.

The Platform SRE team owns three core mandates
1. Technical debt reduction :

systematically identifying and remediating aging firmware, kernels, operating systems, and software across the managed hosting fleet and managed application environments

2. Automation of critical operational workflows :

provisioning, patching, remediation, and release processes across both managed apps and managed hosting fleets, with automated release planning and execution under human supervision (not “automation for automation’s sake,” but automation with a human checkpoint before production impact)

3. Reliability and incident support :

defining, instrumenting, and tracking SLIs/SLOs for platform engineering and operations, visualized through dashboards and reporting tools, and providing fast, expert frontline response and remediation during service-impacting events (incident command and process ownership sit with the Incident Management team; this team is the technical responder, not the incident owner).

4. AI-augmented operations :

operating agentic AI SRE tooling for proactive anomaly detection, automated investigation, and human-gated remediation across daily operations, so the team scales its impact without over-scaling headcount with the fleet.

This position serves as a senior technical point of contact for platform engineering, driving initiatives around scalability, fault tolerance, automation, and operational excellence across production infrastructure.

Key Responsibilities
Technical Debt & Platform Modernization
  • Own the lifecycle of firmware, kernel, OS, and software patching across the managed hosting fleet and managed application environments
  • Build a standing inventory and risk model of technical debt (end-of-life OS versions, unpatched firmware, deprecated software) and drive prioritized remediation plans
  • Evaluate and implement infrastructure modernization initiatives, replacing manual or legacy processes with supportable, automated alternatives
Automation & Release Engineering
  • Design and build automation for provisioning, deployment, patching, remediation, and configuration management across managed apps and managed hosting fleets
  • Own the design of automated release pipelines : planning, staging, and executing releases with defined human-in-the-loop approval gates
  • Develop self-healing and auto-remediation capability for common failure modes to reduce manual operational load
  • Support and extend CI/CD workflows and infrastructure-as-code practices across the platform
Reliability Engineering, SLIs/SLOs & Observability
  • Define SLIs and SLOs for platform engineering and operations in partnership with engineering and product stakeholders
  • Instrument systems to measure SLIs accurately and build SLO tracking into standard reporting
  • Build and maintain dashboards (e.g., Grafana, Datadog, or equivalent visualization tooling) to make SLI/SLO performance, error budgets, and platform health visible to engineering and leadership
  • Continuously improve platform observability : monitoring, alerting, logging, and tracing : across distributed and containerized environments
  • Serve as the frontline technical responder: acknowledge pages quickly, diagnose, and remediate platform-level issues
  • Partner with the Incident Management team, who own incident command, severity classification, and customer communication : this role provides the technical hands and expertise, not incident ownership
  • Contribute technical findings to blameless root cause analysis (RCA) and own follow-through on corrective actions for platform systems
  • Maintain runbooks and on-call readiness for platform and infrastructure systems
  • Track incident trends on platform systems and feed them back into the technical debt and automation roadmap
AI-Augmented Operation
  • Operate and tune agentic AI SRE tooling (e.g., Harness AI SRE, HolmesGPT, K8sGPT, or equivalent platforms) for proactive anomaly detection, automated investigation, and root-cause drafting across the fleet
  • Apply AIOps-style alert correlation and noise reduction to cut duplicate/low-value pages and protect on-call sustainability
  • Use AI-assisted runbook and chaos-engineering tooling to convert manual procedures into self-executing, testable workflows
  • Maintain the human-in-the-loop gate on all AI-suggested or AI-generated remediation before it reaches production : AI drafts and proposes, this role verifies and approves
  • Evaluate new AI SRE tooling for fit, accuracy, and safety before adoption; retire tools that don’t earn their keep
Collaboration & Technical Leadership
  • Partner with software engineering teams on platform architecture, operational readiness reviews, and scalability initiatives
  • Support platform security, compliance, and operational governance requirements
  • Mentor engineers and contribute to technical leadership and knowledge-sharing across the team
  • Maintain clear operational documentation and contribute to team standards and process improvement
  • Other duties as assigned
Requirements
  • 3–5+ years of experience in platform engineering, systems engineering, SRE, or infrastructure operations (level based on experience and scope)
  • Advanced Linux systems administration and troubleshooting expertise, including kernel and firmware-level familiarity
  • Strong experience with Kubernetes, Docker, and container orchestration/distributed systems
  • Hands-on automation and infrastructure-as-code experience (e.g., Terraform, Ansible, Puppet/Chef, or equivalent)
  • Familiarity with agentic AI SRE/AIOps tooling for proactive detection, investigation, and human-gated remediation (e.g., Harness AI SRE, HolmesGPT, K8sGPT, or equivalent platforms)
  • Experience building or maintaining CI/CD and automated release/deployment pipelines
  • Experience defining and tracking SLIs/SLOs and working with observability/visualization tools (e.g., Grafana, Datadog, Prometheus, or equivalent)
  • Experience supporting enterprise-scale, high-concurrency, or customer-impacting production environments
  • Demonstrated experience as a technical responder in production incidents, including root cause analysis and corrective action follow-through
  • Strong scripting ability (e.g., Python, Bash, Go) for automation and tooling
  • Strong troubleshooting skills across compute, network, storage, and application layers
  • Experience supporting cloud-hosted, managed hosting, or hybrid infrastructure environments
  • Ability to lead technical initiatives, mentor others, and communicate clearly across teams
Preferred Qualifications
  • Experience owning fleet-wide firmware/OS patch management programs at scale
  • Experience designing human-in-the-loop release automation or progressive delivery systems (canary, blue/green)
  • Experience operating or tuning AI-driven observability/AIOps platforms (alert correlation, autonomous RCA drafting, AI-assisted chaos engineering test generation)
  • Experience evaluating or piloting emerging AI SRE agent platforms and setting guardrails for safe adoption
  • Familiarity with error budgets and SLO-driven prioritization frameworks
  • Experience with configuration/patch management at scale across heterogeneous hardware fleets
  • The physical demands described here are representative of those that must be met by an individual to successfully perform the essential duties of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential duties.
  • While performing the essential duties of this job, the individual is regularly required to speak and hear
  • Works at a desk and computer screen for extended periods of time
  • Works in a traditional climate-controlled office environment or from home
  • Works in a highly stressful environment dealing with a wide variety of challenges, deadlines, and diverse employee population
  • Require participation in an on-call rotation, including responding to incidents outside standard business hours
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Technology Manager
Technology Manager

Wolters Kluwer • Pune District

Hybrid
INR 2,800,000 - 5,000,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Platform Engineer
Platform Engineer

United States Digital Space LLC • Maharashtra

On-site
INR 1,500,000 - 2,800,000
Senior Platform SRE
Senior Platform SRE

IG Infotech • Bengaluru

On-site
INR 1,200,000 - 1,600,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Platform SRE
Platform SRE

YASH Technologies • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Engineering Manager
Engineering Manager

WaferWire Cloud Technologies • Hyderabad

On-site
INR 4,000,000 - 7,000,000