Senior Site Reliability Engineer

SourcingXPress

Hyderabad

On-site

INR 3,000,000 - 5,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SourcingXPress in Hyderabad, India is seeking a skilled infrastructure architect to lead the technical direction for AI-driven healthcare solutions. This role includes designing scalable cloud architectures, mentoring engineers, and guiding incident management practices.

The ideal candidate will have extensive experience in multi-cloud environments, particularly with AWS and GCP, as well as advanced Kubernetes expertise. Join us to build impactful infrastructure that enhances patient outcomes and streamlines operations.

Qualifications

  • 5–8 years in Infrastructure, Platform, DevOps, or SRE roles.
  • Deep expertise managing large-scale production systems in AWS and GCP.
  • Advanced Kubernetes skills spanning cluster design, multi-tenancy, and resource quotas.

Responsibilities

  • Design scalable, secure AWS/GCP architectures for AI and healthcare workloads.
  • Lead incident management, RCA, and blameless post-mortems for critical systems.
  • Evaluate and introduce AI tooling, platforms, and automation best practices.

Skills

Infrastructure management
Kubernetes
AWS
GCP
Terraform
AI/ML Ops
Incident management

Job description

Own infrastructure architecture, scalability, and reliability. Set technical direction, define SRE practice, and mentor engineers. Build and operate AI-scale infrastructure powering real-time healthcare workflows.

Company
  • Recruiting Bond
  • Business Type: Startup
  • Company Type: Product
  • Business Model: B2B
  • Funding Stage: Seed
  • Industry: Healthcare
  • Salary Range: ₹ 30-50 Lacs PA
  • Website: Visit Website
Key Responsibilities
Architecture & Platform
  • Design scalable, secure, cost‑efficient AWS/GCP architectures for AI and healthcare workloads
  • Lead IaC standards and reusable modules org‑wide; own GitOps practices
  • Architect optimized Kubernetes (EKS/GKE) for HA, scale, and GPU workload orchestration
Reliability Engineering
  • Define and drive reliability targets (SLOs/SLIs); establish SRE practice
  • Lead incident management, RCA, and blameless post‑mortems for critical systems
  • Partner on infrastructure hardening and compliance (SOC 2, HIPAA)
AI‑First Engineering
  • Lead AI/ML inference and GPU infrastructure scaling in production
  • Evaluate and introduce AI tooling, platforms, and automation best practices
  • Contribute to LLM deployment pipelines and AI systems reliability
Leadership & SDLC
  • Mentor L1/L2 engineers; influence cross‑team infrastructure decisions
  • Drive DevSecOps maturity across the SDLC (policy‑as‑code, SBOM, supply‑chain security)
  • Evaluate and introduce new tooling and platforms
Experience
  • 5–8 years in Infrastructure, Platform, DevOps, or SRE roles.
  • Multi‑Cloud: Deep expertise managing large‑scale production systems in AWS and GCP.
  • Orchestration: Advanced Kubernetes skills spanning cluster design, multi‑tenancy, and resource quotas.
  • AI/ML Ops: Proven track record scaling GPU infrastructure and production AI/ML inference.
  • Automation: Strong Terraform (or equivalent) proficiency integrated with GitOps workflows.
  • Development: Solid scripting and automation capabilities using Python, Go, or Bash.
  • Metrics: Demonstrated history of driving system scale, high reliability, and cost efficiency.
  • Operations: Experience leading incident response management and structuring on‑call practices.
Preferred Qualifications
  • Architecture and optimization experience with LLM Gateways.
  • Knowledge of Zero‑Trust networking and compliance frameworks like SOC 2 and HIPAA.
  • Focus on DevSecOps maturity, including SBOM, supply‑chain security, and policy‑as‑code.
  • Background in platform team leadership or active open‑source community contributions.
Why this opportunity?
  • Build infrastructure and products powering real‑world healthcare AI
  • Work with AI‑native engineering teams using cutting‑edge tools (Claude Code, Cursor, Copilot)
  • Mission‑driven impact — improve patient outcomes and save thousands of staff hours
  • High ownership, rapid learning, and significant career upside
  • Global collaboration across India and the US
  • Seed‑stage growth — scale from Seed to Series A alongside the team
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer
Senior Infrastructure Engineer

SourcingXPress • Hyderabad

On-site
INR 2,000,000 - 3,500,000
High ownership
Rapid learning opportunities
Career advancement potential
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • India

On-site
INR 4,200,000 - 6,200,000
Hybrid work in Hyderabad
Health & life insurance
Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Bengaluru

Hybrid
INR 6,000,000 - 12,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Gurugram District

Hybrid
INR 3,500,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Maharashtra

On-site
INR 700,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Relatient • Pune District

On-site
INR 1,500,000 - 2,600,000
Life insurance
Accident coverage
Education reimbursement
+2
Infrastructure Associate Advisor
Infrastructure Associate Advisor

Evernorth Health Services • Hyderabad

Hybrid
INR 1,000,000 - 2,000,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000