Senior Site Reliability Engineer

BlinkRx

Bengaluru

On-site

INR 3,500,000 - 7,000,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Blink Health is seeking an experienced Site Reliability Engineer to join our cloud-based platform and help establish and evolve SRE practices across the organization.

You will define observability, design automation, mentor engineers, and lead large, ambiguous initiatives with security and product teams. The role requires deep Linux, networking, and AWS/GCP experience and a collaborative mindset.

Qualifications

  • Bachelor's or Master's degree in Computer Science or equivalent practical experience.
  • 10+ years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles, with demonstrated impact at scale.
  • Expert-level, methodical troubleshooting across the entire stack, from application to kernel to network.

Responsibilities

  • Establish and evolve SRE best practices across the organization, including reliability principles, error budgets, incident response, postmortems, and operational readiness standards.
  • Define and drive observability strategy for system health, performance, and reliability, including SLIs/SLOs, alerting quality, dashboards, and service health indicators.
  • Design and implement software-driven solutions within the infrastructure domain, automating manual processes and eliminating operational complexity and toil.
  • Act as a technical leader and force multiplier, helping set priorities and influencing decision-making across core cloud infrastructure, reliability tooling, and platform architecture.
  • Take ownership of large, ambiguous initiatives, driving them from concept to delivery while aligning stakeholders across engineering, security, and product.
  • Combine deep knowledge of software development, infrastructure, and security to improve platform resilience, scalability, performance, and compliance.
  • Proactively identify systemic risks and reliability gaps, recommending and leading platform upgrades and architectural improvements before they become incidents.
  • Partner with engineering teams to improve developer workflows, tooling, and operational maturity, increasing productivity while reducing cognitive load.
  • Provide technical mentorship, architecture guidance, and high-quality design and code reviews for engineers across infrastructure and product teams.
  • Lead by example in documentation and knowledge sharing, ensuring systems and processes are well-understood and not dependent on individual ownership.
  • Participate in and help mature incident response, escalation practices, and post-incident learning across the organization.

Skills

Troubleshooting
Automation
Mentorship
Cross-functional collaboration

Education

Bachelor's or Master's in Computer Science or equivalent

Tools

Kubernetes (EKS/Helm)
Terraform
Pulumi
CloudFormation
Ansible
AWS
GCP

Job description

Blink Health is the fastest growing healthcare technology company that builds products to make prescriptions accessible and affordable to everybody. Our two primary products - BlinkRx and Quick Save - remove traditional roadblocks within the current prescription supply chain, resulting in better access to critical medications and improved health outcomes for patients.

BlinkRx is the world’s first pharma-to-patient cloud that offers a digital concierge service for patients who are prescribed branded medications. Patients benefit from transparent low prices, free home delivery, and world-class support on this first-of-its-kind centralized platform. With BlinkRx, never again will a patient show up at the pharmacy only to discover that they can’t afford their medication, their doctor needs to fill out a form for them, or the pharmacy doesn’t have the medication in stock.

We are a highly collaborative team of builders and operators who invent new ways of working in an industry that historically has resisted innovation. Join us!

Responsibilities
  • Establish and evolve SRE best practices across the organization, including reliability principles, error budgets, incident response, postmortems, and operational readiness standards.
  • Define and drive observability strategy for system health, performance, and reliability, including SLIs/SLOs, alerting quality, dashboards, and service health indicators.
  • Design and implement software-driven solutions within the infrastructure domain, automating manual processes and eliminating operational complexity and toil.
  • Act as a technical leader and force multiplier, helping set priorities and influencing decision-making across core cloud infrastructure, reliability tooling, and platform architecture.
  • Take ownership of large, ambiguous initiatives, driving them from concept to delivery while aligning stakeholders across engineering, security, and product.
  • Combine deep knowledge of software development, infrastructure, and security to improve platform resilience, scalability, performance, and compliance.
  • Proactively identify systemic risks and reliability gaps, recommending and leading platform upgrades and architectural improvements before they become incidents.
  • Partner with engineering teams to improve developer workflows, tooling, and operational maturity, increasing productivity while reducing cognitive load.
  • Provide technical mentorship, architecture guidance, and high-quality design and code reviews for engineers across infrastructure and product teams.
  • Lead by example in documentation and knowledge sharing, ensuring systems and processes are well-understood and not dependent on individual ownership.
  • Participate in and help mature incident response, escalation practices, and post-incident learning across the organization.
Desired Experience
  • Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.
  • 10+ years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles, with demonstrated impact at scale.
Reliability & Troubleshooting
  • Expert-level, methodical troubleshooting across the entire stack, from application to kernel to network.
  • Strong command-line proficiency and deep expertise in Linux systems and operating system fundamentals.
  • Advanced understanding of networking concepts including load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication.
Software & Automation
  • Experience working across multiple languages (e.g., Python, Go, Bash, and familiarity troubleshooting application stacks such as React or similar).
  • Strong track record of automating repetitive and complex operational work to reduce toil and increase reliability.
  • Ability to design and build internal tools (Python or Go) that standardize and scale engineering practices.
  • Comfortable operating in an agile environment, with disciplined testing and quality practices.
  • Deep experience with cloud platforms (AWS preferred, GCP/Azure acceptable), particularly managed services and production-grade architectures.
  • Strong expertise in Kubernetes and container orchestration (EKS, Helm), including lifecycle management and operational best practices.
  • Proven experience designing and implementing observability systems, including metrics, logging, tracing, dashboards, and alerting.
  • Deep understanding of container technologies, security scanning, secrets management, dynamic configuration, and microservices architectures.
  • Familiarity with service meshes and advanced traffic management concepts.
Infrastructure as Code
  • Experience designing and maintaining company-wide IaC codebases using tools such as Terraform, Pulumi, CloudFormation, or Ansible.
  • Ability to think holistically about infrastructure design, cost, reliability, security, and long-term maintainability.
Why Join Us:

It is rare to have a company that both deeply impacts its customers and is able to provide its services across a massive population. At Blink, we have a huge impact on people when they are most vulnerable: at the intersection of their healthcare and finances. We are also the fastest growing healthcare company in the country and are driving that impact across millions of new patients every year. Our business model not only helps people, but drives economics that allow us to build a generational company. We are a relentlessly learning, constantly curious, and aggressively collaborative cross-functional team dedicated to inventing new ways to improve the lives of our customers.

We are an equal opportunity employer and value diversity of all kinds. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Applicants who provide their phone number and consent to receive text messages may receive SMS or MMS updates from Blink Health regarding their application.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff / Lead Site Reliability Engineer
Staff / Lead Site Reliability Engineer

blinkhealth • India

On-site
INR 4,200,000 - 7,000,000
Staff / Lead Site Reliability Engineer
Staff / Lead Site Reliability Engineer

Blink Health • India

Hybrid
INR 4,500,000 - 7,500,000
Senior Cloud Resilience Architect
Senior Cloud Resilience Architect

BlinkRx • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Senior Cloud Resilience Architect
Senior Cloud Resilience Architect

Blink Health • India

On-site
INR 2,500,000 - 3,800,000
Director, Engineering
Director, Engineering

Blink Health • India

On-site
INR 5,000,000 - 7,500,000
Senior Software Engineer II
Senior Software Engineer II

BlinkRx • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Security Engineer II
Senior Security Engineer II

BlinkRx • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Senior IT App Automation Engineer New India
Senior IT App Automation Engineer New India

Blink Health LLc • India

On-site
INR 2,500,000 - 4,500,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Optum India • Hyderabad

On-site
INR 1,200,000 - 2,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UnitedHealth Group • Dadri

On-site
INR 2,500,000 - 4,000,000