Senior Site Reliability Engineer (Night Shift)

Resilinc

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Fully remote
In-person meetups
Comprehensive benefits

Job summary

Resilinc is seeking an experienced Site Reliability Engineer (SRE) to join our remote-first team. You will own production incidents, improve system reliability, and drive automation across a modern cloud-native stack (Azure, Kubernetes, Kafka, Redis, PostgreSQL).

The role establishes dedicated India-based night-time coverage aligned with US business hours, addressing a critical support gap and offering high-impact work on globally used platforms.

Qualifications

  • 6-12 years of experience in SRE/DevOps or related roles.
  • Strong hands-on experience with Azure Cloud services.
  • Solid experience in Linux system administration.
  • Expertise in Docker and Kubernetes (deployment, scaling, troubleshooting).
  • Experience with Kafka, Redis, and PostgreSQL.
  • Working knowledge of Hadoop ecosystem (HDFS, Hadoop).
  • Experience with Cloudflare (CDN, security, DNS management).
  • Proficiency in CI/CD tools (GitHub, GitHub Actions).
  • Experience in Helm Charts and Kubernetes deployments.
  • Strong understanding of monitoring and logging tools (Grafana, etc.)

Responsibilities

  • Design, implement, and manage scalable and highly available systems on Azure Cloud.
  • Monitor system performance, troubleshoot issues, and ensure uptime and reliability.
  • Manage and optimize Kubernetes clusters and containerized workloads (Docker).
  • Build and maintain robust CI/CD pipelines using GitHub Actions and related tools.
  • Implement infrastructure as code and deployment automation using Helm Charts.
  • Work with distributed systems such as Kafka, Redis, PostgreSQL, Hadoop/HDFS.
  • Configure and manage Cloudflare for performance, security, and traffic routing.
  • Set up monitoring, alerting, and observability using tools like Grafana.
  • Collaborate with development teams to improve system reliability and deployment practices.
  • Perform root cause analysis (RCA) and implement preventive measures.
  • Ensure security best practices and compliance across infrastructure.

Skills

SRE/DevOps experience
Cloud-native systems
Linux administration
Docker & Kubernetes proficiency
CI/CD automation
Monitoring & logging

Tools

Azure Cloud
Kubernetes
Docker
Kafka
Redis
PostgreSQL
Hadoop/HDFS
Cloudflare
GitHub Actions
Helm Charts
Grafana

Job description

Join the Future of Supply Chain Intelligence - Powered by Agentic AI

At Resilinc, we're pioneering intelligent, autonomous systems that redefine supply chain risk management. Our agentic AI helps global enterprises predict disruptions, assess impact, and act in real time - before operations are affected. Named a 2025 Gartner® Magic Quadrant™ Leader, we're trusted by top companies in life sciences & pharma, aerospace & defense, high tech, and automotive to protect what matters most. Be part of a team that’s redefining resilience on a global scale.

But the real power behind Resilinc? Our people. We're a fully remote, mission-led team making sure life-saving products and critical goods get where they're needed, fast.We offer the chance to do meaningful work in a collaborative, empowering culture-where you can be an agent of change.Join us to tackle critical global challenges through high-impact work that matters.

Check out this blog to learn more about how we are impacting the world's most critical supply chains. Global Supply Chain Risks 2026: Act Faster | TEC

Resilinc | Innovation with Purpose. Intelligence with Impact.

We are looking for an experienced Site Reliability Engineer (SRE) to join our team. The ideal candidate will be responsible for ensuring high availability, scalability, performance, and reliability of our platform while driving automation and operational excellence across cloud-native environments.

This role offers an opportunity to work at the core of production reliability for a globally used platform. The position is being created to establish dedicated India-based night-time SRE coverage aligned with US business hours, addressing a critical support gap.

The SRE will play a key role in owning production incidents, improving system reliability, reducing MTTR, and driving automation across a modern cloud-native stack (Azure, Kubernetes, Kafka, etc.). This is a high-impact role with direct visibility into business-critical operations and opportunities to work on large-scale distributed.

What You Will Do
  • Design, implement, and manage scalable and highly available systems onAzure Cloud
  • Monitor system performance, troubleshoot issues, and ensure uptime and reliability
  • Manage and optimizeKubernetes clustersand containerized workloads (Docker)
  • Build and maintain robustCI/CD pipelinesusing GitHub Actions and related tools
  • Implement infrastructure as code and deployment automation usingHelm Charts
  • Work with distributed systems such asKafka, Redis, PostgreSQL, Hadoop/HDFS
  • Configure and manageCloudflarefor performance, security, and traffic routing
  • Set up monitoring, alerting, and observability using tools likeGrafana
  • Collaborate with development teams to improve system reliability and deployment practices
  • Perform root cause analysis (RCA) and implement preventive measures
  • Ensure security best practices and compliance across infrastructure
What You Will Bring
  • 6-12 years of experience in SRE/DevOps or related roles
  • Strong hands-on experience withAzure Cloud services
  • Solid experience inLinux system administration
  • Expertise inDocker and Kubernetes(deployment, scaling, troubleshooting)
  • Experience withKafka, Redis, and PostgreSQL
  • Working knowledge ofHadoop ecosystem (HDFS, Hadoop)
  • Experience withCloudflare(CDN, security, DNS management)
  • Proficiency inCI/CD tools(GitHub, GitHub Actions)
  • Experience inHelm Chartsand Kubernetes deployments
  • Strong understanding ofmonitoring and logging tools(Grafana, etc.)
What Will Make You Stand Out
  • Experience with large-scale distributed systems
  • Knowledge of infrastructure automation tools (Terraform, Ansibleetc.)
  • Exposure to security and compliance best practices
  • Strong problem-solving and troubleshooting skills
  • Databricks , Clickhouse and ML Ops exposure
Why You Will Love It Here
  • Opportunity to work on cutting-edge cloud and distributed systems
  • Exposure to large-scale, high-impact platforms
  • Collaborative and innovation-driven environment
What’s in it for you?

At Resilinc, we're fully remote, with plenty of opportunities to connect in person. We provide a culture where ownership, purpose, technical growth and a voice in shaping impactful technology are at our core. Oh, and the perks? Full-stack benefits for health, wealth and wellbeing to keep you thriving. Check in with your talent acquisition contact for a location-specific FAQ.

Curious to know more about us? Dive in at www.resilinc.ai

More great news! Resilinc is backed by Vista Equity Partners

Resilinc is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic protected by law.

If you are a person with a disability needing assistance with the application process please contact HR@resilinc.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Night-Shift SRE: Global Cloud Reliability (Remote)
Senior Night-Shift SRE: Global Cloud Reliability (Remote)

Resilinc • United States

Remote
USD 140,000 - 210,000
Fully remote
In-person meetups
Comprehensive benefits
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • United States

On-site
USD 145,000 - 200,000
Remote-first
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

On-site
USD 140,000 - 170,000
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 130,000 - 180,000
Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
+3
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
3510- Site Reliability Engineer II
3510- Site Reliability Engineer II

Innovaccer • Dallas (TX)

On-site
USD 110,000 - 140,000
Generous Paid Time Off: 22 days per year plus company holidays
Best-in-Class Parental Leave
Comprehensive insurance coverage