Site Reliability Engineer - India

JumpCloud Inc.

United States

Remote

USD 120,000 - 180,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

JumpCloud is seeking a Software Engineer 3 (SRE) to join our Infrastructure & Reliability Engineering team. You will design automation, build observability, and reduce toil across production systems.

You will operate multi-cloud infrastructure (AWS/GCP), implement SLOs and DR processes, and drive reliability improvements through code and tooling.

Qualifications

  • 3–5 years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems.
  • Python/Go proficiency for SRE tools and cloud integrations.
  • Kubernetes production experience with cluster operations, container orchestration, and GitOps pipelines.

Responsibilities

  • Design, deploy, and maintain reliability, availability, and performance of JumpCloud systems and APIs across AWS and GCP.
  • Operationalize SLIs, SLOs, and error budgets with core app teams.
  • Build end-to-end observability across microservices and cloud infra using Datadog.
  • Implement monitoring across Golden Signals to optimize detection and reduce alert fatigue.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews.
  • Manage production Kubernetes (EKS) clusters with GitOps workflows (Argo CD, Kargo).
  • Provision and secure multi-cloud infra using Terraform.
  • Develop DR dashboards, runbooks, and failover automation to meet RTO/RPO targets.
  • Eliminate toil with production-grade Python or Go automation scripts.
  • Leverage AI-assisted tools to accelerate scripting and incident triage.

Skills

Python/Go proficiency
Kubernetes operations
Observability
Incident management
Automation scripting
GitOps pipelines

Tools

Datadog
Argo CD
Terraform
Istio
HAProxy
NGINX

Job description

About the role: We are seeking a Software Engineer 3 (SRE) to join our Infrastructure & Reliability Engineering team. This role sits at the center of platform resilience—ensuring high availability, performance, recoverability, and operational maturity across JumpCloud’s production systems. This is not a traditional operations role. Our SREs are engineers first: designing automation, building observability frameworks, defining reliability standards, and reducing operational toil through code. You will build and scale cloud-native infrastructure, participate in incident management, and implement reliability best practices across our directory platform and microservices.

What you’ll be doing:
  • Design, deploy, and maintain the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP.
  • Operationalize SLIs, SLOs, and error budgets in direct partnership with core application teams.
  • Build and refine end-to-end observability across microservices and cloud infrastructure using tools like Datadog.
  • Implement actionable monitoring across Golden Signals (Latency, Traffic, Errors, Saturation) to optimize detection (MTTD) and minimize alert fatigue.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews to drive continuous systemic improvements.
  • Manage and operationalize production Kubernetes (EKS) clusters utilizing GitOps delivery workflows (Argo CD, Kargo).
  • Provision and secure multi-cloud infrastructure using modular Terraform (Infrastructure-as-Code).
  • Develop and maintain Disaster Recovery (DR) dashboards, runbooks, multi-region failover automation, and validation tests to ensure alignment with defined RTO and RPO targets.
  • Eliminate operational toil by writing production-grade Python or Go scripts and automation tools.
  • Leverage AI-assisted development tools (Cursor, Claude Code, GitHub Copilot) to accelerate scripting, runbook generation, and incident triage.
We’re looking for:
  • 3–5 years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems.
  • Python/Go Proficiency: Hands‑on capabilities writing code for SRE tools, custom automation, and cloud integrations.
  • Kubernetes Ecosystem: Production experience with Kubernetes cluster operations, container orchestration, and GitOps pipelines (Argo CD).
  • Infrastructure as Code: Solid experience writing, maintaining, and modularizing Terraform configurations.
  • Cloud Architecture: Direct experience operating cloud workloads on AWS (EKS, IAM, VPC networking, Route53, ALB/NLB) or GCP.
  • FinOps & Cost Visibility: Practical experience setting up cost‑allocation tagging, resource right‑sizing, and building FinOps dashboards to visualize cloud spend.
  • Disaster Recovery & Monitoring: Experience building DR dashboards, running failover drills, and configuring monitoring tools to track system health and recovery metrics
  • Observability & Incident Management: Practical experience with Datadog (or similar), PagerDuty, alerting hygiene, and working within SLI/SLO frameworks.
  • Solid operational experience configuring and troubleshooting production service meshes (Istio or similar) and managing high‑availability proxy solutions (HAProxy, NGINX, or similar).
  • Problem Solving & Mindset: Strong troubleshooting skills, effective collaboration, and a track record of driving operational efficiency through code.
  • A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.
Preferred Qualifications:
  • Experience with CI/CD tools such as GitHub Actions or GitLab Pipelines.
  • Basic understanding of chaos engineering principles or testing resilience in staging/production.
  • Familiarity with secrets management tools (HashiCorp Vault, AWS Secrets Manager, External Secrets Operator).
  • Basic knowledge of DevSecOps tools and scanning/fixing infrastructure-as-code vulnerabilities.

#LI-MS1

Where you’ll be working/Location:

JumpCloud is committed to being Remote First, meaning that you are able to work remotely within the country noted in the Job Description.

You must be located in and authorized to work in the country noted in the job description to be considered for this role.

Please note: There is an expectation that our engineers participate in on-call shifts. You will be expected commit to being ready and able to respond during your assigned shift, so that alerts don't go unaddressed.

Language:

JumpCloud has teams in 15+ countries around the world and conducts our internal business in English. The interview and any additional screening process will take place primarily in English. To be considered for a role at JumpCloud, you will be required to speak and write in English fluently. Any additional language requirements will be included in the details of the job description.

Why JumpCloud?

If you thrive working in a fast, SaaS‑based environment and you are passionate about solving challenging technical problems, we look forward to hearing from you! JumpCloud is an incredible place to share and grow your expertise! You’ll work with amazing talent across each department who are passionate about our mission. We’re out of the box thinkers, so your unique ideas and approaches for conceiving a product and/or feature will be welcome. You’ll have a voice in the organization as you work with a seasoned executive team, a supportive board and in a proven market that our customers are excited about.

One of JumpCloud's three core values is to “Build Connections.” To us that means creating " human connection with each other regardless of our backgrounds, orientations, geographies, religions, languages, gender, race, etc. We care deeply about the people that we work with and want to see everyone succeed." - Rajat Bhargava, CEO

JumpCloud is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

#LI-Remote #BI-Remote

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - India
Senior Site Reliability Engineer - India

JumpCloud Inc. • United States

Remote
USD 170,000 - 260,000
Remote-first culture
Senior Vice President of Global Customer Success & Support - United States
Senior Vice President of Global Customer Success & Support - United States

JumpCloud Inc. • Atlanta (GA)

Remote
USD 250,000 - 500,000
Site Reliability Engineer (Fully Remote)
Site Reliability Engineer (Fully Remote)

JumpCloud Inc. • United States

Remote
USD 120,000 - 180,000
Senior Quality Engineer - India
Senior Quality Engineer - India

JumpCloud Inc. • United States

Remote
USD 100,000 - 130,000
Senior Vice President of Global Customer Success & Support - United States
Senior Vice President of Global Customer Success & Support - United States

JumpCloud Inc. • San Francisco (CA)

Remote
USD 250,000 - 400,000
Sales Engineer - United States JumpCloud · Denver, CO Full-time · Remote $94,000–172,000 1 hour ago
Sales Engineer - United States JumpCloud · Denver, CO Full-time · Remote $94,000–172,000 1 hour ago

Emploive • Denver (CO)

Hybrid
USD 94,000 - 172,000
Health plans (medical, dental, vision)
HSA plan with employer contribution
FSA
+4
Sales Engineer - United States
Sales Engineer - United States

Lever, Inc. • Denver (CO)

Remote
USD 90,000 - 130,000
Sales Engineer - United States
Sales Engineer - United States

Wwshemi • Northern (KY)

Remote
USD 90,000 - 140,000
Account Executive - United States
Account Executive - United States

Lever, Inc. • Denver (CO)

Remote
USD 175,000 - 225,000
Medical plans
Dental plans
Vision plans
+2
Account Executive - United States
Account Executive - United States

JumpCloud Inc. • Denver (CO)

On-site
USD 175,000 - 225,000