Senior Site Reliability Engineer - India

JumpCloud Inc.

United States

Remoto

USD 170.000 - 260.000

Tempo pieno

10 giorni fa
Generatore di candidature

Una candidatura completa in un minuto — curriculum e lettera di presentazione personalizzati, pronti da inviare.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Remote-first culture

Descrizione del lavoro

JumpCloud Inc. is seeking an experienced SRE/Platform Engineer to scale and secure our global, multi-region infrastructure. You will own reliability, observability, and automation across Kubernetes and cloud resources, driving cost efficiency and resilience.

You will mentor engineers, design disaster recovery, and advance GitOps practices with Argo CD in a remote-first, high-availability environment.

Competenze

  • 5–8 years of experience in SRE/DevOps/Platform Engineering for 24/7 distributed systems.
  • Bachelor's degree in Computer Science, Software Engineering, or equivalent.
  • Strong Python/Go capabilities for internal tools and integrations.
  • Deep Kubernetes expertise with production EKS/GKE lifecycles and GitOps tooling.
  • Advanced IaC with Terraform across multi-account AWS/GCP environments.
  • Experience driving FinOps, cost-optimization dashboards, and resource right-sizing.
  • Disaster Recovery design, cross-region failover, and DR dashboards.
  • Observability: define SLI/SLOs, manage PagerDuty, improve telemetry.
  • Experience with service meshes (Istio/Linkerd) and ingress systems.
  • Technical mentorship and architectural documentation skills.
  • Strong problem-solving and collaboration.

Mansioni

  • Architect, scale, and improve reliability of multi-region microservices and APIs.
  • Lead DR planning, failover automation, and business continuity strategies.
  • Define and enforce SLI/SLOs and error budgets across teams.
  • Drive end-to-end observability and reduce MTTR/MTTD.
  • Lead on-call escalation and incident management at scale.
  • Scale production Kubernetes clusters and GitOps workflows.
  • Create modular IaC using Terraform across accounts and regions.
  • Build FinOps dashboards for multi-cloud spend and unit economics.
  • Develop automation tools in Python/Go to reduce toil.
  • Mentor engineers and document architecture decisions.

Conoscenze

Python
Go
Kubernetes
Terraform
GitOps
Argo CD
Datadog
AWS/GCP
Istio/Linkerd
EKS/GKE

Formazione

Bachelor's degree in CS/CE/Equivalent

Strumenti

Terraform
Argo CD
Kubernetes
Datadog
NGINX/HAProxy

Descrizione del lavoro

All roles at JumpCloud® are Remote unless otherwise specified in the Job Description.

About JumpCloud®

JumpCloud® is the AI-powered unified IT management platform designed to secure the modern workforce. By consolidating identity, device, and access management, JumpCloud provides intelligent, secure IT that scales from human users to autonomous AI agents. We help organizations around the globe eliminate complexity and turn AI risk into an optimized advantage, ensuring the right people and agents have secure access to the right resources at all times.

JumpCloud is Intelligent, Secure IT.

What You’ll Be Doing:
  • - Architect, scale, and continuously improve the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication infrastructure (AWS/GCP).
  • - Architect, build, and maintain Disaster Recovery (DR) process, multi-region failover automation, and business continuity strategies to ensure rapid recovery against strict RTO and RPO objectives.
  • - Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks across multi-disciplinary engineering teams.
  • - Drive end-to-end observability strategy using Datadog, implementing actionable Golden Signals monitoring to drastically reduce MTTD/MTTR and eliminate alert fatigue.
  • - Lead on-call escalation, major incident management, and drive strict adherence to 99.99% availability SLAs.
  • - Facilitate blameless post-incident reviews, executing systemic root-cause remediations to prevent recurring failure modes.
  • - Architect, manage, and scale production Kubernetes (EKS) clusters, implementing advanced GitOps workflows (Argo CD, Kargo) and deployment patterns.
  • - Design and maintain modular, enterprise-grade Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments.
  • - Design, build, and maintain interactive FinOps and cost-optimization dashboards to provide engineering and leadership teams with actionable insights into multi-cloud spend, unit economics, and resource utilization.
  • - Eliminate complex operational toil by writing production-grade Python or Go tooling, platform automation, and custom integrations.
  • - Champion AI-assisted software development workflows (Cursor, Claude Code, GitHub Copilot) to accelerate automation, runbook creation, and incident triage across the team.
  • - Author operational runbooks, architecture decision records, and mentor mid-level/junior engineers to raise the overall technical bar.
We’re Looking For:
  • - 5–8 years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems.
  • - Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline.
  • - Strong Python/Go Capabilities: Advanced software engineering skill set for writing internal SRE platforms, tools, and API integrations.
  • - Deep Kubernetes Expertise: Hands-on experience with production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling (Argo CD).
  • - Advanced IaC & AWS/GCP: Deep Terraform proficiency (module architecture, state management refactoring) across complex multi-account AWS environments (IAM, VPCs, Transit Gateway, ALB/NLB, Route53).
  • - FinOps & Cost Optimization Leadership: Demonstrated experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and building FinOps dashboards to embed financial accountability into engineering workflows.
  • - Disaster Recovery & High Availability: Proven background in designing and testing multi-region Disaster Recovery architectures, automating failover systems, and monitoring recovery health via DR dashboards.
  • - Observability & Reliability Architecture: Track record of defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms.
  • - Experience designing and operating enterprise service meshes (Istio, Linkerd, or similar) and production ingress/proxy systems (HAProxy, NGINX, or similar).
  • - Technical Mentorship: Demonstrated ability to lead technical discussions, write architectural design docs/RFCs, and mentor engineering peers.
  • - Strong problem-solving, communication, and collaboration skills with a passion for solving complex distributed systems challenges at scale.
  • - A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.
Preferred Qualifications:
  • - Basic understanding of chaos engineering principles or testing resilience in staging/production.
  • - Experience with secrets management architectures (Vault, AWS Secrets Manager, External Secrets Operator, Cert-Manager).
  • - Background in DevSecOps practices, service meshes (Istio), and automated vulnerability remediation within cloud infrastructure code.
  • - Background supporting identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions.
Where you’ll be working/Location:

JumpCloud is committed to being Remote First, meaning that you are able to work remotely within the country noted in the Job Description.

You must be located in and authorized to work in the country noted in the job description to be considered for this role.

Please note: There is an expectation that our engineers participate in on-call shifts. You will be expected commit to being ready and able to respond during your assigned shift, so that alerts don't go unaddressed.

Language:

JumpCloud has teams in 15+ countries around the world and conducts our internal business in English. The interview and any additional screening process will take place primarily in English. To be considered for a role at JumpCloud, you will be required to speak and write in English fluently. Any additional language requirements will be included in the details of the job description.

Why JumpCloud?

If you thrive working in a fast, SaaS-based environment and you are passionate about solving challenging technical problems, we look forward to hearing from you! JumpCloud is an incredible place to share and grow your expertise! You’ll work with amazing talent across each department who are passionate about our mission. We’re out of the box thinkers, so your unique ideas and approaches for conceiving a product and/or feature will be welcome. You’ll have a voice in the organization as you work with a seasoned executive team, a supportive board and in a proven market that our customers are excited about.

One of JumpCloud's three core values is to “Build Connections.” To us that means creating " human connection with each other regardless of our backgrounds, orientations, geographies, religions, languages, gender, race, etc. We care deeply about the people that we work with and want to see everyone succeed." - Rajat Bhargava, CEO

JumpCloud is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Site Reliability Engineer - India
Site Reliability Engineer - India

JumpCloud Inc. • Stati Uniti

Remoto
USD 120.000 - 180.000
Senior Vice President of Global Customer Success & Support - United States
Senior Vice President of Global Customer Success & Support - United States

JumpCloud Inc. • Atlanta (GA)

Remoto
USD 250.000 - 500.000
Senior Vice President of Global Customer Success & Support - United States
Senior Vice President of Global Customer Success & Support - United States

JumpCloud Inc. • San Francisco (CA)

Remoto
USD 250.000 - 400.000
Senior Quality Engineer - India
Senior Quality Engineer - India

JumpCloud Inc. • Stati Uniti

Remoto
USD 100.000 - 130.000
Sales Engineer - United States JumpCloud · Denver, CO Full-time · Remote $94,000–172,000 1 hour ago
Sales Engineer - United States JumpCloud · Denver, CO Full-time · Remote $94,000–172,000 1 hour ago

Emploive • Denver (CO)

Ibrido
USD 94.000 - 172.000
Health plans (medical, dental, vision)
HSA plan with employer contribution
FSA
+4
Sales Engineer - United States
Sales Engineer - United States

Lever, Inc. • Denver (CO)

Remoto
USD 90.000 - 130.000
Sales Engineer - United States
Sales Engineer - United States

Wwshemi • Northern (KY)

Remoto
USD 90.000 - 140.000
Account Executive - United States
Account Executive - United States

Lever, Inc. • Denver (CO)

Remoto
USD 175.000 - 225.000
Medical plans
Dental plans
Vision plans
+2
Account Executive - United States
Account Executive - United States

JumpCloud Inc. • Denver (CO)

In loco
USD 175.000 - 225.000
Sales Engineer - United States
Sales Engineer - United States

JumpCloud Inc. • Stati Uniti

Remoto
USD 94.000 - 172.000
HSA plan with employer contribution
Dental plans
Vision insurance
+6