Senior Site Reliability Engineer

Remote Jobs

United States

Remote

USD 150,000 - 210,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Quarterly Bonuses
Unlimited PTO
Unlimited Paid Sick Leave
Remote & Flexible Working
Private Medical Insurance
Financial Support for Life Events
Professional Development Budget
International Exposure
Regular Company Events

Job summary

Embedded Shishya is seeking a Senior Site Reliability Engineer to join the Infrastructure Squad, a lean, senior team with high ownership. You will maintain reliability, performance, and stability in a fast-paced, real-time production environment.

Responsibilities include managing alerts, on-call rotations, incident RCA, and building observability across a Kubernetes (EKS) stack, while advancing CI/CD pipelines and IaC practices.

Qualifications

  • 5+ years of hands-on SRE/DevOps experience in large-scale systems.
  • Deep knowledge of Kubernetes and cloud infrastructure.
  • Experience with GitOps tooling (FluxCD/ArgoCD) and incident response.
  • Strong scripting in Python/Go/Node.js.
  • Excellent on-call and incident management skills.

Responsibilities

  • Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time.
  • Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic environment.
  • Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence.
  • Build and continuously improve monitoring, alerting, and observability across the Kubernetes ecosystem.
  • Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools.
  • Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues.
  • Maintain and evolve CI/CD pipelines and infrastructure-as-code practices.
  • Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment.
  • Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance.
  • Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load.

Skills

Kubernetes
GitOps
Incident response
AWS
Terraform
Docker
CI/CD
Datadog
Prometheus
Grafana
ELK/CloudWatch
Networking
Scripting (Python/Go/Node.js)
Git
PagerDuty/Opsgenie

Tools

Kubernetes
FluxCD
ArgoCD
AWS
Terraform
Docker
Datadog
Prometheus
Grafana
ELK
CloudWatch
Networking
Python/Go/Node.js scripting

Job description

About the Role

We’re looking for a Senior Site Reliability Engineer to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands‑on role at the core of a high‑traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast‑paced environment.

You’ll be working on real‑time production challenges, handling incidents, managing alerts, and being part of a critical on‑call rotation. This role requires resilience, strong decision‑making under pressure, and a proactive mindset to continuously improve systems operating at scale.

If you thrive in high‑load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.

Key Responsibilities
  • Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time
  • Participate in 24/7 on‑call rotations, taking full ownership of production stability in a high‑traffic (5–7k RPS) environment
  • Investigate incidents, perform root cause analysis, and implement long‑term fixes to prevent recurrence
  • Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem
  • Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)
  • Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues
  • Maintain and evolve CI/CD pipelines and infrastructure‑as‑code practices
  • Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment
  • Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance
  • Handle environment‑specific requests and ensure smooth day‑to‑day platform operations under constant load
Requirements
  • Strong hands‑on experience with Kubernetes (deployment, scaling, troubleshooting) in high‑load environments
  • Experience with GitOps tools such as FluxCD or ArgoCD
  • Proven experience in incident response, root cause analysis, and postmortems in production systems
  • Solid experience with AWS, Terraform, Docker, and CI/CD pipelines
  • Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch
  • Strong understanding of networking concepts and protocols
  • Proficiency in at least one scripting language (e.g. Python, Go, Node.js)
  • Experience working with version control systems (Git)
  • Familiarity with incident management tools like PagerDuty, Opsgenie, or similar
  • Ability to operate effectively in a fast‑paced, high‑pressure environment with strong ownership and accountability
  • Proactive, resilient mindset with a focus on continuous improvement and system stability
What We Offer
  • Competitive Salary
  • Quarterly Bonuses
  • Unlimited Paid Time Off
  • Unlimited Paid Sick Leave
  • Remote & Flexible Working
  • Private Medical Insurance
  • Financial Support for Life Events
  • Professional Development Budget
  • International Exposure
  • Regular Company Events

*Benefits may vary depending on location and contractual agreement

Recruitment Process
  1. HR Interview (30-45 min)
  2. Technical interview (90 min)
  3. Final Interview with C-level (60 min)

By submitting your application, you acknowledge that your personal data will be processed in accordance with our Privacy Policy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY
Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY

Ardan Labs • United States

Remote
USD 140,000 - 210,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 240,000
Hybrid schedule
Equity
Health, dental, and vision
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 240,000
Hybrid schedule 3 days/week in NYC
Equity in growing company
Health/dental/vision + 401k
+2
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • San Francisco (CA)

On-site
USD 140,000 - 180,000