Senior Site Reliability Engineer

Justjoin

United States

Remote

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health benefits
Financial planning
Family benefits
Work-life balance
Time for other pursuits

Job summary

Justjoin seeks a Senior SRE to own reliability workstreams for a serverless inference platform, build automation, and drive architecture and operational decisions. You will partner with product engineering to scale GPU infrastructure and AI workloads using Kubernetes at scale.

You will lead observability efforts, implement automation to reduce toil, and contribute to incident management with runbooks and blameless post-mortems.

Qualifications

  • Expertise in SRE, infra, or platform engineering for large-scale distributed systems.
  • Kubernetes and large-scale containerization experience.
  • Define SLOs and use observability tools (Prometheus, Grafana, tracing).
  • Proficiency in Python or Go for automation and IaC (Terraform).
  • Interest in AI/ML infra, model serving, or GPU workloads.
  • Independent problem solver with accountability.
  • Collaborate with teams unfamiliar with SRE practices.

Responsibilities

  • Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed
  • Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
  • Integrating AI workloads into our existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems
  • Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation
  • Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases
  • Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Skills

SRE practices
Automation
Observability
On-call experience

Tools

Kubernetes
Terraform
Prometheus
Grafana
Distributed tracing

Job description

Do you enjoy solving complex reliability challenges for cutting-edge technology?Do you have a passion for automation and building systems that scale?As a Senior SRE, responsibilities include owning reliability workstreams for the company's serverless inference platform, building automation and tooling, and contributing to architecture and operational decisions. Opportunities exist to take ownership of critical reliability problems end-to-end, partner with product engineering teams, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads.

Responsibilities
  • Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed
  • Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
  • Integrating AI workloads into our existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems
  • Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation
  • Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases
  • Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure
Requirements
  • Demonstrate expertise in SRE, infrastructure, or platform engineering, managing large-scale distributed systems with extensive operational experience.
  • Demonstrate expertise in Kubernetes and large-scale containerization systems.
  • Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
  • Demonstrate proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code like Terraform.
  • Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads
  • Resolve issues independently while maintaining accountability throughout the process.
  • Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.

We will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life:

  • Your health
  • Your finances
  • Your family
  • Your time at work
  • Your time pursuing other endeavors

Our benefit plan options are designed to meet your individual needs and budget, both today and in the future.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 130,000 - 180,000
Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
+3
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • United States

Remote
USD 140,000 - 170,000
Medical, dental, and vision
Equity / Stock Options
Remote equipment stipend
+3
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and snacks
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • United States

On-site
USD 150,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic Technologies • Dunwoody (GA)

On-site
USD 200,000 - 225,000