Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment

Mount Laurel Township (NJ)

Remote

USD 130,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
Annual learning and development budget
Flexible PTO
Career Growth Within a Rapidly Scaling

Job summary

Motion Recruitment is seeking a Site Reliability Engineer for a remote, full-time role supporting a cloud platform used by millions of developers. You will help improve reliability, scalability, and performance across Linux, Kubernetes, distributed systems, GPU infrastructure, observability, and automation.

You’ll partner with Infrastructure, Product Engineering, and Support teams to define SLOs, reduce toil through automation, and lead incident response.

Qualifications

  • 5+ years in a major public cloud environment (AWS or GCP).
  • Strong Linux systems administration experience.
  • Solid networking fundamentals and troubleshooting skills.
  • Experience with containerized environments (Kubernetes).
  • Experience with monitoring/observability tools and reliability metrics.
  • Incident response and postmortem experience.
  • Programming/scripting in Python, Go, Bash, or similar.

Responsibilities

  • Hands-on engineering across Linux and Kubernetes administration.
  • Define and manage SLIs, SLOs, and reliability metrics.
  • Lead incident response and postmortems to drive improvements.
  • Automate reliability tasks to reduce operational toil.
  • Collaborate with Infra, Product Engineering, and Support teams.

Skills

AWS
GCP
Linux administration
Networking fundamentals
Kubernetes
Monitoring observability
SLIs SLOs
Incident response
Python Go Bash
Distributed systems

Tools

Prometheus
Grafana
Terraform
CI/CD tooling

Job description

Remote (USA) | Full-Time | Site Reliability Engineer

Join a rapidly growing B2B AI infrastructure company powering large-scale machine learning and AI workloads for more than one million developers worldwide. As a Site Reliability Engineer, you'll help improve the reliability, scalability, and performance of a cloud platform built on Linux, Kubernetes, distributed systems, GPU infrastructure, observability, and automation technologies. This full-time remote opportunity offers the chance to work on critical infrastructure supporting AI applications on a global scale.

As the company continues to scale its AI infrastructure platform, reliability has become a critical business function. This role sits at the center of that effort, partnering with Infrastructure, Product Engineering, and Support teams to improve uptime, strengthen observability, establish SLOs, reduce operational toil through automation, and lead incident response initiatives. The ideal candidate brings experience supporting large-scale production environments and enjoys solving complex reliability challenges while influencing engineering practices across a rapidly growing organization. This is an opportunity to gain exposure to cutting-edge AI and GPU infrastructure, take ownership of high-impact initiatives, and help shape the reliability strategy of a platform relied upon by more than one million developers.

Required Skills & Experience
  • 5+ years of experience within major public cloud environment like AWS, GCP
  • Strong Linux systems administration experience
  • Strong networking fundamentals and troubleshooting skills
  • Experience supporting containerized environments (Kubernetes preferred)
  • Experience with monitoring, alerting, and observability tools
  • Experience defining and managing SLIs, SLOs, and reliability metrics
  • Incident response and postmortem experience
  • Scripting or programming experience. Python, Go, Bash, or similar technologies
  • Distributed systems and failure scenarios
Desired Skills & Experience
  • Kubernetes
  • Prometheus, Grafana, or similar monitoring platforms
  • Experience supporting GPU infrastructure or AI/ML platforms
  • Infrastructure as Code experience (Terraform preferred)
  • CI/CD pipeline experience
What You Will Be Doing
Tech Breakdown
  • 40% Linux & Kubernetes Administration
  • 25% Monitoring, Observability & Incident Response
  • 20% Automation & Reliability Engineering
  • 15% Distributed Systems & Cloud Infrastructure
Daily Responsibilities
  • 80% Hands-On Engineering
  • 5% Management Duties
  • 15% Team Collaboration
The Offer
  • medical, dental, and vision benefits
  • Equity / Stock Options
  • Remote equipment stipend
  • Annual learning and development budget
  • Flexible PTO
  • Career Growth Within a Rapidly Scaling AI Infrastructure Company

Applicants must be currently authorized to work in the US on a full-time basis now and in the future. Sponsorship is not available for this position

#LI-JG2

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • United States

Remote
USD 140,000 - 170,000
Medical, dental, and vision
Equity / Stock Options
Remote equipment stipend
+3
Remote Senior SRE - AI Infra, Kubernetes & Terraform
Remote Senior SRE - AI Infra, Kubernetes & Terraform

Motion Recruitment • United States

Remote
USD 140,000 - 170,000
Medical, dental, and vision
Equity / Stock Options
Remote equipment stipend
+3
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

On-site
USD 140,000 - 170,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer
Site Reliability Engineer

Inclusion Services S.A • Chicago (IL)

On-site
USD 90,000 - 130,000
100% company-covered health insurance
401k plan with 4% match
15 days paid time off
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Remote Senior Site Reliability Engineer (US or Canada)
Remote Senior Site Reliability Engineer (US or Canada)

MAP SSG • United States

Remote
USD 160,000 - 210,000
Equity compensation
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Washington

On-site
USD 175,000 - 250,000
Senior SRE: Cloud, Kubernetes & Terraform (Remote)
Senior SRE: Cloud, Kubernetes & Terraform (Remote)

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 130,000 - 180,000
Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
+3