AI SRE: Scale, Resilience & GPU Inference

Seekr

Austin (TX)

Hybrid

USD 140,000 - 200,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity Ownership – RSUs
Unlimited PTO
14 paid company holidays
Hybrid work environment
401(k) with Company Match
Comprehensive Health & Wellness
Parental Leave

Job summary

Seekr is seeking an AI Site Reliability Engineer to ensure reliability, scalability, and safe operation of AI-powered services and infrastructure. You will blend software engineering with AI/ML operational expertise to deploy, monitor, and operate models, APIs, data pipelines, and platform services in production.

You will partner with AI/ML, Platform, Security, and Product teams to build dependable, performant, resilient systems ready to scale, with a focus on SLIs/SLOs, automation, and

Qualifications

  • Bachelor's degree or equivalent practical experience in CS/engineering.
  • 4+ years in Site Reliability Engineering, DevOps, or similar roles.
  • Strong knowledge of distributed systems, cloud infra, networks, containers, and Kubernetes.
  • Experience building and operating CI/CD pipelines and progressive deployment strategies.
  • Proficiency with infrastructure-as-code and tools like Terraform.
  • Experience with monitoring, logging, tracing, alerting, and incident management (Prometheus, Grafana, OpenTelemetry).
  • Strong scripting and software development in Python or Go.
  • Knowledge of SLIs/SLOs, capacity planning, and production-readiness reviews.

Responsibilities

  • Build safe release processes with CI/CD automation, progressive deployments, validation, rollback, and health metrics.
  • Define and operate SLIs/SLOs for AI APIs and critical services.
  • Develop observability through metrics, logs, traces, dashboards, and SLO alerts.
  • Lead on-call rotation, incident response, and blameless postmortems.
  • Reduce toil with infrastructure as code, automation, and DR planning.
  • Design automated load, stress, and scalability tests for AI production traffic.
  • Set performance baselines, release thresholds, and capacity forecasts.
  • Validate autoscaling, rate limiting, backpressure, graceful degradation, and recovery.
  • Scale GPU-backed model inference on Kubernetes with batching and load balancing.
  • Optimize inference engines and infra across GPUs, compute, networking, storage, and cost.
  • Support reliable training workloads including distributed training and checkpointing.
  • Collaborate with AI/ML, Platform, Security, and Product teams for production readiness.

Skills

Distributed systems
Cloud infrastructure
Kubernetes
CI/CD pipelines
Terraform
Monitoring & tracing
Prometheus
Grafana
OpenTelemetry
Python
Go
SRE / DevOps
Blameless postmortems

Education

Bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience

Tools

Terraform
Prometheus
Grafana
OpenTelemetry
Kubernetes
NVIDIA Triton
PyTorch
DeepSpeed
FSDP
vLLM

Job description

Seekr is seeking an AI Site Reliability Engineer to ensure reliability, scalability, and safe operation of AI-powered services and infrastructure. You will blend software engineering with AI/ML operational expertise to deploy, monitor, and operate models, APIs, data pipelines, and platform services in production.

You will partner with AI/ML, Platform, Security, and Product teams to build dependable, performant, resilient systems ready to scale, with a focus on SLIs/SLOs, automation, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
SRE for AI Platform & ML Inference Infra
SRE for AI Platform & ML Inference Infra

Cohere • New York (NY)

Hybrid
USD 140,000 - 200,000
Weekly lunch stipend
Health and dental benefits
RRSP matching
+3
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior AI Infrastructure Engineer: Scalable GPU & Kubernetes
Senior AI Infrastructure Engineer: Scalable GPU & Kubernetes

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+3
Senior AI GPU Infra SRE - Scale, Automation & Equity
Senior AI GPU Infra SRE - Scale, Automation & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior AI Infra Architect: Scalable GPU & Edge Platforms
Senior AI Infra Architect: Scalable GPU & Edge Platforms

Seekr • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Equity ownership – RSUs
Unlimited PTO
Flexible hybrid work environment
+3
Senior AI Infra Architect — Scale Enterprise AI
Senior AI Infra Architect — Scale Enterprise AI

Seekr • Washington

Hybrid
USD 170,000 - 260,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+2
Senior SRE: GPU-Driven, Global Scale & Causal AI
Senior SRE: GPU-Driven, Global Scale & Causal AI

Crossing Hurdles • San Francisco (CA)

On-site
USD 180,000 - 240,000