Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI

Austin (TX)

On-site

USD 140,000 - 190,000

Full time

20 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Future Secure AI is seeking a Site Reliability Engineer to design, build, and operate the platforms powering AI Co‑Workers. This hands‑on role requires ownership of reliability end‑to‑end and close collaboration with product, AI, and engineering teams.

You will design and run production infrastructure, manage Kubernetes workloads across EKS/AKS/GKE, use Terraform and Helm, and drive reliability metrics with SLIs, SLOs, and SLAs. On-call rotations and post‑mortems are part of the role.

Qualifications

  • Bachelor's degree in Computer Science, Information Systems, or related field
  • 5+ years of professional experience in Site Reliability Engineering or DevOps Engineering
  • Kubernetes experience designing, building, or operating workloads on EKS, AKS, GKE, or self‑managed Kubernetes
  • Terraform experience for infrastructure provisioning and automation
  • Helm experience for Kubernetes application deployment
  • Professional experience using at least two programming or scripting languages such as Python, Go, Java, Bash, PowerShell, or Ruby
  • Direct experience with reliability engineering, on‑call rotations, incident response, post‑mortems, and toil reduction
  • DevOps or DevSecOps experience, including CI/CD ownership, infrastructure automation, and security considerations

Responsibilities

  • Design, build, and operate reliable production infrastructure supporting AI Co‑Workers
  • Own Kubernetes‑based platforms used to deploy and run AI workloads
  • Build and maintain infrastructure as code using Terraform
  • Implement and maintain Helm‑based deployment workflows
  • Define, measure, and improve system reliability using SLIs, SLOs, and SLAs
  • Participate in on‑call rotation, incident response, root cause analysis, and post‑mortems
  • Reduce operational toil through automation and engineering improvements
  • Build and improve observability across monitoring, logging, and alerting
  • Partner closely with engineers to ensure systems are resilient, scalable, and secure
  • Operate across build, deploy, and operate phases of the software lifecycle

Skills

SRE/DevOps
Kubernetes
Terraform
Helm
CI/CD ownership
Python/Go/Java
On-call support
Security practices

Education

Bachelor's degree
Masters preferred

Tools

Terraform
Helm
ArgoCD
GitOps

Job description

Future Secure AI is seeking a Site Reliability Engineer to design, build, and operate the platforms powering AI Co‑Workers. This hands‑on role requires ownership of reliability end‑to‑end and close collaboration with product, AI, and engineering teams.

You will design and run production infrastructure, manage Kubernetes workloads across EKS/AKS/GKE, use Terraform and Helm, and drive reliability metrics with SLIs, SLOs, and SLAs. On-call rotations and post‑mortems are part of the role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer – AI Platform
Senior Site Reliability Engineer – AI Platform

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 130,000
Competitive salary
Flexible work environment
High-performance culture
Site Reliability Engineer Austin, TX
Site Reliability Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Site Reliability Engineer
Site Reliability Engineer

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Platform DevOps Engineer Austin, TX
Platform DevOps Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 130,000
Competitive salary
Flexible work environment
High-performance culture
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior Site Reliability Engineer - AI Cloud Platform
Senior Site Reliability Engineer - AI Cloud Platform

Lambda • United States

Hybrid
USD 160,000 - 220,000
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Lake Buena Vista (FL)

On-site
USD 120,000 - 150,000