Senior AI Platform SRE: Scale Cloud Infra & Kubernetes

GCS Recruitment

Mount Laurel Township (NJ)

On-site

USD 110,000 - 170,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

GCS Recruitment is hiring a Site Reliability Engineer III to design, implement, and maintain scalable cloud infrastructure for an AI platform. The role emphasizes Kubernetes environments, IaC with Terraform, and robust observability across production systems.

You'll deploy AI/ML applications in multi-cloud contexts, build CI/CD pipelines with GitHub Actions, and collaborate with engineering teams to improve reliability and performance.

Qualifications

  • Bachelor's degree in Computer Science or a related technical field (or equivalent experience).
  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
  • Strong understanding of distributed systems, software design, algorithms, and system performance.
  • Experience programming or scripting with Python, Bash, Go, Java, or C/C++.

Responsibilities

  • Design, implement, and support highly available, scalable cloud infrastructure.
  • Maintain and improve Kubernetes-based production environments.
  • Deploy and support AI/ML applications across cloud platforms.
  • Automate infrastructure provisioning and operational processes using Terraform and scripting.
  • Build and maintain CI/CD pipelines using GitHub Actions.
  • Monitor system health and performance using Prometheus, Grafana, Datadog, CloudWatch, Elasticsearch.
  • Troubleshoot production issues across distributed systems and cloud infrastructure.
  • Partner with development teams to improve application reliability and performance.
  • Participate in incident response, root cause analysis, and continuous service improvements.
  • Support large-scale Kubernetes clusters and cloud-native workloads.

Skills

Python
Bash
Go
Java
C/C++
Distributed systems
Troubleshooting
Communication
Collaboration

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Kubernetes
Docker
Amazon EKS
Google GKE
GitHub Actions
Prometheus

Job description

GCS Recruitment is hiring a Site Reliability Engineer III to design, implement, and maintain scalable cloud infrastructure for an AI platform. The role emphasizes Kubernetes environments, IaC with Terraform, and robust observability across production systems.

You'll deploy AI/ML applications in multi-cloud contexts, build CI/CD pipelines with GitHub Actions, and collaborate with engineering teams to improve reliability and performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Cloud Infra, Kubernetes & Terraform (Remote)
Senior SRE: AI Cloud Infra, Kubernetes & Terraform (Remote)

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 140,000 - 190,000
Remote equipment stipend
Annual learning and development budget
Equity / Stock Options
+1
Site Reliability Engineer
Site Reliability Engineer

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Kubernetes SRE for AI Infra & GPU Clusters
Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
SRE: AI/ML Infra on Kubernetes, AWS & Terraform
SRE: AI/ML Infra on Kubernetes, AWS & Terraform

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Senior Cloud Platform SRE for Scalable AI Systems
Senior Cloud Platform SRE for Scalable AI Systems

Mistral AI • Germany (OH)

On-site
USD 130,000 - 210,000
Healthcare coverage
Relocation support
Retirement plans
+2
AI Inference Platform SRE — Build Reliable Cloud Operations
AI Inference Platform SRE — Build Reliable Cloud Operations

BBG Ventures, LLC • Bellevue (WA)

On-site
USD 133,000 - 220,000
Medical insurance
Dental insurance
Vision insurance
+5
Senior Platform Engineer, AI Scale & Cloud Foundations
Senior Platform Engineer, AI Scale & Cloud Foundations

Scale AI, Inc. • San Francisco (CA)

On-site
USD 216,000 - 270,000
Equity
Health, dental & vision
Retirement benefits
+3
Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000