Senior SRE - AI Platform & Cloud Infra

GCS Recruitment

Philadelphia (Philadelphia County)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

GCS Recruitment seeks a Site Reliability Engineer III to support an AI Platform Engineering team responsible for building and maintaining the infrastructure behind enterprise AI and machine learning applications. This role focuses on large-scale distributed systems, Kubernetes, and cloud-native platforms that power next-generation AI solutions.

The ideal candidate has strong cloud infrastructure, Kubernetes, Infrastructure as Code, observability, and automation experience.

Qualifications

  • Bachelor's degree in Computer Science or related field, or equivalent experience.
  • 4+ years of experience as Site Reliability Engineer, DevOps, Platform Engineer, or Cloud Infrastructure.
  • Strong understanding of distributed systems, software design, algorithms, and system performance.
  • Experience programming or scripting with Python, Bash, Go, Java, or C/C++.
  • Experience with Kubernetes-based production environments.

Responsibilities

  • Design, implement, and support highly available, scalable cloud infrastructure.
  • Maintain and improve Kubernetes-based production environments.
  • Deploy and support AI/ML applications across cloud platforms.
  • Automate infrastructure provisioning and operational processes using Terraform and scripting.
  • Build and maintain CI/CD pipelines using GitHub Actions.
  • Monitor system health and performance with Prometheus, Grafana, Datadog, CloudWatch, Elasticsearch, and logging tools.
  • Troubleshoot production issues across distributed systems, cloud infrastructure, networking, and containerized applications.
  • Partner with development teams to improve application reliability, scalability, and performance.
  • Participate in incident response, root cause analysis, and continuous service improvements.
  • Support large-scale Kubernetes clusters and cloud-native workloads.

Skills

Python
Go
Bash
Java
C/C++
Observability
Communication

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Docker
Terraform
GitHub Actions
Prometheus
Grafana
Datadog
CloudWatch
Elasticsearch
MySQL
Kafka
AWS
GKE
EKS

Job description

GCS Recruitment seeks a Site Reliability Engineer III to support an AI Platform Engineering team responsible for building and maintaining the infrastructure behind enterprise AI and machine learning applications. This role focuses on large-scale distributed systems, Kubernetes, and cloud-native platforms that power next-generation AI solutions.

The ideal candidate has strong cloud infrastructure, Kubernetes, Infrastructure as Code, observability, and automation experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

GCS Recruitment • Philadelphia

On-site
USD 120,000 - 160,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Kubernetes SRE for AI Infra & GPU Clusters
Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits