Site Reliability Engineer

GCS Recruitment

Mount Laurel Township (NJ)

On-site

USD 110,000 - 170,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

GCS Recruitment is hiring a Site Reliability Engineer III to design, implement, and maintain scalable cloud infrastructure for an AI platform. The role emphasizes Kubernetes environments, IaC with Terraform, and robust observability across production systems.

You'll deploy AI/ML applications in multi-cloud contexts, build CI/CD pipelines with GitHub Actions, and collaborate with engineering teams to improve reliability and performance.

Qualifications

  • Bachelor's degree in Computer Science or a related technical field (or equivalent experience).
  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
  • Strong understanding of distributed systems, software design, algorithms, and system performance.
  • Experience programming or scripting with Python, Bash, Go, Java, or C/C++.

Responsibilities

  • Design, implement, and support highly available, scalable cloud infrastructure.
  • Maintain and improve Kubernetes-based production environments.
  • Deploy and support AI/ML applications across cloud platforms.
  • Automate infrastructure provisioning and operational processes using Terraform and scripting.
  • Build and maintain CI/CD pipelines using GitHub Actions.
  • Monitor system health and performance using Prometheus, Grafana, Datadog, CloudWatch, Elasticsearch.
  • Troubleshoot production issues across distributed systems and cloud infrastructure.
  • Partner with development teams to improve application reliability and performance.
  • Participate in incident response, root cause analysis, and continuous service improvements.
  • Support large-scale Kubernetes clusters and cloud-native workloads.

Skills

Python
Bash
Go
Java
C/C++
Distributed systems
Troubleshooting
Communication
Collaboration

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Kubernetes
Docker
Amazon EKS
Google GKE
GitHub Actions
Prometheus

Job description

Site Reliability Engineer III (AI Platform)

Location: Mount Laurel, NJ (Onsite)
Duration: Contract
Experience: 4+ years

About the Role

We are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible for building and maintaining the infrastructure behind enterprise AI and machine learning applications. This is an exciting opportunity to work on large-scale distributed systems, Kubernetes environments, and cloud-native platforms that power next-generation AI solutions.

The ideal candidate has a strong background in cloud infrastructure, Kubernetes, Infrastructure as Code, observability, and automation. Experience supporting production environments at scale is essential.

Responsibilities
  • Design, implement, and support highly available, scalable, and secure cloud infrastructure.
  • Maintain and improve Kubernetes-based production environments.
  • Deploy and support AI/ML applications across cloud platforms.
  • Automate infrastructure provisioning and operational processes using Terraform and scripting.
  • Build and maintain CI/CD pipelines using GitHub Actions.
  • Monitor system health and performance using Prometheus, Grafana, Datadog, CloudWatch, Elasticsearch, and logging tools.
  • Troubleshoot production issues across distributed systems, cloud infrastructure, networking, and containerized applications.
  • Partner with development teams to improve application reliability, scalability, and performance.
  • Participate in incident response, root cause analysis, and continuous service improvements.
  • Support large-scale Kubernetes clusters and cloud-native workloads.
Required Qualifications
  • Bachelor's degree in Computer Science or a related technical field (or equivalent experience).
  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
  • Strong understanding of distributed systems, software design, algorithms, and system performance.
  • Experience programming or scripting with one or more of the following:
    • Python
    • Bash
    • Go
    • Java
    • C/C++
  • Strong troubleshooting and problem-solving skills.
  • Excellent communication and collaboration abilities.
Required Technical Skills
Cloud Platforms
  • AWS (required)
  • Azure or GCP
Containers & Orchestration
  • Kubernetes (required)
  • Docker
  • Amazon EKS
  • Google GKE/GKS
Infrastructure as Code
  • Terraform (required)
CI/CD
  • GitHub Actions
  • CI/CD pipeline automation
Monitoring & Observability
  • Prometheus (required)
  • Grafana
  • Datadog
  • CloudWatch
  • Elasticsearch
  • Centralized logging solutions
Additional Technologies
  • MySQL
  • Kafka
Preferred Qualifications
  • Experience supporting AI/ML platforms or machine learning infrastructure.
  • Experience with GPU-based workloads.
  • Knowledge of AI model deployment or MLOps.
  • Experience automating operational processes.
  • Familiarity with enterprise-scale Kubernetes environments.
Nice to Have
  • Python or Go development experience.
  • Experience with ETL workflows.
  • Exposure to AI/ML technologies, LLMs, or agentic AI platforms.
Keywords

Site Reliability Engineer | SRE | DevOps | Platform Engineer | Kubernetes | AWS | Terraform | Docker | EKS | GKE | GitHub Actions | Prometheus | Grafana | Datadog | CloudWatch | Elasticsearch | Kafka | Python | Bash | Go | Infrastructure as Code | Observability | AI Platform | Machine Learning | Cloud Infrastructure | Distributed Systems

GCS is acting as an Employment Business in relation to this vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
AI DevOps Engineer
AI DevOps Engineer

Alignity Solutions • New York (NY)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Cloud Platform SRE Engineer #11145
Cloud Platform SRE Engineer #11145

ECCO Select • Dallas (TX)

Hybrid
USD 110,000 - 150,000