Director-Level SRE for GenAI Platform & Cloud

474 MS Services Group, Inc.

Alpharetta (GA)

On-site

USD 125,000 - 195,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Morgan Stanley Alpharetta seeks an experienced Site Reliability Engineer to join the AI Platform team, supporting GenAI workloads in a regulated financial environment. You will build automation, manage IaC, and ensure reliability across training, inference, and data pipelines.

The role requires strong coding, Kubernetes and cloud experience, plus a track record of driving reliability improvements in large-scale systems. On-site in Alpharetta with cross-team collaboration and on-call duties.

Qualifications

  • 5+ years of production experience in SRE / Infrastructure / ops for large-scale systems.
  • Strong programming/scripting skills (Python, Go, Java, or equivalent).
  • Deep experience with containerization (Docker), orchestration (Kubernetes, etc.).
  • Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.).
  • Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures.
  • Experience with monitoring/observability/logging/alerting tools (Prometheus, Grafana, ELK/EFK, Datadog).
  • Experience in regulated environments (financial services, compliance, audit, security) is a strong plus.

Responsibilities

  • Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving).
  • Design and build automation for core platform capabilities, reducing manual toil.
  • Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
  • Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards.
  • Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation.
  • Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting.
  • Optimize cost vs. performance tradeoffs in large-scale compute environments.
  • Harden systems for security, compliance, auditability, and data governance.
  • Collaborate across teams to ensure safe deployment, rollout, rollback, and integration of new systems.

Skills

Python
Go
Java
Kubernetes
Terraform
CI/CD
Prometheus
Grafana
Linux
Networking

Education

Bachelor's or Master’s in CS

Tools

Docker
Kubernetes
Terraform
CloudFormation
OpenTelemetry

Job description

Morgan Stanley Alpharetta seeks an experienced Site Reliability Engineer to join the AI Platform team, supporting GenAI workloads in a regulated financial environment. You will build automation, manage IaC, and ensure reliability across training, inference, and data pipelines.

The role requires strong coding, Kubernetes and cloud experience, plus a track record of driving reliability improvements in large-scale systems. On-site in Alpharetta with cross-team collaboration and on-call duties.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE Lead for GenAI Platform & Cloud
Senior SRE Lead for GenAI Platform & Cloud

PowerToFly • Alpharetta (GA)

On-site
USD 130,000 - 160,000
Comprehensive employee benefits
Flexible work environment
Senior SRE Lead: AI/ML Reliability & Automation
Senior SRE Lead: AI/ML Reliability & Automation

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 210,000
Senior AI-Driven Platform SRE Lead
Senior AI-Driven Platform SRE Lead

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 170,000 - 210,000
SRE III: AI-Driven Cloud Reliability Engineer
SRE III: AI-Driven Cloud Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 140,000 - 170,000
Senior AI Platform Engineer & Architect (GenAI)
Senior AI Platform Engineer & Architect (GenAI)

Morgan-Stanley • Town of Islip (NY)

On-site
USD 155,000 - 215,000
Comprehensive employee benefits
Senior SRE: AI-Driven Reliability for Scalable Systems
Senior SRE: AI-Driven Reliability for Scalable Systems

JPMorgan Chase & Co. • Wilmington (DE)

On-site
USD 120,000 - 160,000
Senior Platform SRE | AI-Driven Reliability Leader
Senior Platform SRE | AI-Driven Reliability Leader

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 230,000
AWS Certifications
Lead Platform Engineer: Cloud-Native & Platform Automation
Lead Platform Engineer: Cloud-Native & Platform Automation

Morgan Stanley • Alpharetta (GA)

On-site
USD 115,000 - 225,000
Medical benefits
401(k) plan
Paid time off
+1
Senior SRE: AI-Driven Reliability & Cloud Infra
Senior SRE: AI-Driven Reliability & Cloud Infra

Next Frontier Capital • Houston (TX)

On-site
USD 120,000 - 170,000
Health care coverage
On-site health and wellness centers
Retirement savings plan
+3
Senior Lead SRE: AI/ML Data Platforms & Reliability
Senior Lead SRE: AI/ML Data Platforms & Reliability

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 100,000 - 150,000