SRE: AI/ML Platforms & GenAI Reliability

Kforce Inc

Maryland Heights (MO)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Kforce is seeking an SRE (AI/ML) in Maryland Heights, MO to design, operate, and scale reliable cloud platforms for production AI/ML and GenAI workloads. The role focuses on reliability, automation, and cost/performance of AWS infrastructure, collaborating with data science and application teams to productionize models, ML pipelines, and LLM/RAG solutions.

The ideal candidate blends classic SRE with modern MLOps and LLMOps practices, handling incident management and on-call duties while

Qualifications

  • 8+ years total experience with at least 3+ years as SRE/DevOps/Cloud Admin on AWS in production.
  • Strong hands-on with core AWS services: EC2, VPC, IAM, S3, RDS, EKS/ECS, Lambda, CloudWatch, CloudTrail
  • Proficient in Python and one of Bash; comfortable writing reusable modules and operational tooling

Responsibilities

  • Design, operate, and scale reliable cloud platforms powering production AI/ML workloads.
  • Own reliability, automation, and cost/performance of AWS infrastructure.
  • Partner with data science and application teams to productionize models, ML pipelines, and LLM/RAG solutions.
  • Blend classical SRE with modern MLOps and LLMOps practices.

Skills

Python
Bash
AWS
Terraform
CloudFormation/CDK
Kubernetes (EKS)
ML platforms
LLM/GenAI deployment
Monitoring/Observability
Incident management
Postmortem culture
Cost optimization

Tools

EC2
VPC
IAM
S3
RDS
EKS/ECS
Lambda
CloudWatch
CloudTrail
Terraform
CloudFormation/CDK

Job description

Kforce is seeking an SRE (AI/ML) in Maryland Heights, MO to design, operate, and scale reliable cloud platforms for production AI/ML and GenAI workloads. The role focuses on reliability, automation, and cost/performance of AWS infrastructure, collaborating with data science and application teams to productionize models, ML pipelines, and LLM/RAG solutions.

The ideal candidate blends classic SRE with modern MLOps and LLMOps practices, handling incident management and on-call duties while

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE (AI/ML)
SRE (AI/ML)

Kforce Inc • Maryland Heights (MO)

On-site
USD 150,000 - 210,000
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Staff SRE: Cloud, AI/ML & Platform Reliability
Staff SRE: Cloud, AI/ML & Platform Reliability

TransUnion • Chicago (IL), Northern (KY)

Hybrid
USD 113,000 - 188,000
Medical, dental, vision coverage
HSA/FSA options
Life and AD&D insurance
+5
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
Lead SRE—AI/ML Platform Reliability & Automation
Lead SRE—AI/ML Platform Reliability & Automation

Compunnel, Inc. • Austin (TX)

On-site
USD 120,000 - 160,000
Senior SRE: ML Infra at Scale, Multi-Cloud K8s
Senior SRE: ML Infra at Scale, Multi-Cloud K8s

The Consensus • New York (NY)

On-site
USD 110,000 - 140,000
Competitive compensation
100% insurance coverage
Flexible PTO policy
+3
SRE Principal: AI-Driven Reliability Leader
SRE Principal: AI-Driven Reliability Leader

UnitedHealth Group • Minnetonka (MN)

Remote
Confidential
Comprehensive benefits package
Incentive and recognition programs
Equity stock purchase
+1
Senior SRE for AI Platform Reliability & MLOps
Senior SRE for AI Platform Reliability & MLOps

Tiger Analytics • Washington

Hybrid
USD 110,000 - 160,000
SRE Architecture Lead: Reliability & Cloud Platform
SRE Architecture Lead: Reliability & Cloud Platform

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000