AI Platform SRE: Reliability, Observability & Scale

Schonfeld

New York (NY)

On-site

USD 175,000 - 225,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

You’ll define reliability standards, own incident management, and drive improvements for agents, proxies, and RAG pipelines. Strong Python, AWS, Kubernetes, and CI/CD experience are required.

Qualifications

  • 5+ years of professional cloud automation or site reliability engineering experience.
  • Experience with Python
  • Experience building REST APIs and event-driven architectures.
  • Experience with relational and NoSQL databases (Postgres, MySQL, DynamoDB, ElasticSearch).
  • Experience with AWS, Kubernetes, GitHub Actions, and modern CI/CD pipelines.

Responsibilities

  • Set reliability standards: SLOs, error budgets, incident runbooks.
  • Own observability, incident response, reliability, and scalability of our AI platform.
  • Ensure high availability, accuracy and efficiency of agents, gateways, LLM proxies, and RAG pipelines.
  • Assist users via help channel: investigate root causes, provide solutions, monitor dependencies, communicate updates.
  • Improve core platform and developer workflows; enforce best practices across teams.

Skills

Python
REST APIs
Async architectures
AWS
Kubernetes
CI/CD
Observability
DataDog
Problem solving

Tools

DataDog
Postgres
MySQL

Job description

You’ll define reliability standards, own incident management, and drive improvements for agents, proxies, and RAG pipelines. Strong Python, AWS, Kubernetes, and CI/CD experience are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, AI Platform & Observability
Site Reliability Engineer, AI Platform & Observability

Schonfeld • Northern (KY), New York (NY)

Hybrid
USD 175,000 - 225,000
Performance bonus
Competitive benefits package
Senior AI SRE: Build Reliable GenAI Platforms
Senior AI SRE: Build Reliable GenAI Platforms

Charles Schwab • Austin (TX)

On-site
USD 150,000 - 190,000
401(k) match
Sabbatical after 5 years of service
Parental leave
+2
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
AWS DevOps & SRE Engineer for Agentic AI Platform
AWS DevOps & SRE Engineer for Agentic AI Platform

Strategic Business Systems • Chantilly (VA)

On-site
USD 180,000 - 270,000
Comprehensive medical, dental, and V/V
401(k) with company match
Paid Time Off and holidays
+1
Remote SRE for AI Platform - Scale & Reliability
Remote SRE for AI Platform - Scale & Reliability

United States Digital Space LLC • Paris (TX)

On-site
USD 80,000 - 111,000
Principal SRE - AI-Driven Platform & AIOps
Principal SRE - AI-Driven Platform & AIOps

Oracle • United States

Remote
USD 86,000 - 200,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Flexible Vacation
Lead SRE—AI/ML Platform Reliability & Automation
Lead SRE—AI/ML Platform Reliability & Automation

Compunnel, Inc. • Austin (TX)

On-site
USD 120,000 - 160,000
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2