AI Platform Reliability Engineer

Remote Jobs

United States

Remote

USD 101,000 - 188,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, Dental, Vision
401(k) match
Paid time off

Job summary

The AI Factory is seeking a Site Reliability Engineer to enhance the reliability of platforms and services that support AI development and deployment. You will leverage software development and automation to identify operational problems, reduce repetitive work, and help teams deliver dependable systems.

You will collaborate with engineers and partners across the AI Factory to improve observability, incident response, release validation, and platform health, contributing to reliability platforms

Qualifications

  • Experience designing, developing, and maintaining production software or automation using Go, Python, or a comparable language.
  • Experience operating or engineering Kubernetes-based platforms, including troubleshooting complex service or infrastructure issues.
  • Experience building or improving CI/CD, GitOps, infrastructure automation, or deployment workflows.
  • Experience using observability data—including metrics, logs, or traces—to diagnose problems and improve system health.

Responsibilities

  • Improve the scalability, resilience, and reliability of existing platform services and middleware to ensure they remain dependable as usage and demand grow.
  • Develop and maintain automation and operational tooling that improve platform reliability and reduce recurring manual work.
  • Help teams investigate incidents, identify contributing factors, and implement fixes that prevent repeat issues.
  • Build and improve dashboards, alerts, and other observability capabilities using metrics, logs, and traces.
  • Create automated tests and validation workflows for upgrades, releases, and changes to platform services.
  • Contribute to CI/CD and GitOps workflows that support consistent, reliable deployments.
  • Assess platform health, document findings, and work with partner teams on practical reliability improvements.
  • Participate in design reviews, code reviews, testing, and incident reviews.
  • Contribute to reliability improvements for AI Factory services, including AIF Up and tools that support health validation and investigation.

Skills

Go/Python automation
Kubernetes platforms
CI/CD/GitOps
Observability data usage
OpenShift
GitLab CI/CD
Argo CD
Argo Rollouts
Prometheus
Grafana
OpenTelemetry

Tools

OpenShift
GitLab CI/CD
Argo CD
Argo Rollouts
Prometheus
Grafana
OpenTelemetry

Job description

The AI Factory is seeking a Site Reliability Engineer to enhance the reliability of platforms and services that support AI development and deployment. You will leverage software development and automation to identify operational problems, reduce repetitive work, and help teams deliver dependable systems.

You will collaborate with engineers and partners across the AI Factory to improve observability, incident response, release validation, and platform health, contributing to reliability platforms

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Staff Site Reliability Engineer — AI-Driven Reliability
Staff Site Reliability Engineer — AI-Driven Reliability

EarnIn • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Equity
Hybrid work model
Senior Backend Engineer - AI-Driven Reliability Platform
Senior Backend Engineer - AI-Driven Reliability Platform

Affirm • Miami (FL)

On-site
USD 173,000 - 233,000
Health and wellness benefits
Remote-first culture
Competitive equity
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)

Affirm • Riverside (OH)

Remote
USD 173,000 - 233,000
Health coverage
FSA Wallets
Time off
+1
Senior SRE Platform Engineer – AI-Powered Reliability
Senior SRE Platform Engineer – AI-Powered Reliability

UiPath • Denver (CO)

Hybrid
USD 160,000 - 210,000
Remote Backend Reliability Engineer - AI-Driven Platform
Remote Backend Reliability Engineer - AI-Driven Platform

Affirm • Boulder (CO)

On-site
USD 173,000 - 255,000
Health care coverage
Flexible Spending Wallets
Time off
+1
Senior Platform Reliability Engineer – AI-Driven FinServ
Senior Platform Reliability Engineer – AI-Driven FinServ

interface.ai • San Francisco (CA)

On-site
USD 150,000 - 210,000
100% paid health, dental & vision
401(k) & financial wellness
Daily meals on us
+3
Staff SRE: AI-Driven Reliability & Platform Architect
Staff SRE: AI-Driven Reliability & Platform Architect

Devopsroles • Northern (KY)

Remote
USD 150,000 - 225,000
Equity
Benefits program
AI Platform SRE: Reliability, Observability & Scale
AI Platform SRE: Reliability, Observability & Scale

Schonfeld • New York (NY)

On-site
USD 175,000 - 225,000
Platform Reliability Leader for AI Infrastructure
Platform Reliability Leader for AI Infrastructure

Etched.ai, Inc. • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental and vision coverage
Housing subsidy
Relocation support
+3