AI Reliability Engineering Lead — Production-Grade AI

Socket.dev

Cincinnati (OH)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

84.51° is seeking a Senior Manager, AI Reliability Engineering to lead the new discipline ensuring enterprise AI is trustworthy at scale. You will define standards, build the reliability function, and partner with Platform, Model and Applied AI teams to embed resilience from architecture forward.

The role focuses on establishing SLIs/SLOs, on-call readiness, and automated quality controls while guiding a diverse team across Platform Engineering, AI, and Security to enable responsible AI

Qualifications

  • 10+ years in software/production engineering with leadership experience (3+ years leading engineering teams).
  • Track record of building capabilities, standards, automation or platforms that improve system reliability at scale.
  • Deep technical depth in distributed systems, Kubernetes, cloud infrastructure and modern observability engineering.

Responsibilities

  • Build the Discipline (0-to-1) and define production-grade AI standards.
  • Stand up the AI Reliability Engineering function with charter, operating model and roadmap.
  • Position reliability as an enabler of AI adoption and velocity across the enterprise.
  • Embed resilience, testability, observability and safe failure modes into AI systems from architecture forward.
  • Build reliability tooling and automation, including self-healing, quality-regression detection and safe deployment controls.
  • Establish SLOs, error budgets, and reliability scorecards guiding architecture and roadmap decisions.
  • Own production readiness and agent onboarding, with lifecycle controls for agents and runtimes.
  • Own live observability and production-quality signals across model and agent behavior.
  • Drive cost and performance optimization for AI workloads at scale.
  • Lead hiring, coaching and growth of a multidisciplinary reliability team.

Skills

Distributed systems
Kubernetes
Cloud infrastructure (GCP/Azure)
SRE leadership
LLM/AI in production
Executive communication
Production readiness

Tools

LangSmith
MLflow
Arize
Fiddler

Job description

84.51° is seeking a Senior Manager, AI Reliability Engineering to lead the new discipline ensuring enterprise AI is trustworthy at scale. You will define standards, build the reliability function, and partner with Platform, Model and Applied AI teams to embed resilience from architecture forward.

The role focuses on establishing SLIs/SLOs, on-call readiness, and automated quality controls while guiding a diverse team across Platform Engineering, AI, and Security to enable responsible AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Reliability Engineering Leader — Production-Grade AI
AI Reliability Engineering Leader — Production-Grade AI

84.51˚ • Cincinnati (OH)

On-site
USD 190,000 - 270,000
Head of Applied AI & Systems Reliability
Head of Applied AI & Systems Reliability

Relativity • Washington

Hybrid
USD 208,000 - 312,000
Senior AI Reliability Engineer - Scale & Resilience
Senior AI Reliability Engineer - Scale & Resilience

Anthropic • San Francisco (CA)

Hybrid
USD 325,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Senior AI Reliability Engineer — Platform & Observability
Senior AI Reliability Engineer — Platform & Observability

Flatiron Health • Berlin (NH)

On-site
USD 140,000 - 190,000
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

Socket.dev • Cincinnati (OH)

On-site
USD 180,000 - 240,000
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

84.51˚ • Cincinnati (OH)

On-site
USD 190,000 - 270,000
Platform Reliability Leader for AI Infrastructure
Platform Reliability Leader for AI Infrastructure

Etched.ai, Inc. • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental and vision coverage
Housing subsidy
Relocation support
+3
Senior Manager, AI-Driven Software & Reliability
Senior Manager, AI-Driven Software & Reliability

Anduril Industries • Costa Mesa (CA)

On-site
USD 220,000 - 292,000
Equity incentives
Top-tier benefits
AI Performance & Production Reliability Manager
AI Performance & Production Reliability Manager

Devoted Health Services, Inc • Massachusetts

On-site
USD 110,000 - 172,000
Employer sponsored health, dental and vision plan
Generous paid time off
$100 monthly mobile or internet stipend
+4
Senior AI Science Leader - Scalable, Reliable Systems
Senior AI Science Leader - Scalable, Reliable Systems

Relativity • Illinois

Hybrid
USD 208,000 - 312,000