Senior AI Reliability Engineer for Production Systems

Meta

Richmond (VA)

On-site

USD 271,000 - 347,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Meta is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability at scale.

The role sits at the intersection of large-scale distributed systems, incident response, and applied AI, leading the development of an AI agent that can autonomously investigate production incidents, identify root causes, create safe mitigation plans, and execute them under human supervision.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 12+ years of software engineering experience, including experience building and operating large-scale distributed or infrastructure systems.
  • Experience setting technical direction and leading complex, multi-year engineering efforts across organizational boundaries.
  • Experience applying AI or machine learning systems to production problems.
  • Demonstrated experience moving from technical concept to production deployment and measurable impact.
  • Experience diagnosing complex production systems using telemetry, code, configuration, and dependency information.
  • Experience coding in languages such as C++, Java, Python, Rust, or equivalent.

Responsibilities

  • Define the technical vision and architecture for agentic reliability systems across Meta.
  • Personally design, code, and ship production agentic systems for complex infrastructure problems.
  • Lead the development of an AI agent for production incident investigation and mitigation.
  • Develop major improvements in agent reasoning, context, tool use, planning, evaluation, learning, and safe execution.
  • Explore and apply techniques including automated hill climbing, fine-tuning, reinforcement learning, model routing, distillation, pruning, and inference optimization.
  • Identify new high-value applications of AI across incident prevention, detection, mitigation, observability, and infrastructure operations.
  • Move rapidly from ambiguous problems to prototypes, validate them against real production workloads, and develop successful approaches into reliable systems at scale.
  • Build evaluation and experimentation systems that connect agent quality to outcomes such as investigation accuracy, successful mitigation, incident duration, and reduced operational work.
  • Build closed-loop improvement systems that turn production outcomes into evaluations, experiments, and better agent behavior.
  • Establish architectures and guardrails for production actions, including authorization, independent validation, auditability, rollback, and human oversight.
  • Partner with Infrastructure, AI, Product, and Reliability leaders to integrate agentic capabilities into Meta’s production ecosystem.
  • Influence technical strategy across organizations and mentor other engineers working on distributed systems and applied AI.

Skills

Architectural design
Hands-on coding
AI agent development
Agent reasoning
RL & tuning
AI application discovery
Prototype development
Experimentation systems
Feedback loops
Production guardrails
Cross-functional collaboration
Mentorship

Education

Bachelor's degree in CS/engineering

Tools

C++
Python
Java
Rust

Job description

Meta is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability at scale.

The role sits at the intersection of large-scale distributed systems, incident response, and applied AI, leading the development of an AI agent that can autonomously investigate production incidents, identify root causes, create safe mitigation plans, and execute them under human supervision.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI-Driven Reliability Engineer for Production
Senior AI-Driven Reliability Engineer for Production

Meta • United States

Remote
USD 180,000 - 240,000
Lead Software Engineer, AI-Driven Production Reliability
Lead Software Engineer, AI-Driven Production Reliability

Meta • Olympia (WA)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer for Production Autonomy
Principal AI Reliability Engineer for Production Autonomy

Meta • Columbus (OH)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer
Principal AI Reliability Engineer

Meta • Madison (WI)

On-site
USD 271,000 - 347,000
Lead AI Agentic Reliability Engineer
Lead AI Agentic Reliability Engineer

Meta • Santa Fe (NM)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer Autonomous Incident Agent
Principal AI Reliability Engineer Autonomous Incident Agent

Meta • Hartford (CT)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • United States

Remote
USD 180,000 - 240,000
Senior AI Engineer for Production Agentic Systems
Senior AI Engineer for Production Agentic Systems

Bot Auto • Houston (TX)

On-site
USD 160,000 - 220,000
Health insurance
Senior AI Infra & Production Reliability Lead
Senior AI Infra & Production Reliability Lead

The Mutual Group • Iowa (LA), Northern (KY)

Hybrid
USD 130,000 - 150,000
Hybrid/remote options
Educational assistance
401(K) with company match
+1
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Columbus (OH)

On-site
USD 271,000 - 347,000