Principal AI Reliability Engineer for Production Autonomy

Meta

Columbus (OH)

On-site

USD 271,000 - 347,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Meta is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability. You will lead the development of an AI agent that autonomously investigates production incidents, identifies root causes, and creates safe mitigation plans under human supervision.

This role sits at the intersection of large-scale distributed systems, incident response, and applied AI, requiring hands-on coding across architecture, experimentation, and deployment at Meta scale.

Qualifications

  • Bachelor's degree or equivalent practical experience in CS/Engineering or related field.
  • 12+ years of software engineering experience with large-scale distributed or infrastructure systems.
  • Experience leading complex multi-year engineering efforts across teams.
  • Experience applying AI or ML to production problems.

Responsibilities

  • Define the technical vision and architecture for agentic reliability systems at Meta.
  • Design, code, and ship production agentic systems for complex infrastructure problems.
  • Lead development of an AI agent for production incident investigation and mitigation.
  • Improve agent reasoning, context, tool use, planning, evaluation, learning, and safe execution.
  • Explore techniques like fine-tuning, reinforcement learning, distillation, and inference optimization.

Education

Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience

Tools

C++
Java
Python
Rust

Job description

Meta is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability. You will lead the development of an AI agent that autonomously investigates production incidents, identifies root causes, and creates safe mitigation plans under human supervision.

This role sits at the intersection of large-scale distributed systems, incident response, and applied AI, requiring hands-on coding across architecture, experimentation, and deployment at Meta scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Reliability Engineer
Principal AI Reliability Engineer

Meta • Madison (WI)

On-site
USD 271,000 - 347,000
Senior AI Reliability Engineer for Production Systems
Senior AI Reliability Engineer for Production Systems

Meta • Richmond (VA)

On-site
USD 271,000 - 347,000
Senior AI-Driven Reliability Engineer for Production
Senior AI-Driven Reliability Engineer for Production

Meta • United States

Remote
USD 180,000 - 240,000
Lead Software Engineer, AI-Driven Production Reliability
Lead Software Engineer, AI-Driven Production Reliability

Meta • Olympia (WA)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer Autonomous Incident Agent
Principal AI Reliability Engineer Autonomous Incident Agent

Meta • Hartford (CT)

On-site
USD 271,000 - 347,000
Lead AI Agentic Reliability Engineer
Lead AI Agentic Reliability Engineer

Meta • Santa Fe (NM)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • United States

Remote
USD 180,000 - 240,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Columbus (OH)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Santa Fe (NM)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Hartford (CT)

On-site
USD 271,000 - 347,000