Principal AI Reliability Engineer — Production Systems

Meta

Denver (CO)

On-site

USD 271,000 - 347,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

Meta seeks an experienced Principal Software Engineer to build next-gen AI-driven reliability systems at scale. You will prototype, ship, and operate an AI agent that autonomously investigates production incidents under human supervision, while driving safety, context engineering, and guardrails across the platform.

You will lead architecture, experimentation, and deployment, collaborating with Infra, AI, Product, and Reliability teams to push the boundaries of agentic systems in production.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 12+ years of software engineering experience, including experience building and operating large-scale distributed or infrastructure systems.
  • Experience setting technical direction and leading complex, multi-year engineering efforts across organizational boundaries.
  • Experience applying AI or machine learning systems to production problems.
  • Demonstrated experience moving from technical concept to production deployment and measurable impact.
  • Experience diagnosing complex production systems using telemetry, code, configuration, and dependency information.
  • Experience coding in languages such as C++, Java, Python, Rust, or equivalent.
  • Recent hands-on experience building and shipping complex production systems, with the ability to move directly between architecture, experimentation, debugging, and implementation.
  • Experience influencing senior engineers and leaders without direct organizational authority.

Responsibilities

  • Define the technical vision and architecture for agentic reliability systems across Meta.
  • Personally design, code, and ship production agentic systems for complex infrastructure problems.
  • Lead the development of an AI agent for production incident investigation and mitigation, advancing it toward accurate, trusted, and safely supervised autonomous action.
  • Develop major improvements in agent reasoning, context, tool use, planning, evaluation, learning, and safe execution.
  • Explore and apply techniques including automated hill climbing, fine-tuning, reinforcement learning, model routing, distillation, pruning, and inference optimization, selecting the simplest approach that produces measurable gains.
  • Identify new high-value applications of AI across incident prevention, detection, mitigation, observability, and infrastructure operations.
  • Move rapidly from ambiguous problems to prototypes, validate them against real production workloads, and develop successful approaches into reliable systems at scale.
  • Build evaluation and experimentation systems that connect agent quality to outcomes such as investigation accuracy, successful mitigation, incident duration, and reduced operational work.
  • Build closed-loop improvement systems that turn production outcomes into evaluations, experiments, and better agent behavior.
  • Establish architectures and guardrails for production actions, including authorization, independent validation, auditability, rollback, and human oversight.
  • Partner with Infrastructure, AI, Product, and Reliability leaders to integrate agentic capabilities into Meta’s production ecosystem.
  • Influence technical strategy across organizations and mentor other engineers working on distributed systems and applied AI.

Skills

Distributed systems
AI/ML in production
C++/Java/Python/Rust
Telemetry/Observability
Mentoring/leadership
Autonomous systems

Education

Bachelor's degree in CS/Engineering

Job description

Meta seeks an experienced Principal Software Engineer to build next-gen AI-driven reliability systems at scale. You will prototype, ship, and operate an AI agent that autonomously investigates production incidents under human supervision, while driving safety, context engineering, and guardrails across the platform.

You will lead architecture, experimentation, and deployment, collaborating with Infra, AI, Product, and Reliability teams to push the boundaries of agentic systems in production.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Reliability Engineer – Production Agentics
Principal AI Reliability Engineer – Production Agentics

Meta • Lincoln (NE)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer, Production Systems
Principal AI Reliability Engineer, Production Systems

Meta • Honolulu (HI)

On-site
USD 271,000 - 347,000
Lead AI Reliability Engineer for Production Systems
Lead AI Reliability Engineer for Production Systems

Meta • Boise (ID)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer
Principal AI Reliability Engineer

Meta • Montpelier (VT)

On-site
USD 271,000 - 347,000
Lead Principal: AI-Driven Production Reliability Engineer
Lead Principal: AI-Driven Production Reliability Engineer

Meta • Helena (MT)

On-site
USD 271,000 - 347,000
Lead AI Reliability Engineer for Production Agents
Lead AI Reliability Engineer for Production Agents

Meta • Augusta (ME)

On-site
USD 271,000 - 347,000
Principal AI Agentic Reliability Engineer
Principal AI Agentic Reliability Engineer

Meta • Topeka (KS)

On-site
USD 271,000 - 347,000
Equity
Bonus
Benefits
Senior AI-Driven Reliability Engineer for Production
Senior AI-Driven Reliability Engineer for Production

Meta • United States

Remote
USD 180,000 - 240,000
Principal AI Reliability Engineer
Principal AI Reliability Engineer

Meta • Salem (OR)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer — Autonomous Incident Agent
Principal AI Reliability Engineer — Autonomous Incident Agent

Meta • Carson City (NV)

On-site
USD 271,000 - 347,000