Principal AI Reliability Engineer Autonomous Incident Agent

Meta

Hartford (CT)

On-site

USD 271,000 - 347,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Meta seeks an experienced Software Engineer to build next-generation AI systems for production reliability. You will work at the intersection of large-scale distributed systems, incident response, and applied AI, leading the development of an autonomous AI agent that investigates incidents, identifies root causes, and executes safe mitigations under human supervision.

The role combines deep infrastructure expertise with practical AI deployment, requiring hands-on prototyping, evaluation, and

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 12+ years of software engineering experience with large-scale distributed or infrastructure systems.
  • Experience setting technical direction and leading complex multi-year engineering efforts.
  • Experience applying AI or machine learning systems to production problems.
  • Demonstrated experience moving from concept to production deployment and measurable impact.
  • Experience diagnosing complex production systems using telemetry, code, configuration, and dependency information.
  • Experience coding in languages such as C++, Java, Python, Rust, or equivalent.
  • Hands-on experience building and shipping complex production systems across architecture, experimentation, debugging, and implementation.
  • Experience influencing senior engineers and leaders without direct authority.

Responsibilities

  • Define the technical vision and architecture for agentic reliability systems across Meta.
  • Design, code, and ship production agentic systems for complex infrastructure problems.
  • Lead development of an AI agent for production incident investigation and mitigation.
  • Improve agent reasoning, context, tool use, planning, evaluation, learning, and safe execution.
  • Explore techniques like fine-tuning, reinforcement learning, distillation, pruning, and inference optimization.
  • Identify new high-value applications of AI across incident prevention, detection, mitigation, observability, and infrastructure operations.
  • Prototype rapidly, validate against real workloads, and scale successful approaches.
  • Build evaluation/experimentation systems connecting agent quality to outcomes.

Skills

Technical leadership
Distributed systems
AI/ML in production
Incident response
Evaluation & experimentation

Education

Bachelor's degree in Computer Science or equivalent

Tools

C++
Java
Python
Rust

Job description

Meta seeks an experienced Software Engineer to build next-generation AI systems for production reliability. You will work at the intersection of large-scale distributed systems, incident response, and applied AI, leading the development of an autonomous AI agent that investigates incidents, identifies root causes, and executes safe mitigations under human supervision.

The role combines deep infrastructure expertise with practical AI deployment, requiring hands-on prototyping, evaluation, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Reliability Engineer for Production Autonomy
Principal AI Reliability Engineer for Production Autonomy

Meta • Columbus (OH)

On-site
USD 271,000 - 347,000
Lead AI Agentic Reliability Engineer
Lead AI Agentic Reliability Engineer

Meta • Santa Fe (NM)

On-site
USD 271,000 - 347,000
Principal AI Reliability Engineer
Principal AI Reliability Engineer

Meta • Madison (WI)

On-site
USD 271,000 - 347,000
Senior AI-Driven Reliability Engineer for Production
Senior AI-Driven Reliability Engineer for Production

Meta • United States

Remote
USD 180,000 - 240,000
Senior AI Reliability Engineer for Production Systems
Senior AI Reliability Engineer for Production Systems

Meta • Richmond (VA)

On-site
USD 271,000 - 347,000
Lead Software Engineer, AI-Driven Production Reliability
Lead Software Engineer, AI-Driven Production Reliability

Meta • Olympia (WA)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • United States

Remote
USD 180,000 - 240,000
Director, AI-Powered Incident Response
Director, AI-Powered Incident Response

Meta • Menlo Park (CA)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Santa Fe (NM)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Hartford (CT)

On-site
USD 271,000 - 347,000