Senior AI Agent Reliability Engineer

Meta

Austin (TX)

On-site

USD 271,000 - 347,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Meta in Austin seeks an experienced Software Engineer to build AI systems for production reliability. You will lead the development of an autonomous AI agent to investigate production incidents, identify root causes, craft safe mitigation plans, and execute them under human supervision.

This is a deeply hands-on senior IC role with a long-term technical vision for reliability at Meta. The role requires deep infrastructure expertise, experience applying AI to production problems, and the ability

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • 12+ years of software engineering experience in large-scale distributed or infrastructure systems.
  • Experience setting technical direction and leading multi-year engineering efforts.
  • Experience applying AI or ML systems to production problems.
  • Experience moving from concept to production deployment with measurable impact.
  • Experience diagnosing complex production systems using telemetry, code, and configuration.
  • Experience coding in C++, Java, Python, Rust, or equivalent.
  • Hands-on experience building and shipping complex production systems.
  • Influencing senior engineers and leaders without direct authority.

Responsibilities

  • Define the technical vision and architecture for agentic reliability systems across Meta.
  • Design, code, and ship production agentic systems for complex infrastructure problems.
  • Lead development of an AI agent for production incident investigation and mitigation.
  • Improve agent reasoning, context, tool use, planning, evaluation, learning, and safe execution.
  • Apply techniques like reinforcement learning, fine-tuning, and model optimization; choose simple effective approaches.
  • Identify new high-value AI applications across incident prevention, detection, and observability.
  • Prototype rapidly, validate with production workloads, and scale successful approaches.
  • Build evaluation and experimentation systems linking agent quality to outcomes.
  • Create closed-loop improvement systems from production data to agent behavior.
  • Establish architectures and guardrails for production actions with oversight.
  • Collaborate with Infrastructure, AI, Product, and Reliability leaders to integrate capabilities.
  • Mentor engineers in distributed systems and applied AI.

Skills

Large-scale distributed systems
AI systems in production
Technical leadership
Programming languages

Education

Bachelor's degree in CS/CE or equivalent

Tools

C++
Java
Python
Rust

Job description

Meta in Austin seeks an experienced Software Engineer to build AI systems for production reliability. You will lead the development of an autonomous AI agent to investigate production incidents, identify root causes, craft safe mitigation plans, and execute them under human supervision.

This is a deeply hands-on senior IC role with a long-term technical vision for reliability at Meta. The role requires deep infrastructure expertise, experience applying AI to production problems, and the ability

Get your free, confidential resume review.

or drag and drop your file here.