Software Engineer - Infrastructure (Technical Leadership)

Meta Careers

United States

Remote

USD 180,000 - 280,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Meta Platforms, Inc. is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability.

You will lead the development of an AI agent that autonomously investigates production incidents, identifies root causes, and executes safe mitigation plans with human supervision. You will work at the intersection of distributed systems, incident response, and applied AI, prototyping, building, and shipping production-grade systems that learn from outcomes and

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering or equivalent.
  • 12+ years of software engineering experience in large-scale distributed or infrastructure systems.
  • Experience applying AI/ML to production problems and moving concepts to production deployments.

Responsibilities

  • Prototype, build, evaluate, and ship AI systems with production data.
  • Lead development of an AI agent to autonomously investigate incidents and mitigate risks.
  • Improve agent accuracy and autonomy, and push improvements across incident prevention and observability.

Skills

Distributed systems
AI/ML in production
Technical leadership
Telemetry analysis
Reliability engineering
Observability
System safety & rollback
Self-improving systems
Experimentation & debugging
Production coding
Programming languages

Education

Bachelor's degree in CS/CE or related field

Tools

C++
Java
Python
Rust

Job description

Meta is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability. This role sits at the intersection of large-scale distributed systems, incident response, and applied AI. You will lead the development of an AI agent that can autonomously investigate production incidents, identify likely root causes, create safe mitigation plans, and execute them while humans supervise and can intervene. You will improve the agent's accuracy, autonomy, and real-world impact, while identifying other opportunities to apply agentic systems across incident prevention, detection, mitigation, observability, and infrastructure operations. This may include improving the agent's architecture, context, and tooling, or adapting foundation models through fine-tuning, reinforcement learning, distillation, pruning, and inference optimization.

This is an applied engineering role, not a research position. The ideal candidate combines deep infrastructure expertise with sound judgment about when and how to apply AI to production problems, a long-term technical vision, and the ability to move quickly from an ambitious idea to a production system operating safely at Meta scale.

This is also a deeply hands-on senior IC role. You will personally prototype, build, evaluate, and ship AI systems, working directly in the code and with production data. A central challenge is creating systems that improve from experience: learning from outcomes, identifying their own failure modes, testing changes, and safely increasing their effectiveness over time.

Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 12+ years of software engineering experience, including experience building and operating large-scale distributed or infrastructure systems
  • Experience setting technical direction and leading complex, multi-year engineering efforts across organizational boundaries
  • Experience applying AI or machine learning systems to production problems
  • Demonstrated experience moving from technical concept to production deployment and measurable impact
  • Experience diagnosing complex production systems using telemetry, code, configuration, and dependency information
  • Experience coding in languages such as C++, Java, Python, Rust, or equivalent
  • Recent hands-on experience building and shipping complex production systems, with the ability to move directly between architecture, experimentation, debugging, and implementation
  • Experience influencing senior engineers and leaders without direct organizational authority
  • Experience with autonomous or semi-autonomous production actions and their safety, authorization, and rollback mechanisms
  • Experience building systems that operate at hyperscale under strict reliability and latency requirements
  • Experience improving agent quality through evaluation, context engineering, fine-tuning, reinforcement learning, distillation, or model-serving optimization
  • Track record of identifying unconventional opportunities, rapidly prototyping solutions, and changing the technical direction of a large organization
  • Deep expertise in observability, incident response, change safety, distributed systems, or production infrastructure
  • Experience building self-improving or self-evolving systems, including automated experimentation, feedback loops, hill climbing, reinforcement learning, or recursive self-improvement
  • Experience building production AI agents that reason across telemetry, code, configuration, deployments, and operational knowledge
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • United States

Remote
USD 180,000 - 240,000
Software Engineer - Infrastructure (Technical Leadership)
Software Engineer - Infrastructure (Technical Leadership)

Meta • United States

Remote
USD 250,000 - 350,000
Software Engineer, Infrastructure (Technical Leadership)
Software Engineer, Infrastructure (Technical Leadership)

Meta Careers • Bellevue (WA), Menlo Park (CA)

On-site
USD 240,000 - 320,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Columbia (SC)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Indianapolis (IN)

On-site
USD 271,000 - 347,000
Principal Software Engineer, Systems
Principal Software Engineer, Systems

Meta • Baton Rouge (LA)

On-site
USD 271,000 - 347,000
Equity
Bonus
Lead AI Systems Engineer for Production Reliability
Lead AI Systems Engineer for Production Reliability

Meta • Indianapolis (IN)

On-site
USD 271,000 - 347,000
Senior Software Engineer - Agentic Infra & Production AI
Senior Software Engineer - Agentic Infra & Production AI

Meta • United States

Remote
USD 250,000 - 350,000
Senior AI-Driven Reliability Engineer for Production
Senior AI-Driven Reliability Engineer for Production

Meta • United States

Remote
USD 180,000 - 240,000
Senior AI Reliability Engineer for Production Systems
Senior AI Reliability Engineer for Production Systems

Meta • Baton Rouge (LA)

On-site
USD 271,000 - 347,000
Equity
Bonus