Software Engineering Director, AI for Incident Response

Meta

Menlo Park (CA)

On-site

USD 271,000 - 347,000

Full time

21 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Meta seeks a Software Engineering Director to lead the organization building the future of incident response. This leader will drive Meta's AI-powered incident response agent in a human-on-the-loop model, balancing autonomous actions with required human judgment.

You will define the technical, product, and organizational strategy, lead an engineering organization with real production systems, and collaborate with infrastructure and product groups to deliver reliable, scalable, and safety-first

Qualifications

  • 10+ years of software engineering experience in leadership roles.
  • 7+ years in engineering management directing multiple teams.
  • Experience with large-scale distributed systems and production AI.
  • Proven ability to build and operate reliability-focused platforms.
  • Experience hiring and mentoring engineers and managers.

Responsibilities

  • Define and drive the engineering strategy and roadmap for AI-powered incident response.
  • Lead architecture and delivery of an agent that reasons across telemetry, code, configurations, deployments, dependencies, and incident history to identify root causes and take action.
  • Establish a human-on-the-loop operating model with boundaries, validation, auditability, rollback, and escalation paths.
  • Build Opsmate as an extensible platform with shareable capabilities and domain-specific skills.
  • Establish evaluation systems linking agent quality to reliability outcomes (investigation accuracy, adoption, time to mitigation).
  • Partner with product groups, infrastructure teams, product management, and data science to define success and shared outcomes.
  • Ensure platform reliability, latency, scale, security, privacy, and cost in production environments.
  • Grow an engineering organization by attracting, developing, and retaining engineers and managers across distributed systems and applied AI.
  • Create a culture of technical depth, rapid learning, operational excellence, and accountability for customer impact.
  • Contribute to company-wide engineering and AI initiatives, including recruiting and standards.

Skills

Software engineering
Technical leadership
Engineering management
Team leadership
Engineering strategy
Distributed systems
AI systems
Production systems
Platform building
Automated systems
Hiring
Cross-functional collaboration

Job description

Meta is seeking a Software Engineering Director to lead the organization building the future of incident response. As agents increasingly build and operate infrastructure, Meta needs reliability systems that can reason and act at machine speed. This leader will drive Meta's AI-powered incident response agent, Meta’s AI agent for incident response, toward a human-on-the-loop model: the agent autonomously gathers context, investigates incidents, identifies root causes, creates a safe mitigation plan, and executes the appropriate action, while humans set boundaries, supervise consequential actions, and intervene when judgment is required.This role combines frontier agent capabilities with one of the world’s largest and most complex production environments. You will define the technical, product, and organizational strategy, lead an engineering organization with real production systems, and partner across infrastructure and product groups to make incident response faster, safer, and dramatically less labor-intensive. Few roles offer the opportunity to define both the technology and operating model for autonomous incident mitigation at this scale.

Software Engineering Director, AI for Incident Response Responsibilities:
  • Define and drive the engineering strategy and roadmap for AI-powered incident response, progressing from investigation and diagnosis to safe, end-to-end mitigation and continuous learning
  • Lead the architecture and delivery of an agent that reasons across telemetry, code, configurations, deployments, service dependencies, and incident history to identify root causes and take appropriate action
  • Establish a human-on-the-loop operating model with progressive autonomy, explicit policy and permission boundaries, independent validation, auditability, rollback, and clear escalation paths
  • Build Opsmate as an extensible platform that combines shared, out-of-the-box capabilities with domain-specific skills and agents contributed by product groups and individual on-call engineers
  • Establish trusted evaluation and measurement systems that connect agent quality to customer and reliability outcomes, including investigation accuracy, adoption, time to mitigation, and reduced operational effort
  • Partner deeply with product groups, infrastructure teams, product management, and data science to understand what success means for their services and build toward shared outcomes
  • Ensure the platform itself meets a high bar for reliability, latency, scale, security, privacy, and cost in Meta’s production environment
  • Build and grow an engineering organization that consistently delivers measurable results by attracting, developing, and retaining engineers and engineering managers across distributed systems, reliability, and applied AI
  • Create a culture of technical depth, rapid learning, operational excellence, and accountability for realized customer impact
  • Contribute to company-wide engineering and AI initiatives, including recruiting, technical standards, and responsible AI practices
Minimum Qualifications:
  • 10+ years of software engineering experience, including experience in technical leadership roles
  • 7+ years of engineering management experience, including leading multiple teams or managers and delivering products or systems with measurable impact
  • Experience defining and executing engineering strategy for large-scale distributed systems, reliability platforms, developer infrastructure, or production AI systems
  • Experience building and operating systems that support critical, always-on production workloads
  • Track record of building broadly adopted platforms and driving outcomes across organizational boundaries
  • Experience evaluating complex automated systems and establishing safeguards for trustworthy production operation
  • Experience building engineering teams through hiring, onboarding, and developing engineers and managers at multiple levels
  • Experience partnering cross-functionally with product, data science, and infrastructure organizations
Preferred Qualifications:
  • Experience with incident management, SRE, observability, oncall operations, or automated remediation at significant scale
  • Experience designing safe execution systems, including identity, authorization, sandboxing, validation, rollback, and auditability
  • Experience building extensible platforms that allow domain teams to contribute specialized capabilities while preserving a coherent user experience and control plane
  • Track record of taking a technically ambitious product from early adoption to trusted, broad production use
  • Experience managing other engineering managers and building a deep technical and management bench across multiple levels
  • Experience operating LLM-based or tool-using agents in production, including evaluation, orchestration, and progressive autonomy
About Meta:

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.

Meta is proud to be an Equal Employment Opportunity and Aff…?

Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

$271,000/year to $347,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure (Technical Leadership)
Software Engineer, Infrastructure (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Bonus
Equity
Benefits
Security Engineer, Applied AI
Security Engineer, Applied AI

Meta • New York (NY)

On-site
USD 154,000 - 217,000
Software Engineer, Product (Technical Leadership)
Software Engineer, Product (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Software Engineering Manager
Software Engineering Manager

WhatsApp • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Software Engineer, Systems
Software Engineer, Systems

Meta • Menlo Park (CA)

On-site
USD 184,000 - 257,000
Competitive compensation
Software Engineer, Systems
Software Engineer, Systems

Meta • New York (NY)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Software Engineer (Technical Leadership) — Central Security
Software Engineer (Technical Leadership) — Central Security

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Software Engineer, Systems
Software Engineer, Systems

Meta • Washington

On-site
USD 184,000 - 257,000
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Software Developer, Scaled Ops AI Acceleration Team
Software Developer, Scaled Ops AI Acceleration Team

Meta • Austin (TX)

On-site
USD 171,000 - 240,000
Bonus
Equity
Benefits