Director of Engineering, SRE (AI Security / Startup)

Jobot

New York (NY)

Remote

USD 220,000 - 260,000

Full time

35 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options
10% annual bonus

Job summary

Jobot is seeking a Director of Engineering specializing in Site Reliability Engineering to lead a highly technical team focused on reliability, scalability, and operational excellence. This remote US role demands deep expertise in cloud infrastructure, distributed systems, automation, observability, and production operations, with a track record of engineering leadership beyond people management.

You will mentor senior engineers, set SLIs/SLOs, drive incident postmortems, and balance long-term

Qualifications

  • Extensive leadership experience in Site Reliability Engineering.
  • Technical expertise across cloud infrastructure and distributed systems.
  • Ability to mentor senior engineers and drive reliability initiatives.

Responsibilities

  • Lead and grow a high-caliber SRE organization.
  • Define SLIs/SLOs and incident response improvements.
  • Balance short-term needs with long-term platform investments.
  • Recruit and retain top engineering talent.

Skills

SRE Leadership
Cloud Infrastructure
Distributed Systems
Observability
Production Operations
Automation
Architecture Decisions

Tools

AWS
Kubernetes
Terraform
CI/CD
Linux

Job description

Want to learn more about this role and Jobot? Click our Jobot logo and follow our LinkedIn page!

Job details:

Incredible opportunity to join one of the fastest growing AI Security startups in the world // VC Backed Startup experience is required // Fully Remote

This Jobot Job is hosted by: Craig Rosecrans

Salary: $220,000 - $260,000 per year

A bit about us:

We are partnering with a rapidly growing, venture-backed technology company operating at the intersection of Artificial Intelligence, Cybersecurity, and Cloud Infrastructure to hire a Director of Engineering focused on Site Reliability Engineering.

The company's technology protects mission-critical AI systems and is deployed across sophisticated enterprise and government environments. As adoption continues to accelerate, the organization is looking for an exceptional engineering leader to own and evolve the reliability, scalability, resiliency, and operational excellence of its platform.

This is not a traditional people-management-only Director position.

We are looking for someone who combines exceptional leadership ability with deep, current technical expertise across SRE, cloud infrastructure, distributed systems, platform engineering, automation, observability, and production operations.

You will inherit a highly technical engineering team, and credibility matters.

The engineers reporting to this person need to trust that their leader understands the technology at their level, can challenge their thinking, can make difficult architectural decisions, and—when necessary—can sit beside them during a complex production incident and help solve the problem.

You don't need to write production code every day.

But you absolutely need to be capable of doing it.

Why join us?

This is an opportunity to join a rapidly scaling company tackling one of the most important emerging challenges in technology: protecting the AI systems enterprises and government organizations increasingly depend upon.

You will have significant ownership over the infrastructure and reliability strategy supporting a sophisticated cybersecurity platform while leading an experienced technical team.

For the right engineering leader, this represents a rare combination of:

  • AI + Cybersecurity + Cloud Infrastructure + Distributed Systems + Engineering Leadership + Mission-Critical Reliability.

Compensation: $230,000–$260,000 Base Salary + 10% Annual Bonus + Stock Options

Location: Fully Remote – United States

Job Details
What You'll Own

Lead a Highly Technical SRE Organization

Lead, mentor, develop, and grow a team responsible for the reliability and operational excellence of a sophisticated AI security platform.

You will:

  • Develop and mentor senior-level SRE, Infrastructure, and Platform engineers.
  • Establish clear expectations around ownership, execution, technical quality, and accountability.
  • Build an engineering culture centered around reliability, automation, continuous improvement, and operational excellence.
  • Recruit and retain exceptional engineering talent as the organization scales.
  • Provide meaningful technical mentorship rather than simply managing projects and people.
  • Make thoughtful, decisive engineering decisions in a fast-moving startup environment where perfect information isn't always available.
  • Balance short-term operational requirements with long‑term platform and infrastructure investments.

This organization values leaders who can move quickly, make difficult decisions, and create clarity in ambiguous environments.

Remain Deeply Technical

This Director will remain close to the technology and serve as a senior technical authority across SRE and infrastructure.

You should be capable of contributing meaningfully to conversations involving:

  • Cloud architecture
  • AWS
  • Kubernetes and container orchestration
  • Linux
  • Networking
  • Distributed systems
  • Infrastructure as Code
  • CI/CD
  • Observability
  • Production troubleshooting
  • Reliability engineering
  • Infrastructure automation
  • Security
  • Performance
  • Scalability
  • Resiliency and disaster recovery

You will partner with senior engineers on architectural decisions, identify systemic risks, challenge assumptions, and help troubleshoot particularly complex production problems.

Your engineers should consider you one of the strongest technical resources in the organization—not simply the person managing the strongest technical resources.

Site Reliability & Production Engineering

Own and continuously improve how production systems are designed, deployed, monitored, and operated.

Responsibilities will include:

  • Establishing and evolving SLIs, SLOs, availability objectives, and error budgets.
  • Improving system availability, fault tolerance, scalability, and performance.
  • Building mature observability practices across metrics, logs, traces, dashboards, and alerting.
  • Leading improvements to incident response, escalation, root‑cause analysis, and postmortems.
  • Reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
  • Improving capacity planning and infrastructure forecasting.
  • Identifying systemic reliability risks before they become customer‑impacting incidents.
  • Developing production‑readiness standards across engineering.
  • Reducing operational toil through automation.

The goal is not simply to respond effectively when systems fail. It is to engineer systems so failures become less frequent, less severe, easier to detect, and faster to recover from.

Cloud, Platform & Infrastructure Engineering

Help evolve the infrastructure supporting a rapidly scaling, security‑focused technology platform.

Relevant experience may include:

  • AWS and cloud‑native infrastructure
  • Kubernetes
  • Docker / containerized workloads
  • Infrastructure as Code
  • Terraform or comparable technologies
  • CI/CD and software delivery automation
  • Linux systems engineering
  • Networking and distributed systems
  • Cloud security and IAM
  • Secrets and key management
  • Configuration management
  • Observability and monitoring platforms
  • Automated provisioning
  • Performance optimization
  • Capacity management
  • High‑availability architecture
  • Experience operating complex distributed SaaS platforms at scale is strongly preferred.
Air‑Gapped & Disconnected Environments

Experience supporting air‑gapped, disconnected, restricted, or highly regulated environments will be particularly valuable.

Some customers operate environments where traditional cloud assumptions simply don't apply.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of Engineering, SRE (AI Security / Startup)
Director of Engineering, SRE (AI Security / Startup)

Jobot • Chicago (IL)

Remote
USD 220,000 - 260,000
Director of Engineering, SRE (AI Security / Startup)
Director of Engineering, SRE (AI Security / Startup)

Jobot • Philadelphia

On-site
USD 220,000 - 260,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Leadout Capital • San Francisco (CA)

On-site
USD 180,000 - 240,000
On-site campus amenities
Director of Engineering (AI Security)
Director of Engineering (AI Security)

Leoforce • New York (NY)

On-site
USD 250,000 - 350,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Onebrief • Colorado Springs (CO)

On-site
USD 205,000 - 255,000
Relocation assistance
On-site customer deployments
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Onebrief • Arlington (VA)

On-site
USD 190,000 - 240,000
Relocation assistance
Director of SRE Engineering for AI Security & Cloud
Director of SRE Engineering for AI Security & Cloud

Jobot • Chicago (IL)

Remote
USD 220,000 - 260,000
Director of Infrastructure & Reliability
Director of Infrastructure & Reliability

PracticeSuite, Inc. • Tampa (FL)

On-site
USD 160,000 - 230,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Engineering Manager - Cloud Platform & SRE | Mission-Critical AI Software Platform
Senior Engineering Manager - Cloud Platform & SRE | Mission-Critical AI Software Platform

Techfellow Limited • Boston (MA)

On-site
USD 220,000 - 260,000
Relocation funded