Director of Engineering, SRE (AI Security / Startup)

Jobot

Chicago (IL)

Remote

USD 220,000 - 260,000

Full time

12 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Jobot is hiring a Director of Engineering focused on Site Reliability Engineering for a fast-growing AI security startup. This fully remote role leads a highly technical SRE organization, shaping reliability, scalability, and production excellence across cloud infrastructure.

You will mentor senior engineers, define ownership and discipline across AWS, Kubernetes, IaC, and observability, and drive incident response improvements while balancing short- and long-term platform investments.

Qualifications

  • Experience leading a technical SRE/infra organization.
  • Proven ability to balance reliability with fast-moving startup needs.
  • Strong cloud and distributed systems expertise with production-grade platforms.
  • Able to mentor senior engineers and shape engineering culture.

Responsibilities

  • Lead a highly technical SRE organization focused on reliability and operational excellence.
  • Mentor and grow senior-level engineers; set ownership and accountability.
  • Define and uphold SLIs/SLOs, incident response, postmortems, and capacity planning.
  • Influence architecture decisions across AWS, Kubernetes, and IaC tooling.

Skills

SRE
Cloud infrastructure
Distributed systems
Platform engineering
Automation
Observability
Production operations
Leadership
Mentorship

Tools

AWS
Kubernetes
Docker
Terraform
CI/CD
Linux

Job description

Want to learn more about this role and Jobot? Click our Jobot logo and follow our LinkedIn page!

Job details:

Incredible opportunity to join one of the fastest growing AI Security startups in the world // VC Backed Startup experience is required // Fully Remote

This Jobot Job is hosted by: Craig Rosecrans

Salary: $220,000 - $260,000 per year

A bit about us:

We are partnering with a rapidly growing, venture-backed technology company operating at the intersection of Artificial Intelligence, Cybersecurity, and Cloud Infrastructure to hire a Director of Engineering focused on Site Reliability Engineering.

The company's technology protects mission-critical AI systems and is deployed across sophisticated enterprise and government environments. As adoption continues to accelerate, the organization is looking for an exceptional engineering leader to own and evolve the reliability, scalability, resiliency, and operational excellence of its platform.

This is not a traditional people-management-only Director position.

We are looking for someone who combines exceptional leadership ability with deep, current technical expertise across SRE, cloud infrastructure, distributed systems, platform engineering, automation, observability, and production operations.

You will inherit a highly technical engineering team, and credibility matters.

The engineers reporting to this person need to trust that their leader understands the technology at their level, can challenge their thinking, can make difficult architectural decisions, and—when necessary—can sit beside them during a complex production incident and help solve the problem.

You don't need to write production code every day.

But you absolutely need to be capable of doing it.

Why join us?

This is an opportunity to join a rapidly scaling company tackling one of the most important emerging challenges in technology: protecting the AI systems enterprises and government organizations increasingly depend upon.

You will have significant ownership over the infrastructure and reliability strategy supporting a sophisticated cybersecurity platform while leading an experienced technical team.

For the right engineering leader, this represents a rare combination of:

  • AI + Cybersecurity + Cloud Infrastructure + Distributed Systems + Engineering Leadership + Mission-Critical Reliability.

Compensation: $230,000–$260,000 Base Salary + 10% Annual Bonus + Stock Options

Location: Fully Remote – United States

Job Details
What You'll Own

Lead a Highly Technical SRE Organization

Lead, mentor, develop, and grow a team responsible for the reliability and operational excellence of a sophisticated AI security platform.

You will:

  • Develop and mentor senior-level SRE, Infrastructure, and Platform engineers.
  • Establish clear expectations around ownership, execution, technical quality, and accountability.
  • Build an engineering culture centered around reliability, automation, continuous improvement, and operational excellence.
  • Recruit and retain exceptional engineering talent as the organization scales.
  • Provide meaningful technical mentorship rather than simply managing projects and people.
  • Make thoughtful, decisive engineering decisions in a fast-moving startup environment where perfect information isn't always available.
  • Balance short-term operational requirements with long‑term platform and infrastructure investments.

This organization values leaders who can move quickly, make difficult decisions, and create clarity in ambiguous environments.

Remain Deeply Technical

This Director will remain close to the technology and serve as a senior technical authority across SRE and infrastructure.

You should be capable of contributing meaningfully to conversations involving:

  • Cloud architecture
  • AWS
  • Kubernetes and container orchestration
  • Linux
  • Networking
  • Distributed systems
  • Infrastructure as Code
  • CI/CD
  • Observability
  • Production troubleshooting
  • Reliability engineering
  • Infrastructure automation
  • Security
  • Performance
  • Scalability
  • Resiliency and disaster recovery

You will partner with senior engineers on architectural decisions, identify systemic risks, challenge assumptions, and help troubleshoot particularly complex production problems.

Your engineers should consider you one of the strongest technical resources in the organization—not simply the person managing the strongest technical resources.

Site Reliability & Production Engineering

Own and continuously improve how production systems are designed, deployed, monitored, and operated.

Responsibilities will include:

  • Establishing and evolving SLIs, SLOs, availability objectives, and error budgets.
  • Improving system availability, fault tolerance, scalability, and performance.
  • Building mature observability practices across metrics, logs, traces, dashboards, and alerting.
  • Leading improvements to incident response, escalation, root‑cause analysis, and postmortems.
  • Reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
  • Improving capacity planning and infrastructure forecasting.
  • Identifying systemic reliability risks before they become customer‑impacting incidents.
  • Developing production‑readiness standards across engineering.
  • Reducing operational toil through automation.

The goal is not simply to respond effectively when systems fail. It is to engineer systems so failures become less frequent, less severe, easier to detect, and faster to recover from.

Cloud, Platform & Infrastructure Engineering

Help evolve the infrastructure supporting a rapidly scaling, security‑focused technology platform.

Relevant experience may include:

  • AWS and cloud‑native infrastructure
  • Kubernetes
  • Docker / containerized workloads
  • Infrastructure as Code
  • Terraform or comparable technologies
  • CI/CD and software delivery automation
  • Linux systems engineering
  • Networking and distributed systems
  • Cloud security and IAM
  • Secrets and key management
  • Configuration management
  • Observability and monitoring platforms
  • Automated provisioning
  • Performance optimization
  • Capacity management
  • High‑availability architecture
  • Experience operating complex distributed SaaS platforms at scale is strongly preferred.
Air‑Gapped & Disconnected Environments

Experience supporting air‑gapped, disconnected, restricted, or highly regulated environments will be particularly valuable.

Some customers operate environments where traditional cloud assumptions simply don't apply.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of Engineering, SRE (AI Security / Startup)
Director of Engineering, SRE (AI Security / Startup)

Jobot • New York (NY)

Remote
USD 220,000 - 260,000
Stock options
10% annual bonus
Director of Engineering, SRE (AI Security / Startup)
Director of Engineering, SRE (AI Security / Startup)

Jobot • Philadelphia

On-site
USD 220,000 - 260,000
Director of Engineering (AI Security)
Director of Engineering (AI Security)

Leoforce • New York (NY)

On-site
USD 250,000 - 350,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Onebrief • Colorado Springs (CO)

On-site
USD 205,000 - 255,000
Relocation assistance
On-site customer deployments
Director of SRE Engineering for AI Security & Cloud
Director of SRE Engineering for AI Security & Cloud

Jobot • Chicago (IL)

Remote
USD 220,000 - 260,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Onebrief • Arlington (VA)

On-site
USD 190,000 - 240,000
Relocation assistance
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Principal Site Reliability Engineer
Principal Site Reliability Engineer

HeyGen • Tempe (AZ)

On-site
USD 140,000 - 190,000
Director of Infrastructure & Reliability
Director of Infrastructure & Reliability

PracticeSuite, Inc. • Tampa (FL)

On-site
USD 160,000 - 230,000
Senior Engineering Manager - Cloud Platform & SRE | Mission-Critical AI Software Platform
Senior Engineering Manager - Cloud Platform & SRE | Mission-Critical AI Software Platform

Techfellow Limited • Boston (MA)

On-site
USD 220,000 - 260,000
Relocation funded