Director, AI Platform Reliability

logicmonitor

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

LogicMonitor, based in San Francisco, seeks an accomplished Director of AI Platform Reliability to lead architecture, development, and operation of scalable software platforms. You will manage multiple teams and drive engineering standards across data pipelines, microservices, and cloud‑native services.

The ideal candidate brings deep Java expertise, distributed systems experience, and a track record of building mission‑critical platforms at scale.

Qualifications

  • 10+ years of professional software‑engineering experience, including building large‑scale distributed systems.
  • Experience leading engineering teams, architects, and staff engineers.
  • Proven success delivering and operating platforms that process hundreds of millions of transactions, requests, or events.

Responsibilities

  • Lead multiple engineering teams responsible for high‑volume distributed systems and data platforms.
  • Define technical strategy and architecture for platforms processing hundreds of millions of transactions and terabytes of data.
  • Guide development of Java‑based microservices, APIs, Kafka pipelines, and cloud‑native services.
  • Build scalable data lake/data lakehouse platforms supporting real‑time, near‑real‑time, and batch analytics.

Skills

Java
Distributed systems
Concurrency
Multithreading
Team leadership

Tools

Kafka

Job description

About Us:

We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better everyday. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers.

This role is open to candidates based in or near San Francisco, CA. At LogicMonitor, we hire within our Centers of Energy-vibrant locations where our teams connect, collaborate, and innovate.

To learn more about life at LogicMonitor, check out our Careers Page .

What You'll Do:

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.

Our customers love LogicMonitor’s ability to bring cloud and traditional IT together into one view, as seen in minimal churn rates, expansion business, and exciting new customer references. In fact, LogicMonitor has received the highest Net Promoter Score of any IT Infrastructure Management provider. LogicMonitor also boasts high employee satisfaction. We have been certified as a Great Place To Work®, and named one of BuiltIn’s Best Places to Work for the seventh year in a row!

We are looking for an accomplished and hands‑on Director of AI Platform Reliability to lead the architecture, development, and operation of highly scalable, distributed software platforms.

This leader will be responsible for systems that process hundreds of millions/billions of transactions and events , manage terabytes to petabytes of data , and deliver reliable, low‑latency services to enterprise customers. The ideal candidate combines strong engineering depth in Java, Kafka, distributed systems, and cloud‑native microservices with a demonstrated ability to build and lead high‑performing engineering organizations.

This is a strategic leadership role, but it requires a leader who can remain close to the technology, participate in architecture reviews, challenge design decisions, guide teams through complex production problems, and establish the engineering practices required to operate mission‑critical platforms at scale.

Here’s a closer look at this key role:

  • Lead and scale multiple engineering teams responsible for high-volume, business‑critical distributed systems and data platforms.
  • Define the technical strategy and architecture for platforms processing hundreds of millions of transactions and terabytes of data.
  • Guide the development of Java‑based microservices, APIs, Kafka streaming pipelines, batch‑processing workflows, and cloud‑native services.
  • Build and evolve scalable data lake and Data Lakehouse platforms supporting real‑time, near‑real‑time, and batch analytics workloads.
  • Establish reliable data ingestion, transformation, storage, governance, lineage, retention, and data‑quality practices across streaming and batch pipelines.
  • Build low‑latency, highly available, fault‑tolerant systems with strong scalability, resiliency, and disaster‑recovery capabilities.
  • Define and own operational SLAs, SLOs, availability targets, recovery objectives, and performance metrics for critical services and data pipelines.
  • Drive capacity planning, load testing, throughput optimization, and improvements to p95 and p99 latency.
  • Ensure effective Kafka design, including partitioning, consumer groups, ordering, schema evolution, replay, and lag management.
  • Establish engineering standards for architecture, coding, testing, security, observability, and production readiness.
  • Partner with Product, Architecture, SRE, Security, Data, and Infrastructure teams to deliver strategic platform initiatives.
  • Strengthen operational excellence through monitoring, incident management, on‑call practices, root‑cause analysis, and continuous reliability improvements.
  • Recruit, mentor, and develop engineering managers, architects, and senior technical leaders.
  • Improve developer productivity, CI/CD automation, deployment safety, and release predictability.
  • Manage technical debt, platform modernization, cloud costs, and long‑term scalability investments.
What You’ll Need:

10+ years of professional software‑engineering experience, including significant experience building large‑scale distributed systems.

Experience leading engineering teams, architects, and staff engineers.

Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events.

Deep technical expertise in Java, JVM performance, concurrency, multithread

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, AI Platform Reliability - LogicMonitor
Director, AI Platform Reliability - LogicMonitor

OpenTalent • San Francisco (CA)

On-site
USD 180,000 - 260,000
Head of AI Platform Reliability & Scale
Head of AI Platform Reliability & Scale

logicmonitor • San Francisco (CA)

On-site
USD 180,000 - 260,000
Director, AI Platform Reliability & Scale
Director, AI Platform Reliability & Scale

OpenTalent • San Francisco (CA)

On-site
USD 180,000 - 260,000
Head of Platform Engineering
Head of Platform Engineering

Soni • Hazlet Township (NJ)

On-site
USD 216,000 - 264,000
Senior Staff Engineer
Senior Staff Engineer

Harnham • California (MO)

Hybrid
USD 180,000 - 210,000
Software Development Engineer
Software Development Engineer

Encore Talent Solutions • San Jose (CA)

Hybrid
USD 140,000 - 180,000
Platform Engineering Architect (DevOps Architect, Webfarm, Kafka, Security: Authentication, Inf[...]
Platform Engineering Architect (DevOps Architect, Webfarm, Kafka, Security: Authentication, Inf[...]

Infojini Inc • Alpharetta (GA)

On-site
USD 100,000 - 140,000
(Confluent) Senior Principal Engineer, Kafka
(Confluent) Senior Principal Engineer, Kafka

IBM • Tucson (AZ)

On-site
USD 210,000 - 290,000
(Confluent) Senior Principal Engineer, Kafka
(Confluent) Senior Principal Engineer, Kafka

IBM • Lowell (MA)

On-site
USD 180,000 - 240,000
(Confluent) Senior Principal Engineer, Kafka
(Confluent) Senior Principal Engineer, Kafka

IBM • Poughkeepsie (AR)

On-site
USD 180,000 - 280,000