Engineering Manager – Production Infrastructure & Reliability

Harvey, Inc.

San Francisco (CA)

On-site

USD 260,000 - 340,000

Full time

28 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Harvey is seeking an Engineering Manager to lead the Infrastructure Production Engineering organization. You will own the reliability, scalability, security, and efficiency of Harvey's infrastructure platform, guiding a team responsible for compute fleet management, capacity planning, automation, and production operations.

You'll partner with Product Engineering, Security, AI Infrastructure, and Platform teams to scale quickly with growth, reporting to the Head of Infrastructure, and helping

Qualifications

  • 7+ years of software or infrastructure engineering experience, including 5+ years leading engineering teams.
  • Deep expertise operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
  • Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
  • Experience building and operating large-scale distributed systems with strong reliability, scalability, and performance characteristics.
  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
  • Experience designing and operating observability platforms, including monitoring, logging, alerting, and incident response processes.
  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
  • Demonstrated success leading complex cross-functional technical initiatives and influencing engineering strategy across organizations.
  • Excellent communication skills with the ability to communicate technical concepts clearly to both engineering and executive audiences.
  • A systems-thinking mindset and passion for building simple, reliable, and scalable infrastructure platforms.

Responsibilities

  • Lead, mentor, and grow a team of high-performing infrastructure engineers responsible for Harvey's production infrastructure foundation.
  • Foster a culture of operational excellence, engineering quality, customer ownership, and continuous improvement.
  • Partner with Engineering, Security, Product, and AI Infrastructure leaders to define long-term infrastructure strategy and execution priorities.
  • Drive technical direction for compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
  • Lead cross-functional initiatives to improve reliability, scalability, security, operational efficiency, and infrastructure cost optimization.
  • Own and operate Harvey's global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.
  • Manage compute resources to maximize utilization, performance, and service availability while supporting rapidly growing AI workloads.
  • Lead capacity planning, demand forecasting, and fleet lifecycle management to ensure infrastructure scales efficiently with business growth.
  • Operate and continuously improve Harvey's Kubernetes platform, including cluster provisioning, upgrades, monitoring, reliability, performance, and operational automation.
  • Own Harvey's Temporal-based workflow orchestration platform, ensuring reliable, scalable, and observable execution of distributed application workflows.
  • Drive infrastructure cost optimization through capacity management, resource rightsizing, workload efficiency improvements, and utilization monitoring.
  • Build and maintain secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.
  • Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.
  • Establish comprehensive observability, monitoring, alerting, incident response, and operational readiness practices across the infrastructure platform.

Skills

Cloud architecture
Leadership
Production troubleshooting
Cross-functional collaboration
Security best practices

Tools

Terraform
Pulumi
Kubernetes
AWS/Azure/GCP

Job description

Harvey is seeking an Engineering Manager to lead the Infrastructure Production Engineering organization. You will own the reliability, scalability, security, and efficiency of Harvey's infrastructure platform, guiding a team responsible for compute fleet management, capacity planning, automation, and production operations.

You'll partner with Product Engineering, Security, AI Infrastructure, and Platform teams to scale quickly with growth, reporting to the Head of Infrastructure, and helping

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Infrastructure Engineer — Scale & Reliability
Senior Cloud Infrastructure Engineer — Scale & Reliability

Harvey • San Francisco (CA)

On-site
USD 161,300 - 241,900
Engineering Manager, Production Engineering
Engineering Manager, Production Engineering

Harvey, Inc. • San Francisco (CA)

On-site
USD 260,000 - 340,000
Senior Technical Program Manager – Platform & Infrastructure
Senior Technical Program Manager – Platform & Infrastructure

Harvey, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 188,000 - 278,000
Engineering Manager - AI Model Infrastructure & Platform
Engineering Manager - AI Model Infrastructure & Platform

Harvey • San Francisco (CA)

On-site
USD 260,000 - 340,000
Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Harvey • San Francisco (CA)

On-site
USD 161,300 - 241,900
Technical Program Manager, Quality & Reliability
Technical Program Manager, Quality & Reliability

Harvey • United States

Remote
USD 140,000 - 180,000
Head of Enterprise App Engineering & AI Enablement
Head of Enterprise App Engineering & AI Enablement

Harvey • United States

Remote
USD 120,000 - 180,000
Staff Site Reliability Engineer: Global Infra & Automation
Staff Site Reliability Engineer: Global Infra & Automation

Harvey, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 238,000 - 290,000
Relocation assistance
In-person work model
Technical Program Manager, Platform & Infrastructure
Technical Program Manager, Platform & Infrastructure

Harvey, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 188,000 - 278,000
Senior Integrations & Product Operations Lead
Senior Integrations & Product Operations Lead

Harvey, Inc. • New York (NY)

On-site
USD 155,000 - 233,000