Senior Cloud Infrastructure Engineer — Scale & Reliability

Harvey

San Francisco (CA)

On-site

USD 161,300 - 241,900

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Harvey is building the AI platform trusted by the world’s leading law firms and enterprises. We’re seeking a Production Engineer to design, operate, and scale Harvey’s global compute and networking infrastructure, Kubernetes platform, and production foundations.

You will improve reliability, capacity planning, automation, and observability, partnering with Product Engineering, Security, and Platform teams to support rapid growth while maintaining high availability and security.

Qualifications

  • 5+ years of software, infrastructure, site reliability, or production engineering experience.
  • Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
  • Strong hands‑on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
  • Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.
  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
  • Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.
  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
  • A track record of driving complex, cross‑functional technical initiatives and influencing engineering decisions without relying on formal authority.
  • Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders.
  • A systems‑thinking mindset and a passion for building simple, reliable, and scalable infrastructure platforms.

Responsibilities

  • Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads.
  • Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
  • Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost.
  • Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions.
  • Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely.
  • Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship.
  • Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.
  • Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads.
  • Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth.
  • Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation.
  • Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring.
  • Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.
  • Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.
  • Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform.
  • Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements.

Skills

5+ years experience
Kubernetes in production
Infrastructure automation
Terraform or Pulumi
Observability systems
Cloud infrastructure (AWS/Azure/GCP)
Security best practices
Strong communication
Multi-cloud experience
Cross-functional collaboration

Tools

Terraform
Pulumi
Kubernetes

Job description

Harvey is building the AI platform trusted by the world’s leading law firms and enterprises. We’re seeking a Production Engineer to design, operate, and scale Harvey’s global compute and networking infrastructure, Kubernetes platform, and production foundations.

You will improve reliability, capacity planning, automation, and observability, partnering with Product Engineering, Security, and Platform teams to support rapid growth while maintaining high availability and security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer – Cloud, Kubernetes & SRE
Senior Infrastructure Engineer – Cloud, Kubernetes & SRE

Harvey • United States

On-site
USD 161,000 - 242,000
Senior Production Engineer: Scale-Driven Infra & Kubernetes
Senior Production Engineer: Scale-Driven Infra & Kubernetes

Neura Market • San Francisco (CA)

On-site
USD 161,000 - 242,000
Senior Infrastructure Engineer - Scalable Cloud & Kubernetes
Senior Infrastructure Engineer - Scalable Cloud & Kubernetes

Harvey • New York (NY)

On-site
USD 161,000 - 242,000
Senior Production Infrastructure Engineer
Senior Production Infrastructure Engineer

Harvey • United States

On-site
USD 231,000 - 340,000
Engineering Manager, Production Infrastructure
Engineering Manager, Production Infrastructure

Harvey • United States

On-site
USD 260,000 - 340,000
Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Harvey • United States

On-site
USD 161,000 - 242,000
Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Harvey • New York (NY)

On-site
USD 161,000 - 242,000
Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Neura Market • San Francisco (CA)

On-site
USD 161,000 - 242,000
Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Harvey • San Francisco (CA)

On-site
USD 161,300 - 241,900
Senior Core Infra Engineer — Multi-Cloud & AI
Senior Core Infra Engineer — Multi-Cloud & AI

Harvey • San Francisco (CA)

On-site
USD 200,000 - 250,000