Senior Software Engineer, Production Engineering

Neura Market

San Francisco (CA)

On-site

USD 161,300 - 241,900

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Harvey is building the AI platform trusted by the world’s leading law firms and enterprises. Our infrastructure powers every customer interaction, model inference, and production workload.

We’re seeking a Production Engineer to design, operate, and scale Harvey’s compute, networking, Kubernetes, and workflow platforms. You’ll improve reliability, security, and efficiency while partnering with cross-functional teams to translate product needs into resilient infrastructure.

Qualifications

  • 5+ years in software, infrastructure, SRE, or production engineering.
  • Extensive experience with cloud providers (AWS/Azure/GCP).
  • Hands-on Kubernetes in production and cluster lifecycle.
  • Distributed systems with high reliability and performance.
  • Infra automation with Terraform or Pulumi.
  • Strong compute, networking, capacity planning, and ops knowledge.
  • Security-focused infrastructure, IAM, and secrets management.
  • Excellent communicator and able to influence without authority.
  • Systems-thinking mindset to build simple, scalable platforms.
  • Experience supporting AI/ML or GPU infra is a plus.

Responsibilities

  • Design, build, and operate Harvey’s production infrastructure powering products and AI workloads.
  • Drive technical direction across compute, networking, Kubernetes, workflow orchestration, and production operations.
  • Lead cross-functional initiatives that improve reliability, scalability, security, and efficiency.
  • Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate requirements into resilient infrastructure solutions.
  • Establish reusable patterns, tooling, and paved paths for safe production shipping.
  • Raise engineering standards through design reviews, documentation, and mentorship.
  • Operate Harvey’s global compute and network infrastructure with high availability.
  • Improve compute utilization, performance, and service availability amid growing workloads.
  • Develop capacity models and fleet automation to scale with business growth.
  • Maintain and improve the Kubernetes platform, including provisioning and upgrades.
  • Drive cost efficiency via capacity management and monitoring.
  • Build secure foundations (IAM, network isolation, secrets, auditing).
  • Develop IaC and automation using Terraform or Pulumi.
  • Enhance observability, monitoring, alerting, and incident response.
  • Participate in on-call rotations and turn incidents into durable improvements.

Skills

5+ years exp
Multi-cloud infra
Kubernetes in prod
Distributed systems
Terraform/Pulumi
Observability
IAM & security
Cross-functional leadership
Strong communication
Systems thinking

Job description

Why Harvey

At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come.

This is a rare chance to help build a generational company at a true inflection point. We have strong product-market fit and world-class investor support. We’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth — personal, professional, and financial — is unmatched.

Our team moves fast, takes ownership, and is deeply committed to the mission — operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you.

At Harvey, the future of professional services is being written today — and we’re just getting started.

Role Overview

Harvey is building the AI platform trusted by the world’s leading law firms and enterprises. Our infrastructure is the foundation that powers every customer interaction, every model inference, and every production workload.

We’re looking for a Production Engineer to help build and operate Harvey’s core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. You’ll work on the systems that enable engineering teams to move quickly and operate reliable services at scale.

In this role, you’ll improve the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform. You’ll solve complex production challenges across compute fleet management, capacity planning, infrastructure automation, and production operations. You’ll partner closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.

At Harvey, we value Decisiveness, Simplicity, and the belief that Job’s Not Finished. We move quickly, prioritize clarity, and continuously raise the bar for engineering excellence.

What You'll Do
Infrastructure Engineering & Technical Leadership
  • Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads.

  • Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.

  • Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost.

  • Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions.

  • Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely.

  • Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship.

Infrastructure Foundation & Production Operations
  • Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.

  • Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads.

  • Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth.

  • Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation.

  • Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring.

  • Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.

  • Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.

  • Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform.

  • Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements.

What You Have
  • 5+ years of software, infrastructure, site reliability, or production engineering experience.

  • Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.

  • Strong hands‑on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.

  • Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.

  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.

  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.

  • Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.

  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.

  • A track record of driving complex, cross‑functional technical initiatives and influencing engineering decisions without relying on formal authority.

  • Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders.

  • A systems‑thinking mindset and a passion for building simple, reliable, and scalable infrastructure platforms.

Nice to Have
  • Experience supporting AI/ML or LLM infrastructure at scale.

  • Experience operating GPU fleets, high-performance compute infrastructure, or large-scale capacity planning.

  • Experience with multi-cloud infrastructure or hybrid cloud environments.

  • Experience building internal platforms or developer tooling that improves engineering velocity and production safety.

Compensation

$161,300 - $241,900 USD

Depending on your location, an Applicant Privacy Notice may apply to you. You can find all of our Applicant Privacy Notices here.

#LI-AN2

Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made by emailing accommodations@harvey.ai

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Harvey • United States

On-site
USD 161,000 - 242,000
Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Harvey • San Francisco (CA)

On-site
USD 161,300 - 241,900
Staff Software Engineer, Production Engineering
Staff Software Engineer, Production Engineering

Harvey • United States

On-site
USD 231,000 - 340,000
Senior Software Engineer, Production Engineering
Senior Software Engineer, Production Engineering

Harvey • New York (NY)

On-site
USD 161,000 - 242,000
Staff Software Engineer, Production Engineering Harvey AI New York $231,000 - $340,000/yr
Staff Software Engineer, Production Engineering Harvey AI New York $231,000 - $340,000/yr

Neura Market • New York (NY)

On-site
USD 231,000 - 340,000
Engineering Manager, Production Engineering
Engineering Manager, Production Engineering

Harvey • United States

On-site
USD 260,000 - 340,000
Senior Engineering Manager, Production Engineering
Senior Engineering Manager, Production Engineering

Neura Market • San Francisco (CA)

On-site
USD 272,000 - 355,000
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

Harvey • New York (NY)

On-site
USD 201,000 - 264,000
Senior Software Engineer, Core Infrastructure
Senior Software Engineer, Core Infrastructure

Harvey • San Francisco (CA)

On-site
USD 200,000 - 250,000
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

Harvey • San Francisco (CA)

On-site
USD 236,000 - 290,000