Infrastructure Engineer, Database

Neura Market

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 230,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

medical coverage
dental coverage
vision coverage
flexible vacation
401(k) plan
meals on in‑office days in the US

Job summary

LangChain is building SmithDB, an internal database system focused on AI observability and evaluation. We are hiring a Database Infrastructure Engineer to own the infrastructure running SmithDB across cloud environments, ensuring reliability at scale and seamless upgrades.

You will build and maintain tooling, manage Kubernetes-based services, and drive automation in deployment pipelines and capacity planning. Join a fast, autonomous team building production-grade infrastructure.

Qualifications

  • 5+ years of experience in infrastructure, platform engineering, or SRE with hands‑on.
  • Strong hands‑on experience with Kubernetes and cloud infrastructure (AWS/GCP/Azure).
  • Solid scripting/systems programming ability (Go, Python, or similar).
  • Experience with infrastructure‑as‑code and CI/CD tooling (Terraform, Helm, ArgoCD, or similar).
  • Deep familiarity with at least one major cloud provider and the primitives used to run stateful workloads reliably.
  • IaC fluency — Terraform or Pulumi/CDK as primary language.
  • Strong operational instincts — on‑call for high‑traffic data systems, runbooks that are used.

Responsibilities

  • Own the deployment and operations of SmithDB across cloud environments — cluster lifecycle management, upgrades, and automated failover.
  • Build and maintain the infrastructure tooling (Terraform, Kubernetes, Helm, or equivalent).
  • Own the Kubernetes infrastructure that runs our distributed database services (multi‑tenant, high throughput, low latency).
  • Build and improve deployment pipelines and infrastructure‑as‑code for the storage layer.
  • Drive reliability engineering: incident response, postmortems, SLOs, and disaster recovery for massive scale.
  • Manage capacity planning and cost efficiency — model growth and avoid manual intervention.
  • Build the CI/CD pipeline for database infrastructure changes — safe, tested, fast promotion from dev to prod.
  • Collaborate with SmithDB engineers to translate engine features into production infrastructure.

Skills

Kubernetes
Go/Python
Cloud infrastructure
Automation
On-call

Tools

Terraform
Kubernetes
Helm
ArgoCD
Pulumi/CDK
Python
Go

Job description

About Us

At LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. We began as widely adopted open-source tools and have grown to also offer a platform for building, evaluating, deploying, and operating agents at scale.

With $125M raised at Series B from IVP, Sequoia, Benchmark, CapitalG, and Sapphire Ventures, we’re at a stage where we’re continuing to develop new products, growth is accelerating, and all team members have meaningful impact on what we build and how we work together. LangChain is a place where your contributions can shape how this technology shows up in the real world.

Today, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. We have 100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of the Fortune 500 overall), including teams at Klarna, Clay, Coinbase, Workday, Lyft, Cloudflare, Harvey, Rippling, Vanta, LinkedIn, Monday.com, Nvidia, and Bridgewater.

About the team

SmithDB is LangChain's internal database team. We're building a storage and query layer purpose‑built for AI observability and evaluation. Within six months we went from idea to a production system that offers industry leading performance and scalability for agent observability data. We're a small, fast team of systems engineers tackling genuinely hard problems: storage layout, query execution, compaction, and scaling toward trillions of agent traces. We develop in Rust, run on Kubernetes, and integrate tightly with S3/GCS/Azure Blob. There are no legacy constraints; this is a greenfield system with real production load and ambitious engineering goals.

About the role

We're building a database specifically designed for AI observability and evaluation, and we need someone to own the infrastructure layer that keeps it running reliably at scale. As a Database Infra Engineer on the SmithDB team, you won't be designing the storage engine — you'll be making sure the engine never goes down, scales seamlessly as our customer base grows, and is operationally excellent across cloud environments.

What you'll do

  • Own the deployment and operations of SmithDB across cloud environments — including cluster lifecycle management, blue/green and rolling upgrades, and automated failover

  • Build and maintain the infrastructure tooling (Terraform, Kubernetes, Helm, or equivalent) that provisions, configures, and scales SmithDB nodes

  • Own the Kubernetes infrastructure that runs our distributed database services (multi-tenant, high throughput, low latency)

  • Build and improve deployment pipelines, rollout strategies, and infrastructure‑as‑code for the storage layer

  • Drive reliability engineering efforts: incident response, postmortems, SLOs, and disaster recovery for a system operating at massive scale

  • Manage capacity planning and cost efficiency — model growth, rightsize resources, and ensure SmithDB can absorb traffic spikes from our largest customers without manual intervention

  • Build the CI/CD pipeline for database infrastructure changes — safe, tested, and fast promotion from dev through staging to production

  • Collaborate closely with SmithDB internals engineers to translate new engine features into production‑ready infrastructure changes and ensure safe, low‑risk rollouts

What you'll bring

  • 5+ years of experience in infrastructure, platform engineering, or SRE with hands‑on

  • Strong hands‑on experience with Kubernetes and cloud infrastructure (AWS/GCP/Azure)

  • Solid scripting/systems programming ability (Go, Python, or similar);

  • Experience with infrastructure‑as‑code and CI/CD tooling (Terraform, Helm, ArgoCD, or similar)

  • Deep familiarity with at least one major cloud provider (AWS, GCP, or Azure) and the primitives used to run stateful workloads reliably — persistent volumes, managed node groups, cloud storage, etc.

  • Infrastructure‑as‑code fluency — you write Terraform (or Pulumi/CDK) as your primary language, not an afterthought

  • Strong operational instincts — you've been on‑call for high‑traffic data systems, you know how to triage under pressure, and you write runbooks that actually get used

  • Experience with container orchestration (Kubernetes) and deploying stateful workloads in production

  • A bias for automation — if you've done something manual twice, you're already thinking about how to make it never happen again

  • Strong written and oral communication skills, with the ability to translate infrastructure health into language product and business stakeholders understand

  • The DNA to thrive in a fast‑moving, high‑autonomy environment — you see gaps as opportunities and own them end to end

Nice to Have

  • Ownership of production database systems (Postgres, ClickHouse, Redis, or similar)

  • Comfort reading and reasoning about Rust is a plus, as it's the language our database is written in

  • Understanding of database reliability concepts — replication, backups, point‑in‑time recovery, connection pooling, and graceful degradation under load

Compensation

Salary Range: $180,000-$230,000 USD

Compensation Philosophy:

We offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.

Benefits
  • medical coverage
  • dental coverage
  • vision coverage
  • flexible vacation
  • a 401(k) plan
  • meals on in‑office days in the US
  • and more
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Manager, Database (SmithDB)
Software Engineering Manager, Database (SmithDB)

LangChain • San Francisco (CA)

On-site
USD 215,000 - 260,000
Medical, dental, and vision coverage
Flexible vacation
Meals on in-office days
Software Engineer, SmithDB
Software Engineer, SmithDB

Langchain • San Francisco (CA)

On-site
USD 175,000 - 240,000
Medical, dental, vision coverage
Flexible vacation
401(k) plan
+1
Principal Software Engineer, AI Observability & Evals Platform
Principal Software Engineer, AI Observability & Evals Platform

LangChain • Cambridge (MA)

On-site
USD 230,000 - 270,000
Medical, dental, and vision coverage
401(k) plan with company match
Meals on in-office days (US)
Senior Backend Engineer, LangSmith Deployments
Senior Backend Engineer, LangSmith Deployments

LangChain • Boston (MA)

On-site
USD 175,000 - 225,000
Health and dental coverage
Flexible vacation
401(k) plan
+1
Deployed Engineer (Bay Area)
Deployed Engineer (Bay Area)

LangChain • San Francisco (CA)

On-site
USD 165,000 - 315,000
Medical, dental, and vision coverage
401(k) plan
Meals on in-office days in the US
IT Systems Engineer
IT Systems Engineer

LangChain • San Francisco (CA)

On-site
USD 150,000 - 170,000
Medical,Dental,Vision coverage
401(k) plan
Flexible vacation
+1
Senior Backend Engineer, LangSmith Deployments
Senior Backend Engineer, LangSmith Deployments

LangChain • San Francisco (CA)

On-site
USD 175,000 - 225,000
Health and dental coverage
Flexible vacation
401(k) plan
+1
Deployed Architect, Post-Sales (Dallas)
Deployed Architect, Post-Sales (Dallas)

LangChain • Town of Texas (WI)

On-site
USD 180,000 - 240,000
Medical, dental and vision coverage
Flexible vacation
401(k) plan
+1
Deployed Architect, Post-Sales (Remote)
Deployed Architect, Post-Sales (Remote)

LangChain • Seattle (WA)

Remote
USD 170,000 - 190,000
Medical coverage
Dental coverage
Vision coverage
+3
Solutions Engineer (Chicago)
Solutions Engineer (Chicago)

LangChain, Inc. • Chicago (IL), Northern (KY)

Hybrid
USD 200,000 - 250,000
Medical,dental, and vision coverage
401(k) plan
Meals on in-office days in the US