Software Engineer, Infrastructure

descript

United States

On-site

USD 220,000 - 292,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Competitive benefits

Job summary

descript owns the platform that the company runs on: compute and deployment, reliability and on-call, CI/CD and monorepo health, and the infrastructure that model training and inference run on, with security boundaries around it.

You will contribute to how models are trained and served here, as well as how agents work inside our codebase: the environments they run in, the verification that makes their output trustworthy, and the review paths that keep it all legible.

Qualifications

  • 8+ years building and operating production distributed systems or equivalent server-side engineering with heavy infrastructure focus.
  • Experience leveraging agents to multiply impact and integrating AI into workflows.
  • Experience with reliability, deployment decisions, and incident management.
  • Experience carrying a pager, handling incidents, and performing rollbacks.

Responsibilities

  • Own the platform including GCP, Kubernetes, Temporal, and GPU fleet infrastructure.
  • Own AI enablement substrate: capacity, training and inference pipelines, and production serving costs.
  • Balance cost with performance; apply metering and attribution behind usage.
  • Ensure security with identity, secrets management, and least-privilege controls.
  • Write infrastructure-as-code and maintain runbooks and in-repo context for onboarding.

Skills

Distributed systems
Production deployment
Incident response
SLOs and error budgets
Kubernetes in production
Infrastructure as code
Cloud platforms

Tools

Kubernetes
Terraform
IaC

Job description

Platform owns the foundation the company runs on: compute and deployment, reliability and on-call, CI/CD and monorepo health, developer environments, the infrastructure that model training and inference run on, and the security boundaries around all of it.

AI and agent tooling is the clearest example. You will contribute to how models are trained and served here, as well as how agents work inside our codebase: the environments they run in, the verification that makes their output trustworthy, and the review paths that keep it all legible. There is no industry standard and you'll help form our opinions rather than inheriting one.

Our users are other engineers. Expect to spend time with all the other engineering teams, understanding their needs and building a roadmap. The scope is large, so you'll be choosing what to leave alone as much as build.

Architecture decisions here last: this is a small team covering a large surface. You'll have real room to decide things, and you'll stay close to the systems you decide about.

What you'll own
  • Own our platform: GCP, Kubernetes, Temporal, the GPU fleet behind cloud export, and the deploy and rollback machinery everything ships through. You'll be in the on-call rotation, and we'll expect you to make it quieter and more actionable.
  • The AI enablement substrate: GPU capacity, training and inference pipelines, and the reliability and cost of the systems serving models in production.
  • Cost is an engineering constraint: infrastructure decisions have a number attached and you're expected to make smart trade-offs. As inference grows with usage, the metering and attribution behind those numbers sit with this team.
  • Security comes with the systems you run : identity and access, secrets management, least-privilege boundaries, and supply-chain integrity.
  • Make what you build legible. Infrastructure-as-code that explains itself. Runbooks and in-repo context written for someone jumping in to help. Observability that tells you what failed and why.
  • Improve how the team learns and ships. Form hypotheses, instrument your work, release incrementally, and read results honestly. Strengthen the tooling, standards, tests, observability, and release practices that help the team move quickly without compromising quality.
  • Raise the team's technical ambition. Provide architectural direction, thoughtful reviews, mentoring, and clear human writing.
What you bring
Required:
  • You have 8+ years building and operating production distributed systems, or equivalent server-side engineering with a heavy infrastructure focus.
  • You effectively leverage agents to multiply your impact and think critically about how and when to harness AI in your work.
  • You’ve run systems where failure was expensive, and your opinions about reliability and deployment come from consequences rather than reading.
  • You’ve carried a pager, commanded an incident, and rolled back before you understood why.
  • You use SLOs and error budgets as operating tools.
  • You’ve used a major cloud provider and Kubernetes in production, with infrastructure-as-code as your default.
  • You’ve owned an architecture or migration, from planning through launch, whose consequences outlived the project, and you can say what you'd do differently.
  • You’ve found important unowned work, scoped it, earned buy-in, and delivered it without being handed a spec.
  • You build a minimal repro, read the logs, and write targeted checks to prove a fix works instead of trusting output.
Experience that helps:
  • GPU and ML infrastructure: capacity planning, training or inference pipelines, serving cost and latency. If you haven’t run a GPU fleet, experience with expensive capacity-constrained systems transfers well.
  • Production security engineering: IAM, secrets, supply chain, least privilege. You don't need to have held a security title, but you should have owned these boundaries for systems you ran.
  • Cloud cost modeling: commitment strategy, reservations, unit economics.
  • CI/CD: at monorepo scale, and developer-environment work.
  • Intricacies of Video: Media, video, or GPU-backed workloads.
  • Small teams owning a large surface: whether at a Series B to D company or on an internal platform team.
Compensation and benefits
  • Base salary: $220,000 to $292,000, plus equity and benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Platform Engineer
Platform Engineer

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Meals provided (breakfast, lunch, and/
Snacks, drinks and coffee daily
+2
Senior / Lead / Principal Platform Engineer
Senior / Lead / Principal Platform Engineer

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Significant technical ownership
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Long-term career growth opportunities
Founding Infrastructure Engineer
Founding Infrastructure Engineer

Matterhaul Inc. • San Francisco (CA)

On-site
USD 200,000 - 260,000
Equity options
Hardware + AI coding budget
Real office in SF
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Founding Engineer - Platform
Founding Engineer - Platform

uRun • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary and equity
Full health, dental, and vision coverage
401(k) retirement savings
+4
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

Psi • Boston (MA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Meaningful equity
Competitive compensation
Benefits
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Observable Intuition • New York (NY)

On-site
USD 180,000 - 280,000
Principal Product Engineer, Cloud Platform
Principal Product Engineer, Cloud Platform

Verdigris Technologies Inc • Palo Alto (CA)

On-site
USD 130,000 - 180,000