Platform Support Engineer

Braintrust

Seattle (WA)

On-site

USD 110,000 - 150,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Lunch & beverages
Flexible time off
Equity
Wifi stipend
Cellphone stipend

Job summary

Braintrust is hiring Platform Support Engineers to own customer-facing support for hybrid/self-hosted Braintrust deployments on AWS, Azure, and GCP. You’ll debug infrastructure, diagnose backend performance, and lead incident response with a focus on concrete fixes and reliable operations.

You will work with Kubernetes, Terraform, and cloud tooling, collaborating with Cloud Infrastructure and Engineering teams to keep customers running smoothly.

Qualifications

  • Experience in a customer-facing technical role.
  • Strong Kubernetes fundamentals.
  • Hands-on Terraform, and depth in at least one major cloud (AWS preferred).
  • Comfort in a backend codebase — Python, TypeScript, or Go.
  • Fluency with observability tooling.
  • Clear, calm, direct communication under pressure.
  • Ownership and follow-through until customer is running again.

Responsibilities

  • Own customer-facing support for hybrid and self-hosted Braintrust deployments across AWS, Azure, and GCP.
  • Debug real infrastructure problems: Kubernetes workloads, Terraform state, networking and VPC configuration, IAM and permissions, TLS, and cloud-provider quirks.
  • Diagnose performance and reliability issues in the backend — ingest throughput, query latency, database and object-store behavior.
  • Lead incident response for customer-impacting issues: triage, communicate clearly while it's on fire, and drive to resolution.
  • Ship fixes by submitting PRs to backend services, Terraform modules, and deployment tooling.
  • Build tooling for easier future incidents — diagnostics, health checks, preflight validation, self-service paths.
  • Write and maintain runbooks and deployment docs to turn hard-won answers into permanent ones.
  • Feed patterns back to Engineering and Product to reduce recurring failures.
  • Participate in an on-call rotation for critical customer issues.

Skills

Kubernetes
Python
TypeScript
Go
Observability tooling
Customer-facing
Ownership
Communication

Tools

Terraform
AWS
Docker

Job description

About The Company

Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them.

About The Company

Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them.

Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve.

About The Role

Our largest customers don't just use Braintrust — they run it. They deploy our stack inside their own AWS, Azure, and GCP accounts, behind their own VPCs, under their own compliance requirements, at their own scale. When a hybrid deployment stalls, when ingest backs up, when a query that was fast last week isn't, they come to us. Platform Support is the team that owns that. We're the technical front line for infrastructure, performance, and reliability.

We're hiring Platform Support Engineers at both mid and senior levels to join a small, high-ownership team. You'll work shoulder to shoulder with our Cloud Infrastructure and Engineering teams, and alongside our Developer Support Engineers, who own the SDK and API side of the customer experience. If you like hard infrastructure problems, and you like them more when a real customer is on the other end, this is the role.

What you'll do
  • Own customer-facing support for hybrid and self-hosted Braintrust deployments across AWS, Azure, and GCP — from first install through steady-state operation.
  • Debug real infrastructure problems: Kubernetes workloads, Terraform state, networking and VPC configuration, IAM and permissions, TLS, and cloud-provider quirks.
  • Diagnose performance and reliability issues in the backend — ingest throughput, query latency, database and object-store behavior — using logs, metrics, and traces to get to cause rather than symptom.
  • Lead incident response for customer-impacting issues: triage, communicate clearly while it's still on fire, and drive it to resolution.
  • Ship fixes. Submit PRs to our backend services, Terraform modules, and deployment tooling rather than handing every problem to Engineering.
  • Build the tooling that makes the next one easier — diagnostics, health checks, preflight validation, and self-service paths that let customers unblock themselves.
  • Write and maintain the runbooks and deployment documentation that turn one hard-won answer into a permanent one.
  • Feed patterns back to Engineering and Product, so the recurring failure modes stop recurring.
  • Participate in an on-call rotation for critical customer issues.
What we're looking for
  • Experience in a customer-facing technical role — Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering — or backend/infra engineering experience with real appetite for customer work.
  • Strong Kubernetes fundamentals: you can deploy, debug, and scale actual workloads, and read a failing pod's story from its events and logs.
  • Hands-on Terraform, and depth in at least one major cloud (AWS strongly preferred).
  • Comfort in a backend codebase — Python, TypeScript, or Go — enough to reproduce a bug, trace it to its source, and fix it.
  • Fluency with observability tooling, and the instinct to reach for data before opinion.
  • Clear, calm, direct communication under pressure, especially when the customer is technical, blocked, and losing time.
  • Ownership. You take a problem personally and follow it until the customer is running again.
Bonus points for
  • Supporting self-hosted or on-prem enterprise software, especially in regulated environments.
  • Multi-cloud experience, particularly Azure or GCP alongside AWS.
  • Database and data-infrastructure depth — Postgres, ClickHouse, or similar analytical stores.
  • Experience with observability, ML infrastructure, or developer platforms.
  • Familiarity with LLM APIs and how teams are building and evaluating agents in production.
  • Having built support or diagnostic tooling that measurably reduced ticket volume.
Why join Braintrust
  • Work on genuinely hard infrastructure problems, at the scale and pace of the teams building the best AI products in the world.
  • Join a team early enough to shape how it operates — its standards, its tooling, and its bar.
  • Sit close to both the customer and the code, with the mandate to fix things in either direction.
Benefits include
  • Medical, dental, and vision insurance
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity
  • Wifi & cellphone stipend
Equal opportunity

Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Support Engineer
Platform Support Engineer

braintrust • United States

On-site
USD 120,000 - 160,000
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
+3
Platform Support Engineer
Platform Support Engineer

Braintrust • San Francisco (CA)

Hybrid
USD 140,000 - 190,000
Medical, dental, and vision
Daily lunch, snacks
Flexible time off
+2
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Braintrust • Seattle (WA)

On-site
USD 120,000 - 160,000
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
+2
Cloud Infrastructure Engineer San Francisco, New York City, +more
Cloud Infrastructure Engineer San Francisco, New York City, +more

Braintrust Data, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
+2
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Braintrust • New York (NY)

On-site
USD 100,000 - 150,000
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
+2
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Braintrust • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+3
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Braintrust • San Francisco (CA)

On-site
USD 180,000 - 250,000
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
+2
Founding Data Engineer
Founding Data Engineer

Braintrust • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
+2
Head of Technical Solutions
Head of Technical Solutions

Braintrust • New York (NY)

On-site
USD 180,000 - 260,000
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
+2
Head of Technical Solutions
Head of Technical Solutions

Braintrust • Seattle (WA)

On-site
USD 180,000 - 280,000
Medical insurance
Lunch provided
Flexible time off
+2