Infrastructure Engineer — AI-Powered Reliability

WRITER

San Francisco (CA)

Hybrid

USD 155,000 - 240,000

Full time

36 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Generous PTO
Medical/Dental/Vision
Parental leave
Galleri cancer testing
Health savings account
Learning stipend
401k and stock options

Job summary

WRITER is hiring an Infrastructure Engineer to own reliability and performance for the enterprise AI platform. You will implement scalable, fault-tolerant infrastructure across AWS/GCP/Azure, using Terraform, Helm, and Kubernetes, while embedding AI-assisted workflows into daily operations.

You will work across SRE, DevOps, and Platform teams, balancing on-call demands with long-term observability and cost-efficiency initiatives. This hybrid role is based in our global hubs, including SF and NYC.

Qualifications

  • 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company.
  • Experience running containerisation in production (a real cluster), with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred).
  • Proficiency in Python or Go for automation and tooling.
  • AI is part of how you ship — AI-assisted workflows and agentic tooling in daily practice.
  • Demonstrated ability to reason from first principles and manage blast radius/rollback.

Responsibilities

  • Own the reliability, performance, and efficiency of WRITER's core services end-to-end — define and uphold the SLOs and error budgets, carry the on-call pager, and stand behind the outcomes.
  • Shaping multi-year observability, cost, and reliability investments that align with product and revenue goals.
  • Collaborate across product, security, and engineering to design scalable, fault-tolerant systems.

Skills

Infrastructure
DevOps
Terraform
Python
Go
AI workflows
Containerization
Kubernetes
Monitoring

Tools

Prometheus

Job description

WRITER is hiring an Infrastructure Engineer to own reliability and performance for the enterprise AI platform. You will implement scalable, fault-tolerant infrastructure across AWS/GCP/Azure, using Terraform, Helm, and Kubernetes, while embedding AI-assisted workflows into daily operations.

You will work across SRE, DevOps, and Platform teams, balancing on-call demands with long-term observability and cost-efficiency initiatives. This hybrid role is based in our global hubs, including SF and NYC.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infrastructure Engineer: AI-Driven Platform Reliability (Hybrid)
Infrastructure Engineer: AI-Driven Platform Reliability (Hybrid)

WRITER • Seattle (WA)

Hybrid
USD 155,000 - 240,000
Generous PTO
Medical coverage
Dental coverage
+5
Platform Reliability Engineer — AI Infra & CloudOps
Platform Reliability Engineer — AI Infra & CloudOps

WRITER • New York (NY)

Hybrid
USD 120,000 - 160,000
Generous PTO
Medical, dental, and vision coverage
Paid parental leave (16 weeks)
+5
Backend Engineer – AI Integration Platform
Backend Engineer – AI Integration Platform

WRITER • San Francisco (CA)

Hybrid
USD 132,000 - 240,000
Generous PTO
Medical, dental, and vision coverage
Parental leave
+5
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Staff Site Reliability Engineer — AI-Driven Reliability
Staff Site Reliability Engineer — AI-Driven Reliability

EarnIn • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Equity
Hybrid work model
Senior Infrastructure Engineer — AI‑Native Cloud & Security
Senior Infrastructure Engineer — AI‑Native Cloud & Security

Elicit • Oakland (CA)

Hybrid
USD 185,000 - 260,000
Health insurance
401K with employer match
Mac workstation budget
+2
Platform Engineer – AI Infra, CI/CD & Observability
Platform Engineer – AI Infra, CI/CD & Observability

Outmarket AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Backend Reliability Engineer, AI-Powered Platform (Remote)
Backend Reliability Engineer, AI-Powered Platform (Remote)

Affirm • Madison (WI)

On-site
USD 173,000 - 255,000
Health care coverage
ESPP
Time off
+1
Senior Reliability Engineer — AI Infrastructure & SRE
Senior Reliability Engineer — AI Infrastructure & SRE

Fireworks • San Mateo (CA)

On-site
USD 150,000 - 230,000
Platform Infra Engineer - Cloud Reliability & Automation
Platform Infra Engineer - Cloud Reliability & Automation

CrewAI • United States

On-site
USD 130,000 - 210,000