Lead Infrastructure & Platform Engineer

StudyFetch

Beverly Hills (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) with employer matching
Employer-paid Medical, Dental, and Vis
Daily team dinner provided
Small mission-driven team

Job summary

StudyFetch is the #1 AI-native learning platform globally, transforming how millions learn through personalized AI-powered education. This in-person role in Beverly Hills offers high ownership of the GPU and cloud infrastructure powering our Learn Engine and Honen workforce platform.

You’ll own the platform end-to-end—from IaC and Kubernetes to networking and databases—lead the GPU colo buildout, ensure reliability, security, and on‑call readiness, and set patterns the team will follow as we

Qualifications

  • 7+ years building and operating production infrastructure.
  • Experience with physical or colo infrastructure.
  • Fluent in cloud and Kubernetes; IaC mindset.
  • Ability to design GPU buildout and scalable platforms.
  • Strong security and compliance practices.

Responsibilities

  • Own the platform end-to-end (IaC, Kubernetes, networking, CI/CD, databases).
  • Ensure reliability, incident response, and postmortems.
  • Lead GPU infrastructure buildout in colocated data center.
  • Manage GPU serving platform and model workloads.
  • Shape deployment patterns and security posture.
  • Collaborate with team and scale the platform.

Skills

Kubernetes
Cloud platforms
Networking
Bare-metal provisioning
IaC
GPU infra
Security & compliance
PostgreSQL
MongoDB
Incident response

Tools

PostgreSQL
MongoDB
Redis

Job description

StudyFetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered education. We're growing fast with backing from top-tier investors and a mission that's redefining the future of education and ethical learning.

Why this role exists

We're a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most.

All of that runs on a mature, infrastructure-as-code platform: cloud, Kubernetes, networking, deploy tooling, databases, and the GPUs serving that powers our AI. We're hiring a lead to own that platform and take it further as we scale.

This is a high-ownership role. You'll be the person the rest of engineering depends on to ship, and you'll set the patterns for how we deploy, secure, and operate. You'll also lead the next chapter: we're building out our own GPU infrastructure inside a colocation facility, which means real networking, capacity planning, and bare-metal platform work sitting alongside our cloud footprint. If you want a surface where the decisions are yours and the impact reaches millions of learners, this is it.

What we believe

Every learner deserves the chance to succeed. StudyFetch started with one idea: high-quality, personalized learning should be within reach for anyone, at any stage of life. Honen carries that belief into the workforce.

  • Accessible to everyone. Learning should reach every person, whatever their background, role, and resources.
  • Meet people where they are. Every course adapts to a person's pace, their level, and the way they learn best.
  • Learning never stops. From a first job to a new career, people keep growing at every stage of life.

We hire people who share this conviction. The work is demanding and the hours can be long, and what sustains you through it is caring whether a real student finally understands the material.

What you'll own
  • The platform, end to end. Our infrastructure-as-code (Pulumi/TypeScript on GCP), the Kubernetes clusters, the Shared VPC and networking, the CI/CD and deploy tooling, secrets, and the databases behind them. You own how the whole thing fits together and how the team ships on top of it.
  • Reliability and the on-call that follows. Monitoring, alerting, incident response, and the postmortems that make the next incident less likely. When production has a bad night, you're the person who understands why and makes sure it doesn't repeat.
  • Our own GPU infrastructure, from the ground up. We're standing up GPU capacity in a colocation facility. You'll help design and build it: hardware and capacity planning, physical and virtual networking, the platform layer that makes those GPUs usable for model serving, and the path that connects it cleanly to our cloud environment. This is a meaningful part of the role, and it's new ground for the company.
  • The GPU serving platform. The clusters and pipelines that run our self-hosted models and speech/ASR workloads, keeping latency, cost, and utilization where they need to be for real users at scale.
  • Security and compliance posture. We hold ourselves to a real bar (SOC 2, continuous scanning, least-privilege IAM). You keep the platform audit-ready without slowing the team down.
  • The bar for the team. The patterns you set for how we deploy, secure, and operate are the ones everyone else follows. As the work grows, you'll shape and help grow the team that does it with you.
What we're looking for

You're a strong fit if either of these is true:

  • 7+ years building and operating production infrastructure, with real depth in cloud platform, Kubernetes, and networking, or
  • You were the founding or early infrastructure engineer who owned a significant share of a real platform. Fewer years on paper, but you took production infrastructure from early and messy to reliable and are able to demonstrate it.

Beyond that:

  • You've run production infrastructure that real users depend on, not lab setups. You can walk us through a platform you built or operated, which parts were yours, an incident that went badly, and what you changed because of it.
  • You have hands‑on experience with physical or colocation infrastructure. Bare‑metal provisioning, datacenter or colo networking, hardware and capacity planning, GPU fleets, or standing up a hybrid of on‑prem and cloud. This is the newest part of the role, and experience here is a real differentiator.
  • You're fluent in modern cloud and Kubernetes. Infrastructure‑as‑code is how you work, not a thing you tolerate. You have opinions about ownership boundaries, blast radius, and what belongs where.
  • You use AI everyday and have informed opinions about it. You understand the demands of serving models in production, and genuine curiosity is the one thing we can't teach.
  • You've worked through launch crunch and know how you stay effective and level‑headed under pressure.
  • You're candid about tradeoffs. You can tell us what surprised you last time and what you'd do differently.
  • The mission is why you're here. What sustains you through the hard weeks is the learner on the other end, the one who finally understands because of something you kept running.
The stack you'll work in

You don't need every item below, but you should be deep in most and able to ramp quickly on the rest:

  • Cloud & platform: GCP, Cloudflare, Vercel; infrastructure‑as‑code (Pulumi or Terraform)
  • Orchestration: Kubernetes (GKE), Helm, container build pipelines, CI/CD (Jenkins or similar)
  • Networking: VPC design, IPAM, firewalls, DNS/TLS, zero‑trust access, and — for the colo work — physical and datacenter networking
  • GPU & AI serving: GPU cluster operations, model/inference serving, vLLM, ASR/speech workloads, cost and latency tuning
  • Hardware / colocation: bare‑metal provisioning, capacity planning, hybrid on‑prem + cloud
  • Data: PostgreSQL and MongoDB, managed caches (Redis/Valkey), backups and disaster recovery
  • Observability & security: monitoring and alerting, incident response, least‑privilege IAM, secrets management, SOC 2‑grade compliance
What to expect
  • This is an in‑person role at a fast pace, with periods of intense work around major launches and the GPU buildout.
  • You'll have significant ownership and autonomy with limited oversight. The role suits engineers who do their best work with room to run.
  • It's a strong fit for engineers who have operated production infrastructure that real users depend on. If your experience has been primarily in tutorials, labs, or managed setups you never had to run, this likely isn't the right match.
  • 100% employer‑paid Medical, Dental, and Vision; 75% dependent coverage
  • 401(k) with employer matching
  • Daily team dinner provided in‑office
  • A small, mission‑driven team changing how the world learns
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps & Security Engineer
DevOps & Security Engineer

StudyFetch • Beverly Hills (CA)

On-site
USD 180,000 - 260,000
Medical, dental, and vision coverage
401(k) with employer match
Daily in-office team dinner
+1
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Principal Product Engineer, Cloud Platform
Principal Product Engineer, Cloud Platform

Verdigris Technologies Inc • Palo Alto (CA)

On-site
USD 130,000 - 180,000
Staff Engineer, Applied AI
Staff Engineer, Applied AI

StudyFetch • Beverly Hills (CA)

On-site
USD 170,000 - 270,000
100% employer-paid Medical, Dental, and Vision
75% dependent coverage
401(k) with employer matching
+1
Senior / Lead / Principal Platform Engineer
Senior / Lead / Principal Platform Engineer

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Opportunities for career growth in a high-growth AI company
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Long-term career growth opportunities
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Medical, dental, vision benefits
Unlimited PTO
Generous paid parental leave
+1
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
Founding Engineer - Platform
Founding Engineer - Platform

uRun • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary and equity
Full health, dental, and vision coverage
401(k) retirement savings
+4
Platform Engineer
Platform Engineer

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Meals provided (breakfast, lunch, and/
Snacks, drinks and coffee daily
+2