Senior Software Engineer, LLM Platform

Cacheflow

San Francisco (CA)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
401(k) match
Parental leave
Wellbeing stipend
Learning stipend
Remote & Office benefits
Flexible PTO
Team outings

Job summary

Handshake is hiring a Senior LLM Platform Engineer to join our Data and ML Platform team. This role owns our shared LLM control plane and related infrastructure, enabling production workflows across Handshake’s career marketplace and Handshake AI.

You will collaborate with Backend Platform, HAI engineering, and data science to build scalable, observable, and secure services, participate in on-call, and drive platform improvements from onboarding to deployment.

Qualifications

  • Strong production software engineering experience in Python, TypeScript, Go, or similar language.
  • Experience building or operating high-throughput API gateways, proxies, or multi-tenant platform services.
  • Hands-on experience with Kubernetes, Terraform, CI/CD, and production service ownership.
  • Practical experience with LLM provider APIs, including streaming, long-running requests, retries, timeouts, cancellation, and rate limits.
  • Experience with authentication, quotas, credential management, tenant isolation, and auditability.
  • Experience building observability for distributed systems and leading production incident response.
  • Experience with usage metering, cost attribution, capacity planning, or FinOps.
  • Experience with data or ML platform systems such as BigQuery, Airflow, streaming pipelines, model serving, or ML observability.
  • Strong judgment in ambiguous environments and a bias toward simple, reliable systems that scale.

Responsibilities

  • Build and operate LiteLLM-based AI gateways and shared LLM clients.
  • Own provider and model onboarding, routing, failover, rate limits, and capacity planning.
  • Build self-service workflows for keys, access, budgets, and credentials.
  • Establish SLOs, observability, cost attribution, and alerts for production LLM traffic.
  • Qualify and roll out new models, providers, SDKs, and gateway configurations.
  • Support hosted and self-hosted inference through a consistent platform interface.
  • Partner with product, AI, and FDE teams to turn recurring delivery problems into platform capabilities.
  • Contribute to broader Data and ML Platform including workflow orchestration and model serving.
  • Participate in team on-call and improve runbooks, automation, and reliability.

Skills

Python
TypeScript
Go
API gateways
Kubernetes
Terraform
CI/CD
LLM APIs
Authentication
Observability
FinOps
BigQuery
Airflow
Model serving
Distributed systems

Tools

LiteLLM
vLLM
Modal
Ray
PyTorch

Job description

About Handshake

Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.

In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.

Why join Handshake now:
  • Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel

  • Partner hand-in-hand with world-class AI labs, Fortune 500 partners and the world’s top educational institutions

  • Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders

  • Build a massive, fast-growing business with billions in revenue

About Handshake AI

Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

About the Role

Handshake is hiring a Senior LLM Platform Engineer to join our Data and ML Platform team. This team supports Handshake’s core career marketplace and Handshake AI (HAI) by building the shared data, ML, and LLM infrastructure behind production workflows.

This infrastructure-heavy role primarily owns our shared LLM control plane: LiteLLM gateways, provider integrations, shared clients, access controls, observability, cost attribution, capacity management, and hosted or self-hosted inference. You’ll also contribute to adjacent platform systems for workflow orchestration, model serving, shared cloud infrastructure, and developer enablement—working closely with Backend Platform, HAI engineering, data science, and FDEs.

What You’ll Do
  • Build and operate our LiteLLM-based AI gateways and shared LLM clients.

  • Own provider and model onboarding, routing, failover, rate limits, and capacity planning.

  • Build self-service virtual-key, model-access, budget, and credential-management workflows.

  • Establish SLOs, observability, cost attribution, and alerts for production LLM traffic.

  • Safely qualify and roll out new models, providers, SDKs, and gateway configurations.

  • Support hosted and self-hosted inference through a consistent platform interface.

  • Partner with product, AI, and FDE teams to turn recurring delivery problems into paved-platform capabilities.

  • Contribute to the broader Data and ML Platform, including workflow orchestration, model serving, shared cloud infrastructure, and developer tooling.

  • Participate in team on-call and support, improving runbooks, automation, and reliability across owned platform services.

What We’re Looking For
  • Strong production software engineering experience in Python, TypeScript, Go, or a similar language.

  • Experience building or operating high-throughput API gateways, proxies, or multi-tenant platform services.

  • Hands-on experience with Kubernetes, Terraform, CI/CD, and production service ownership.

  • Practical experience with LLM provider APIs, including streaming, long-running requests, retries, timeouts, cancellation, and rate limits.

  • Experience with authentication, quotas, credential management, tenant isolation, and auditability.

  • Experience building observability for distributed systems and leading production incident response.

  • Experience with usage metering, cost attribution, capacity planning, or FinOps.

  • Experience with data or ML platform systems such as BigQuery, Airflow, streaming pipelines, model serving, or ML observability.

  • Strong judgment in ambiguous environments and a bias toward simple, reliable systems that scale.

Bonus Experience
  • LiteLLM, Portkey, or a comparable multi-provider AI gateway.

  • vLLM, Modal, Ray, Triton, PyTorch, or GPU-backed serving.

  • Temporal or another durable-execution system for long-running LLM requests.

  • Evaluation, batch inference, fine-tuning, RL, or other post-training infrastructure.

  • Agent runtimes, sandbox infrastructure, MCP, tool use, or coding-agent infrastructure.

Our Stack

Python, TypeScript, Go, LiteLLM, FastAPI, PostgreSQL, Redis, GCP, Kubernetes, Terraform, Spacelift, BigQuery, Airflow, Dataflow/Beam, Datastream, OpenAI, Anthropic, Gemini, OpenRouter, vLLM, Modal, Anyscale/Ray, Datadog, Arize, Temporal, Cloudflare, and Tailscale

Perks

Handshake delivers benefits that help you feel supported—and thrive at work and in life.

The below benefits are for full-time US employees.

Ownership: Equity in a fast-growing company

Financial Wellness: 401(k) match, competitive compensation, financial coaching

Family Support: Paid parental leave, fertility benefits, parental coaching

Wellbeing: Medical, dental, and vision, mental health support, $500 wellness stipend

Growth: $2,000 learning stipend, ongoing development

Remote & Office: Internet, commuting, and free lunch/gym in our SF office

Time Off: Flexible PTO, 15 holidays + 2 flex days

Connection: Team outings & referral bonuses

Explore our mission, values, and comprehensive US benefits at joinhandshake.com/careers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Machine Learning Infrastructure
Senior Software Engineer, Machine Learning Infrastructure

Apply • San Francisco (CA)

On-site
USD 170,000 - 250,000
Equity
401(k) match
Parental leave
+3
Senior Data Engineer
Senior Data Engineer

Apply • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
401(k) match
Parental leave
+3
Senior Software Engineer, Agentic Infrastructure
Senior Software Engineer, Agentic Infrastructure

Handshake • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
401(k) match
Parental leave
+6
Strategic Projects Associate
Strategic Projects Associate

Cacheflow • San Francisco (CA)

On-site
USD 70,000 - 90,000
Equity in a fast-growing company
401(k) match
Paid parental leave
+2
Member of Technical Staff, Data AI
Member of Technical Staff, Data AI

Handshake • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
401(k) match
Medical, dental, vision
+3
Manager, Strategic Programs
Manager, Strategic Programs

Apply • San Francisco (CA)

On-site
USD 150,000 - 190,000
Equity
401(k) match
Parental leave
+12
Senior Data Engineer
Senior Data Engineer

Handshake • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Equity ownership
401(k) match
Parental/fertility benefits
+6
Manager, Strategic Programs
Manager, Strategic Programs

Handshake • San Francisco (CA)

On-site
USD 140,000 - 200,000
Equity
401(k) match
Parental leave
+1
Senior Software Engineer, Coding
Senior Software Engineer, Coding

Handshake • San Francisco (CA)

Hybrid
USD 140,000 - 220,000
Ownership
Financial Wellness
Family Support
+5
Software Engineer I, Handshake AI
Software Engineer I, Handshake AI

Cacheflow • San Francisco (CA)

On-site
USD 120,000 - 180,000
Equity
401(k) match
Parental leave
+8