Senior Software Engineer, Machine Learning Infrastructure

TrulyHired

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
401(k) match
Parental leave
Medical/dental/vision
Wellness stipend
Learning stipend
Office perks (SF)
Flexible PTO
Team outings

Job summary

Handshake is seeking a Senior Software Engineer to join the ML Infrastructure & Platform team, building the shared infrastructure behind production ML and AI systems. You will work at the intersection of software engineering, machine learning, and generative AI to move projects from prototype to production quickly and reliably.

You will help design and operate data pipelines, feature stores, training and model serving, while optimizing cost, latency, and scalability for a rapidly expanding AI

Qualifications

  • 5+ years of production software engineering experience using Python, Go, TypeScript, or similar languages.
  • Experience building and operating cloud infrastructure on AWS, GCP, or similar platforms.
  • Strong experience with Kubernetes, Docker, Terraform, CI/CD, and operating production services.
  • Hands-on experience building ML infrastructure, including model serving, training pipelines, feature stores, embeddings, or ML observability.
  • Experience with modern data platforms such as BigQuery, Airflow, Spark, Beam/Dataflow, or streaming pipelines.
  • Practical experience building production systems with LLMs or generative AI, including orchestration, provider APIs, observability, and performance optimization.
  • Strong systems design skills, sound engineering judgment, and the ability to thrive in ambiguous, fast-moving environments.

Responsibilities

  • Build and operate the shared infrastructure behind production ML and AI, including data pipelines, feature stores, training, and model serving.
  • Develop and scale our LLM platform, including provider integrations, orchestration, observability, and controls for cost, latency, and reliability.
  • Build evaluation infrastructure, including LLM eval harnesses, benchmarks, and quality measurement pipelines.
  • Support post-training workflows, including fine-tuning, reinforcement learning pipelines, and supporting data infrastructure.
  • Optimize inference infrastructure for open and fine-tuned models, including GPU serving, batching, and autoscaling.
  • Partner with AI, Data Science, and Product teams to productionize new models and establish best practices for ML infrastructure across Handshake.
  • Improve the reliability, scalability, and developer experience of our ML platform.

Skills

Python
Go
TypeScript
Cloud infra
Kubernetes
Docker
Terraform
CI/CD
ML infra
LLM platforms

Tools

Kubernetes
Docker
Terraform
CI/CD tools
BigQuery
Airflow
Spark
Beam/Dataflow

Job description

Senior Software Engineer, Machine Learning Infrastructure
About Handshake

Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.

In 2025, we started Handshake AI and built the fastest‑growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60 M to over 30 000 individuals every month.

Why join Handshake now
  • Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel
  • Partner hand‑in‑hand with world‑class AI labs, Fortune 500 partners and the world’s top educational institutions
  • Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders
  • Build a massive, fast‑growing business with billions in revenue
About Handshake AI

Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data‑intensive post‑training techniques. We believe that data spend for AI training will increase by 3‑5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

About the Role

We’re looking for a Senior Software Engineer to join our ML Infrastructure & Platform team. This team powers both Handshake’s core career marketplace and Handshake AI by building the shared infrastructure behind our production ML and AI systems.

This is an infrastructure‑heavy role for an engineer who enjoys building scalable platforms at the intersection of software engineering, machine learning, and generative AI. You’ll help teams move quickly from prototype to production while building the reliable, high‑performance systems that power training, evaluation, and inference across Handshake.

What You’ll Do
  • Build and operate the shared infrastructure behind production ML and AI, including data pipelines, feature stores, training, and model serving.
  • Develop and scale our LLM platform, including provider integrations, orchestration, observability, and controls for cost, latency, and reliability.
  • Build evaluation infrastructure, including LLM eval harnesses, benchmarks, and quality measurement pipelines.
  • Support post‑training workflows, including fine‑tuning, reinforcement learning pipelines, and supporting data infrastructure.
  • Optimize inference infrastructure for open and fine‑tuned models, including GPU serving, batching, and autoscaling.
  • Partner with AI, Data Science, and Product teams to productionize new models and establish best practices for ML infrastructure across Handshake.
  • Improve the reliability, scalability, and developer experience of our ML platform.
Desired Capabilities
  • 5+ years of production software engineering experience using Python, Go, TypeScript, or similar languages.
  • Experience building and operating cloud infrastructure on AWS, GCP, or similar platforms.
  • Strong experience with Kubernetes, Docker, Terraform, CI/CD, and operating production services.
  • Hands‑on experience building ML infrastructure, including model serving, training pipelines, feature stores, embeddings, or ML observability.
  • Experience with modern data platforms such as BigQuery, Airflow, Spark, Beam/Dataflow, or streaming pipelines.
  • Practical experience building production systems with LLMs or generative AI, including orchestration, provider APIs, observability, and performance optimization.
  • Strong systems design skills, sound engineering judgment, and the ability to thrive in ambiguous, fast‑moving environments.
Extra Credit
  • Experience with Ray, Anyscale, KubeRay, Ray Serve, vLLM, Triton, PyTorch, or GPU‑backed inference and training.
  • Experience designing LLM evaluation frameworks, benchmarking systems, or quality regression testing.
  • Experience with Vertex AI, Bigtable, Redis, or feature platform infrastructure.
  • Experience with post‑training techniques such as fine‑tuning, RLHF, reinforcement learning, or reward modeling.
  • Experience building agentic systems, MCP integrations, tool use, memory systems, or voice AI applications.
Perks

The below benefits are for full‑time US employees.

  • Equity in a fast‑growing company
  • 401(k) match, competitive compensation, financial coaching
  • Paid parental leave, fertility benefits, parental coaching
  • Medical, dental, and vision, mental health support, $500 wellness stipend
  • $2,000 learning stipend, ongoing development
  • Internet, commuting, and free lunch/gym in our SF office
  • Flexible PTO, 15 holidays + 2 flex days
  • Team outings & referral bonuses

Explore our mission, values, and comprehensive US benefits at joinhandshake.com/careers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Machine Learning Infrastructure
Senior Software Engineer, Machine Learning Infrastructure

Handshake • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity
401k matching
Parental leave
+5
Senior Software Engineer, Machine Learning Infrastructure
Senior Software Engineer, Machine Learning Infrastructure

Apply • San Francisco (CA)

On-site
USD 170,000 - 250,000
Equity
401(k) match
Parental leave
+3
Machine Learning Engineer I, Network
Machine Learning Engineer I, Network

Apply • San Francisco (CA)

On-site
USD 140,000 - 190,000
Ownership
Financial Wellness
Family Support
+5
Staff Machine Learning Engineer
Staff Machine Learning Engineer

AI Chopping Block • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Ownership: Equity
401(k) match
Parental leave
+3
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Handshake • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Ownership: Equity
401(k) match
Parental leave
+4
Senior Software Engineer, Handshake AI
Senior Software Engineer, Handshake AI

Handshake • San Francisco (CA)

On-site
USD 120,000 - 160,000
Equity in a fast-growing company
401(k) match
Flexible PTO
+2
Senior Software Engineer, Coding
Senior Software Engineer, Coding

Apply • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity in a fast-growing company
401(k) match
Parental leave
+4
Machine Learning Engineer I, Growth Relevance
Machine Learning Engineer I, Growth Relevance

Handshake • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Ownership
Financial Wellness
Family Support
+5
Senior Software Engineer, Coding
Senior Software Engineer, Coding

Handshake • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
401(k) match
Parental leave
+7
Machine Learning Engineer I, Network
Machine Learning Engineer I, Network

Handshake • San Francisco (CA)

Hybrid
USD 140,000 - 190,000
Ownership: Equity in a fast-growing co
401(k) match and financial coaching
Parental leave and family benefits
+6