Senior Machine Learning Engineer

Sailplane

San Francisco (CA)

Hybrid

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive Health, Dental, and Vision coverage
Equity grant participation
Flexible PTO with no accrual
Health and Wellness stipend
AI tools stipend

Job summary

An innovative AI infrastructure startup in San Francisco is looking for a Senior ML Engineer to lead the development and operation of LLMs in production. The position requires extensive experience in software engineering and machine learning, particularly in deploying models at scale. Candidates should be proficient in modern programming and monitoring tools, with strong collaboration skills. This hybrid role allows for working from the office 2–3 days a week and offers comprehensive health coverage and flexible PTO.

Qualifications

  • 8+ years of experience in software engineering, preferably in a VC-backed startup.
  • Experience with ML-specific monitoring tools.
  • Proficiency in programming with modern ML frameworks.

Responsibilities

  • Build, deploy, monitor, and operate LLMs in production.
  • Implement MLOps best practices to ensure reliable performance.
  • Collaborate with cross-functional teams to deliver solutions.

Skills

Machine learning algorithms
Software engineering
Communication skills
Containerization
Production infrastructure
Performance benchmarking

Tools

Python
TensorFlow
PyTorch
Docker
Kubernetes
AWS
Prometheus
Grafana

Job description

In an unmarked building somewhere in Silicon Valley, a small team of engineers is working on what could become one of the most transformative technologies in enterprise computing: autonomous infrastructure that manages itself. Sailplane, backed by AI kingmaker Khosla Ventures (OpenAI's first investor) and seed specialist True Ventures, is building a "self-driving cloud" - intelligent agents capable of autonomously managing the largest and most advanced AI infrastructure on the planet.

Sailplane is solving one of the most complex challenges in modern computing: autonomous manufacturing and management of massive AI data centers. We are creating intelligent agents that operate rack-scale systems worth millions of dollars. Think Waymo for cloud infrastructure.

"We're building million-dollar agents," explains co-founder Sam Ramji, who previously led product at Google Cloud Platform and brought Linux to Microsoft. "These aren't consumer-grade chatbots - they're sophisticated autonomous systems managing rack-scale hardware worth millions per unit."

About the Role

Sailplane is an early-stage AI infrastructure startup. Expect to wear many hats (building ML platforms, MLOps tools, data/LLM infrastructure). You will bring a startup mindset, eager to take ownership of projects, navigate ambiguity, and move quickly to solve challenging problems in a fast-paced environment.

As a Senior ML Engineer, you will lead the build and operations of LLMs in production on-premise for Sailplane. This is a senior individual contributor role focused on hands-on coding, systems thinking, and prototyping. You won't manage a team, but you will mentor and amplify those around you.

You should be fluent in models, adept at integrating production infrastructure and observability, and lead performance benchmarking. You’re comfortable working in code and in diverse production environments, and you care deeply about correctness and quality.

This hybrid position reports to the CEO and is expected to work from our downtown San Francisco office 2 to 3 days per week.

What you will do at Sailplane
  • Build, deploy, monitor, and operate LLMs in production on-premises in diverse customer environments
  • Implement MLOps best practices (CI/CD pipelines, containerization, continuous monitoring) to ensure reliable performance
  • Benchmark performance and recommend solutions to improve customer deployments including hardware sizing for target throughout (tokens per second, concurrent user sessions)
  • Experiment and iterate on models by tuning parameters and testing new approaches, continuously improving accuracy and effectiveness through rigorous evaluation
  • Document and ensure reproducibility of ML work, track experiments, code, and model versions to foster knowledge sharing and maintain high standards in the team
  • Collaborate cross-functionally with software engineers, customers, and product stakeholders
What you will bring to Sailplane
  • 8+ years of experience in software engineering, preferably in a VC-backed startup environment
  • Experience with Prometheus, Grafana, distributed tracing, or ML-specific monitoring (Weights & Biases, MLflow for production)
  • Hands-on experience deploying models at scale, including familiarity with containerization (Docker, Kubernetes) and cloud platforms (AWS, GCP, or Azure) to build and operate ML systems in production.Proficiency in programming (especially Python) and experience with modern ML frameworks/libraries such as TensorFlow, PyTorch, etc.
  • Deep understanding of machine learning algorithms and the model development lifecycle (data preprocessing, training, parameter tuning, and evaluation)
  • Proven track record of delivering software that creates real value for users
  • Excellent communication skills with an ability to explain complex ML concepts to non-experts, and a collaborative approach to working with cross-functional teams and partners
Benefits and Perks
  • Comprehensive Health, Dental, and Vision coverage beginning on the first day for employees and their families, paid 100% by Sailplane
  • Equity grant participation
  • Flexible PTO with no accrual or set annual cap, plus 15 paid holidays per year
  • Health and Wellness stipend ($3,000 annually) to help support your personal health goals
  • AI tools stipend ($1,200 annually) to encourage hands-on familiarity with emerging tools
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Product Designer
Senior Product Designer

Sailplane • San Francisco (CA)

Hybrid
USD 200,000 - 300,000
Comprehensive Health, Dental, and Vision coverage
Equity grant participation
Flexible PTO
+3
Member of Technical Staff - Distributed Systems
Member of Technical Staff - Distributed Systems

Sail Research • San Francisco (CA)

On-site
USD 120,000 - 160,000
Free meals
Studio Display at desk
Friendly office environment with a cat
Member of Technical Staff - Distributed Systems
Member of Technical Staff - Distributed Systems

Sail • San Francisco (CA)

On-site
USD 150,000 - 190,000
Staff Software Engineer, Machine Learning
Staff Software Engineer, Machine Learning

Saildrone • Alameda (CA)

On-site
USD 215,000 - 270,000
Generous Time Off
Comprehensive Health Coverage
Equity grants
+1
Senior Software Engineer
Senior Software Engineer

Trayo • San Mateo (CA)

Hybrid
USD 150,000 - 230,000
Competitive salary + equity
Small team, high impact
Hybrid flexibility in San Mateo
Senior AI Agent Engineer - Open Models & Evaluation Systems
Senior AI Agent Engineer - Open Models & Evaluation Systems

Sail Research • San Francisco (CA)

On-site
Senior Software Engineer
Senior Software Engineer

MetAntz • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Senior ML Ops Engineer
Senior ML Ops Engineer

United States Digital Space LLC • New York (NY)

On-site
USD 180,000 - 240,000
Equity
Fully paid health coverage
Dental and vision
+7
Senior Software Engineer, Full Stack
Senior Software Engineer, Full Stack

Parasail AI Inc • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer
Senior Software Engineer

SproutsAI • San Francisco (CA)

On-site
USD 120,000 - 160,000