Staff Engineer, Machine Learning

ActAI

Palo Alto (CA)

On-site

USD 190,000 - 270,000

Full time

5 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ActAI in Palo Alto, CA, is seeking a Staff Engineer, Machine Learning to own the execution layer of our intelligence, turning research and model capabilities into reliable, scalable production systems. You will work across the model lifecycle—data, training, evaluation, inference, and deployment—in a hands-on leadership role at the intersection of research, systems, and product.

We value engineers who ship robust ML systems, optimize latency and resource use, and collaborate across teams to

Qualifications

  • Ship ML systems in production, not just demos.
  • Experience with large-model training, fine-tuning, evaluation, and inference.
  • Strong software engineering and systems fundamentals.
  • Experience operating GPU-based ML workloads at scale.
  • Good technical judgment and independent problem-solving.
  • Bias toward experimentation, measurement, and shipping.
  • High standards for correctness and reliability.

Responsibilities

  • Own end-to-end ML systems powering our company from data to deployment.
  • Build and evolve training and fine-tuning pipelines for large models.
  • Design evaluation systems for capability, robustness, safety, and product performance.
  • Architect high-performance inference systems optimizing latency, GPU utilization, memory, cost, and reliability.
  • Build data pipelines for real-world and synthetic training data.
  • Establish reliable production infrastructure for deploying, monitoring, and improving models.
  • Collaborate with research and application engineering to turn model capabilities into product improvements.
  • Make pragmatic trade-offs and iterate based on real-world performance.

Skills

ML systems design
Production-grade software
GPU-based training/inference
Python
Team leadership

Tools

PyTorch
JAX
TensorFlow

Job description

There are over 5 billion users using basic applications today such email, notes, tasks, calendar and they're not AI-native. Our mission is to build proactive applications for anyone in the world, who are not used to complex prompting. We aim to bring intelligence to conversations, errands, organising and workflows, with minimal to no prompting.

Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. We believe products will greatly reduce hallucinations

Our objective is to organise anyone's life, allowing us all to spend time on valuable and meaningful things

As Staff Engineer, Machine Learning, you own the execution layer of our intelligence, turning research and model capabilities into reliable, scalable production systems.

You will work across the model lifecycle: data, training, evaluation, inference, and deployment. This is a hands-on leadership role for someone who wants to operate at the intersection of research, systems, and product.

What You'll Own
  • Own the end-to-end ML systems powering our company, from data and training to evaluation, inference, and deployment.
  • Build and evolve training and fine-tuning pipelines for large models.
  • Design evaluation systems that measure capability, robustness, safety, and real-world product performance.
  • Architect high-performance inference systems, optimizing latency, GPU utilization, memory, cost, and reliability.
  • Build data pipelines and systems for high-quality real-world and synthetic training data.
  • Establish reliable production infrastructure for deploying, monitoring, and continuously improving models.
  • Partner closely with research and application engineering to turn model capabilities into product improvements.
  • Make pragmatic technical trade-offs and rapidly iterate based on real-world performance.
What We're Looking For
  • Experience building and shipping ML systems used in production, not just research prototypes.
  • Strong understanding of modern large-model training, fine-tuning, evaluation, and inference.
  • Strong software engineering and systems fundamentals.
  • Experience operating ML workloads at meaningful scale, particularly GPU-based systems.
  • Strong technical judgment and the ability to navigate ambiguous problems independently.
  • A bias toward experimentation, measurement, and shipping.
  • High standards for correctness, reliability, and production quality.
Outcomes
  • Research and models reliably translate into production-ready solutions with clear performance and quality targets.
  • ML pipelines, training loops, and inference systems are stable, efficient, and maintainable.
  • Production issues are detected, debugged, and resolved quickly, minimizing user impact.
  • Team members are supported, aligned, and able to deliver high-impact ML work with minimal friction.
  • Iterations on models and systems are measurable, safe, and improve user experience over time.
  • Python
  • PyTorch / JAX
  • GPU-based training and inference system
Ideal Experience
  • You have built or shipped real ML systems used by people, not just demos.
  • You are comfortable working with large models and understanding their failure modes.
  • You write strong, production-grade code and care about system correctness.
How We Work

We are a small, high-talent-density, hands-on team. Engineers have broad ownership and are expected to exercise strong judgment and execute independently.

We make decisions quickly, work closely together, and balance speed with engineering fundamentals. We care less about process and more about building something exceptional.

If there appears to be a fit, we'll reach to schedule 3, but no more than 4 interviews.

Applications are evaluated by our technical team members. Interviews will be conducted via virtual meetings and/or onsite.

We value transparency and efficiency, so expect a prompt decision. If you've demonstrated the exceptional skills and mindset we're looking for, we'll extend an offer to join us. This isn't just a job offer; it's an invitation to be part of a team that's bringing AI to have practical benefits to billions globally.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Machine Learning Engineer
Principal Machine Learning Engineer

A1 • Palo Alto (CA)

On-site
USD 347,000 - 490,000
Technical Product Lead, AI Email App 6 locations · Hybrid · Full-time →
Technical Product Lead, AI Email App 6 locations · Hybrid · Full-time →

ActAI • Northern (KY)

Hybrid
USD 150,000 - 230,000
Lead Engineer, Machine Learning
Lead Engineer, Machine Learning

Salt Digital Recruitment • United States

On-site
USD 180,000 - 260,000
Engineering Lead, AI Email App
Engineering Lead, AI Email App

ActAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Backend Engineer, AI Systems
Backend Engineer, AI Systems

ActAI • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Founding Forward Deployed Machine Learning Engineer
Founding Forward Deployed Machine Learning Engineer

adaption • San Francisco (CA)

On-site
USD 100,000 - 140,000
Flexible work
Annual travel stipend
Weekly meal allowance
+2
Backend Engineer, AI Systems
Backend Engineer, AI Systems

A1 • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Head of Product, AI
Head of Product, AI

A1 • Palo Alto (CA)

On-site
USD 120,000 - 160,000
AIML Engineer
AIML Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Performance bonus (up to 20% of base)
Equity participation
Health, dental, and vision insurance
+3
Backend Engineer, AI
Backend Engineer, AI

A1 • Palo Alto (CA)

On-site
USD 255,000 - 405,000