Staff AI/ML Software Engineer, Model Distillation, Fine-Tuning

Jobtailor

California (MO)

Hybrid

USD 180,000 - 260,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Jobtailor is seeking a senior ML engineer to design and build knowledge distillation pipelines, transferring reasoning, vision, and language capabilities into compact edge-deployable architectures. You will guide ML architecture decisions and lead edge-optimized model strategies across teams.

You will apply parameter-efficient fine-tuning (LoRA/QLoRA), implement quantization-aware training, and collaborate on data collection, dataset curation, and model evaluation for safe, scalable in-vehicle

Qualifications

  • Bachelor's degree in Computer Science, Machine Learning, Data Science, Mathematics, or equivalent practical experience
  • 8+ years of software engineering or applied ML research experience
  • Experience setting technical direction, making architectural decisions on ML systems, and guiding other engineers
  • Based in or willing to work hybrid out of Mountain View, CA or Seattle, WA, reporting to the office at least three days per week
  • Master's degree or Ph.D. in Computer Science, Artificial Intelligence, or related field preferred

Responsibilities

  • Design and build knowledge distillation pipelines transferring reasoning, vision, and language capabilities into compact edge-deployable architectures
  • Apply and scale parameter-efficient fine-tuning techniques such as LoRA and QLoRA
  • Build and own the reinforcement learning flywheel using human-in-the-loop alignment with RLHF/DPO
  • Connect in-cabin data collection with continuous model improvement
  • Curate, evaluate, and synthetically generate datasets for passenger intent and complex in-vehicle visual cues
  • Implement Quantization-Aware Training and related techniques to prevent accuracy degradation during hardware compression
  • Establish evaluation frameworks and benchmarks measuring hallucination rates, domain accuracy, and safety constraints
  • Own the base model strategy and decide which foundation architectures to use
  • Set architectural direction for model optimization pipelines as an individual-contributor technical leader
  • Validate AI capabilities on representative vehicle hardware and chart practical paths to scale

Skills

PyTorch
Knowledge Distillation
Fine-Tuning
Quantization
Reinforcement Learning
Edge Deployment
Model Compression
Open-source
Communication

Education

Bachelor's degree
Master's degree
PhD

Tools

Hugging Face
DeepSpeed
Ray
Megatron

Job description

  • Design and build knowledge distillation pipelines transferring reasoning, vision, and language capabilities into compact edge-deployable architectures
  • Apply and scale parameter-efficient fine-tuning techniques such as LoRA and QLoRA
  • Build and own the reinforcement learning flywheel using human-in-the-loop alignment with RLHF/DPO
  • Connect in-cabin data collection with continuous model improvement
  • Curate, evaluate, and synthetically generate datasets for passenger intent and complex in-vehicle visual cues
  • Implement Quantization-Aware Training and related techniques to prevent accuracy degradation during hardware compression
  • Establish evaluation frameworks and benchmarks measuring hallucination rates, domain accuracy, and safety constraints
  • Own the base model strategy and decide which foundation architectures to use
  • Set architectural direction for model optimization pipelines as an individual-contributor technical leader
  • Validate AI capabilities on representative vehicle hardware and chart practical paths to scale
Requirements
  • Bachelor's degree in Computer Science, Machine Learning, Data Science, Mathematics, or equivalent practical experience
  • 8+ years of software engineering or applied ML research experience
  • Experience setting technical direction, making architectural decisions on ML systems, and guiding other engineers
  • Deep proficiency in PyTorch
  • Hands-on experience fine-tuning large language models or vision-language models
  • Practical experience with at least two of knowledge distillation, parameter-efficient fine-tuning, pruning, or quantization
  • Based in or willing to work hybrid out of Mountain View, CA or Seattle, WA, reporting to the office at least three days per week
  • Master's degree or Ph.D. in Computer Science, Artificial Intelligence, or a related field (preferred)
  • Experience shipping a quantized model to a specific hardware target
  • Familiarity with Hugging Face, DeepSpeed, Ray, or Megatron
  • Experience managing dataset pipelines at scale
  • Domain experience in conversational AI, human-computer interaction, smart spaces, or deploying multimodal models in consumer-facing products
  • Open-source contributions or published research in model compression, distillation, or efficient AI
  • Ability to communicate complex AI training concepts and architectural trade-offs to cross-functional teams
Core Competencies

Demonstrates expertise in designing and optimizing machine learning architectures, particularly in knowledge distillation and reinforcement learning. Proficient in fine-tuning large language models and implementing quantization techniques for edge deployment.

Highest-signal resume keywords
  • Deep Proficiency In Pytorch
  • Knowledge Distillation
  • Parameter-Efficient Fine-Tuning
  • Quantization-Aware Training
  • Architectural Decision-Making
Hard Skills
  • Machine Learning
  • Software Engineering
  • Data Science
  • Model Optimization
  • Reinforcement Learning
  • Dataset Management
  • Quantization
  • Fine-Tuning
  • Human-Computer Interaction
  • Conversational AI
Soft Skills
  • Communication
  • Technical Leadership
Certifications & Qualifications
  • Bachelor's Degree In Computer Science
  • Master's Degree In Artificial Intelligence
  • Ph.D. In Computer Science
Industry Keywords
  • Edge-Deployable Architectures
  • Model Compression
  • AI Training Concepts
  • Multimodal Models
  • Consumer-Facing Products
Tools & Technologies
  • Hugging Face
  • DeepSpeed
  • Ray
  • Megatron
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied Researcher I – AI Foundations, VLM
Applied Researcher I – AI Foundations, VLM

Jobtailor • California (MO)

On-site
USD 180,000 - 240,000
Principal ML Engineer – Embodied AI Scaling Foundations
Principal ML Engineer – Embodied AI Scaling Foundations

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Software Engineer, AI Platform
Software Engineer, AI Platform

Triwill Group • San Francisco (CA)

Hybrid
USD 140,000 - 180,000
AI Architect – Media
AI Architect – Media

Jobtailor • Atlanta (GA)

On-site
USD 180,000 - 240,000
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI

Jobtailor • New York (NY)

On-site
USD 140,000 - 210,000
AI/ML Engineer
AI/ML Engineer

Winaxis LLC • Dallas (TX)

On-site
USD 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

Orbien LLC • Maryland

On-site
USD 150,000 - 230,000
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Jobtailor • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Staff Machine Learning Engineer – ML Frameworks
Staff Machine Learning Engineer – ML Frameworks

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
AI/ML Engineer
AI/ML Engineer

Jobtailor • Colorado

On-site
USD 130,000 - 180,000