Machine Learning Engineer, AI Inference Solutions – Early Career

Jobtailor

Sunnyvale (CA)

On-site

USD 110,000 - 160,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Jobtailor in Sunnyvale, CA is seeking a software engineer intern to contribute production code across the ML deployment platform, model-optimization workflows, and benchmarking infrastructure.

You will pair with senior engineers on deployment workflows, performance investigations, and tool development, building validators and analyzers while learning secure coding and safety practices for autonomous-driving software.

Qualifications

  • Bachelor’s or Master’s degree in CS, ECE, or related field; degree completed before start date.
  • Strong CS fundamentals in data structures, algorithms, OS, and computer architecture.
  • Solid Python and/or C++ coding skills from coursework, internships, or projects.
  • Hands-on AI/ML experience via classes, research, internships, or personal projects.
  • Depth in at least one area: computer architecture, OS, distributed systems, or compilers.
  • Software-engineering experience via internships, coursework, open-source, or competitions.
  • Experience with coding assistants/agents like Cursor, Claude Code, or GitHub Copilot.
  • Familiarity with ML systems stacks: PyTorch, torch.compile, TensorRT, ONNX, Triton, vLLM.
  • Exposure to model optimization or GPU profiling tools.
  • Familiarity with Airflow, Temporal, Flyte, Ray, Kubeflow.
  • Experience building agentic or LLM-powered tools/workflows.
  • Open-source contributions related to PyTorch, TensorRT, vLLM, OpenAI Triton, etc.
  • Coursework, projects or publications touching ML systems.
  • Familiarity with C++ and Linux development.

Responsibilities

  • Contribute production code across the ML deployment platform, model-optimization workflows, and inference benchmarking/profiling infra.
  • Pair with senior engineers on deployment workflows, performance investigations, model-optimization experiments, and platform tooling.
  • Build, test, and maintain platform tools such as validators, performance probes, parity and sensitivity analyzers, and agentic tools.
  • Investigate and help root-cause production deployment or performance issues involving compiler, kernel, runtime, and parity bugs.
  • Collaborate with kernels, compiler, reduced-precision, parity, and model-development teams to plan model deployments to the AV stack.
  • Participate in code reviews, design discussions, and technical documentation.
  • Learn and follow secure coding, safety, and compliance practices for on-vehicle autonomous-driving software.

Skills

Python Programming
C++ Programming
AI/ML Experience
Model Optimization
Collaboration Skills
Data Structures
Algorithms
Operating Systems
Computer Architecture
Distributed Systems
Compilers
GPU Programming
Inference Optimization

Education

Bachelor’s or Master’s degree in CS, ECE, or related field

Tools

PyTorch
TensorRT
ONNX
Triton Inference Server
Cursor
Claude Code
GitHub Copilot
Airflow
Temporal
Kubeflow

Job description

  • Contribute production code across the ML deployment platform, model-optimization workflows, and inference benchmarking/profiling infrastructure
  • Pair with senior engineers on deployment workflows, performance investigations, model-optimization experiments, and platform tooling
  • Build, test, and maintain platform tools such as validators, performance probes, parity and sensitivity analyzers, and agentic specialists
  • Investigate and help root-cause production deployment or performance issues involving compiler, kernel, runtime, and parity bugs
  • Collaborate with kernels, compiler, reduced-precision, parity, and model-development teams to plan and execute model deployments to the AV stack
  • Participate in code reviews, design discussions, and technical documentation
  • Learn and follow secure coding, safety, and compliance practices for on-vehicle autonomous-driving software
Requirements
  • Recently completed or completing a Bachelor’s or Master’s degree by Spring 2026 in Computer Science, ECE, or a related technical field; degree must be completed before start date
  • Strong computer science fundamentals in data structures, algorithms, operating systems, and computer architecture
  • Solid coding skills in Python and/or C++, demonstrated through coursework, internships, or substantial projects
  • Hands-on AI/ML experience through classes, research, internships, or personal projects
  • Depth in at least one of computer architecture, operating systems, distributed systems, or compilers
  • Demonstrated software-engineering experience through internships, coursework, open-source, research code, or competitions
  • Experience with—or strong interest in—coding assistants/agents such as Cursor, Claude Code, or GitHub Copilot
  • Ability to work effectively in collaborative, cross-functional teams and communicate clearly in writing and verbally
  • Preferred: internship, research, or advanced coursework in ML systems, ML compilers, GPU programming, inference optimization, or distributed training/serving infrastructure
  • Preferred: familiarity with PyTorch and modern ML compiler/runtime stacks such as torch.compile, TensorRT, ONNX, Triton Inference Server, or vLLM
  • Preferred: exposure to model optimization or GPU profiling tools
  • Preferred: familiarity with Airflow, Temporal, Flyte, Ray, or Kubeflow
  • Preferred: experience building agentic or LLM-powered tools or workflows
  • Preferred: open-source contributions related to PyTorch, TensorRT, vLLM, OpenAI Triton, or similar projects
  • Preferred: coursework, projects, or publications touching ML systems
  • Preferred: familiarity with C++ and Linux development
Core Competencies

Demonstrates strong coding skills in Python and C++, with hands-on experience in AI/ML and a solid understanding of computer science fundamentals. Capable of collaborating effectively in cross-functional teams while adhering to secure coding and compliance practices for autonomous-driving software.

Highest-signal resume keywords
  • Python Programming
  • C++ Programming
  • AI/ML Experience
  • Model Optimization
  • Collaboration Skills
Hard Skills
  • Data Structures
  • Algorithms
  • Operating Systems
  • Computer Architecture
  • Software Engineering
  • Distributed Systems
  • Compilers
  • Performance Optimization
  • GPU Programming
  • Inference Optimization
Soft Skills
  • Effective Communication
  • Team Collaboration
Industry Keywords
  • ML Deployment
  • Model Development
  • Performance Benchmarking
  • Secure Coding Practices
  • Autonomous Driving Software
Tools & Technologies
  • PyTorch
  • TensorRT
  • ONNX
  • Triton Inference Server
  • Cursor
  • Claude Code
  • GitHub Copilot
  • Airflow
  • Temporal
  • Kubeflow
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, AI Infrastructure – LVM Inference & Evaluation
Software Engineer, AI Infrastructure – LVM Inference & Evaluation

Jobtailor • Redwood City (CA)

On-site
USD 180,000 - 240,000
Principal AI Compiler – Runtime Engineer
Principal AI Compiler – Runtime Engineer

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Principal AI Machine Learning Developer
Principal AI Machine Learning Developer

Jobtailor • Chantilly (VA)

On-site
USD 180,000 - 240,000
Staff AI/ML Software Engineer, Model Distillation, Fine-Tuning
Staff AI/ML Software Engineer, Model Distillation, Fine-Tuning

Jobtailor • California (MO)

Hybrid
USD 180,000 - 260,000
Principal ML Engineer – Embodied AI Scaling Foundations
Principal ML Engineer – Embodied AI Scaling Foundations

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Machine Learning Engineering Graduate Intern
Machine Learning Engineering Graduate Intern

Jobtailor • California (MO)

On-site
USD 34,000 - 48,000
Senior Lead AI Engineer, AI Foundations, VLM Customization
Senior Lead AI Engineer, AI Foundations, VLM Customization

Jobtailor • California (MO)

On-site
USD 180,000 - 240,000
Applied Researcher I – AI Foundations, VLM
Applied Researcher I – AI Foundations, VLM

Jobtailor • California (MO)

On-site
USD 180,000 - 240,000
Senior Manager, AI Deployment
Senior Manager, AI Deployment

General Motors • United States

On-site
USD 180,000 - 240,000
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI

Jobtailor • New York (NY)

On-site
USD 140,000 - 210,000