Senior ML Scientist - Inference & Hardware Acceleration

Netskope

Santa Clara (CA)

On-site

USD 182,500 - 260,500

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Netskope is seeking a Senior Staff Machine Learning Scientist in Santa Clara to own the inference and optimization layer for AI in agentic workflows. You will fine-tune models, push latency and throughput on real hardware, and build a runtime that executes bounded AI tasks with real customer data signals.

You will work on quantization, KV-cache optimization, and hardware acceleration, partnering with systems and backend engineers to ship end-to-end capabilities in production environments.

Qualifications

  • MS or PhD in a technical field with focus on AI/ML research.
  • 10+ years of industry experience with 4+ years in ML/AI roles.
  • Experience with model fine-tuning and inference optimization.
  • Strong Python and capability to interface with C++ for low-level work.
  • Experience with transformer models and production-grade inference.

Responsibilities

  • Build and optimize the model inference path with quantization and KV-cache optimization.
  • Fine-tune models and develop evaluation harnesses for real accuracy and latency.
  • Design and grow the bounded task execution runtime for scalable workflows.
  • Pave the path for hardware acceleration and sparsity for larger models.
  • Collaborate with systems and backend engineers to ship end-to-end capabilities.

Skills

Python
C++ interop
Fine-tuning
Quantization
Inference runtimes
Transformers
KV cache
Hardware acceleration
Agentic AI tools

Education

MS in Computer Science / ML / EE
PhD in a related field (preferred)

Tools

vLLM
ONNX Runtime
TensorRT-LLM
llama.cpp
MLX/CoreML

Job description

Netskope is seeking a Senior Staff Machine Learning Scientist in Santa Clara to own the inference and optimization layer for AI in agentic workflows. You will fine-tune models, push latency and throughput on real hardware, and build a runtime that executes bounded AI tasks with real customer data signals.

You will work on quantization, KV-cache optimization, and hardware acceleration, partnering with systems and backend engineers to ship end-to-end capabilities in production environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Senior ML Engineer, AI Labs - Remote, LLM Inference
Senior ML Engineer, AI Labs - Remote, LLM Inference

Netskope • Santa Clara (CA)

Hybrid
USD 120,000 - 150,000
Family Medical Leave
Flexible Work Schedule
Remote Work Program
+13
Senior ML Inference Systems Engineer - Custom Accelerator
Senior ML Inference Systems Engineer - Custom Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs
401(k) matching
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization
Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization

Netskope • Santa Clara (CA)

On-site
USD 182,000 - 261,000
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Senior ML Inference Engineer - Custom Hardware, LLMs
Senior ML Inference Engineer - Custom Hardware, LLMs

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000