Principal Machine Learning Engineer

Protingent

Washington

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Protingent Staffing is seeking a Principal Machine Learning Engineer to lead the design and evolution of critical ML systems across training, inference, evaluation, and infrastructure in a fully remote role.

You will architect large-scale pipelines, optimize GPU usage, and own production deployment while collaborating with backend/mobile/desktop teams to deliver measurable quality improvements and reliable performance.

Qualifications

  • Strong background in deep learning and transformer architectures.
  • Hands-on experience training, fine-tuning, or deploying large-scale ML models in production.
  • Proficiency with PyTorch or JAX and ability to learn others quickly.
  • Experience with distributed training and inference frameworks (DeepSpeed, FSDP, Megatron, Ray).
  • Strong software engineering fundamentals – robust, production-grade systems.
  • Experience with GPU optimization, memory efficiency, quantization, and mixed precision.
  • Comfort owning end-to-end ML systems in production.
  • Bias toward shipping and iterative improvement.
  • Experience with LLM inference frameworks such as vLLM/TensorRT-LLM or FasterTransformer.
  • Contributions to open-source ML or systems libraries.

Responsibilities

  • Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Design reproducible, high-performance training pipelines across GPU infrastructure.
  • Architect inference systems balancing latency, throughput, cost, and reliability at scale.
  • Design and maintain data systems for synthetic and real training data.
  • Implement evaluation pipelines for performance, robustness, safety, and bias.
  • Own production deployment including GPU optimization and latency reduction.
  • Collaborate with application engineering to integrate ML into backend, mobile and desktop products.
  • Make pragmatic trade-offs and ship improvements quickly from real usage.
  • Work under production constraints: latency, cost, reliability, safety.
  • Ensure ML systems are reliable, scalable and meet performance targets.
  • Deploy models achieving measurable quality improvements and user impact.
  • Monitor, debug, and resolve production issues with root-cause analysis.
  • Provide guidance and scalable ML solutions for teammates.
  • Enable efficient research-to-production cycles improving product experience.

Skills

Deep learning
Transformer models
PyTorch
JAX
Distributed training
GPU optimization
Memory efficiency
Quantization
Mixed precision
RLHF pipelines
Diffusion models
Multimodal models
vLLM
TensorRT-LLM
FasterTransformer
Open-source

Tools

DeepSpeed
FSDP
Megatron
Ray
Spark
Apache Arrow

Job description

Job Title

Principal Machine Learning Engineer

Job Description

Protingent Staffing is offering a direct‑hire Principal Machine Learning Engineer role that is fully remote. This role is a deep technical authority responsible for designing and evolving the most critical ML systems across training, inference, evaluation, and infrastructure.

Responsibilities
  • Architect and build large‑scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Design reproducible, high‑performance training pipelines across GPU infrastructure.
  • Architect inference systems that balance latency, throughput, cost, and reliability at scale.
  • Design and maintain data systems for high‑quality synthetic and real‑world training data.
  • Implement evaluation pipelines for performance, robustness, safety, and bias in partnership with research leadership.
  • Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies.
  • Collaborate closely with application engineering to integrate ML systems cleanly into backend, mobile, and desktop products.
  • Make pragmatic trade‑offs and ship improvements quickly, learning from real usage.
  • Work under real production constraints: latency, cost, reliability, and safety.
  • Ensure ML systems (training, inference, evaluation) are reliable, scalable, and meet defined performance targets.
  • Deploy models that achieve measurable quality improvements and meet user‑impact goals.
  • Proactively monitor, debug, and resolve production issues with clear root‑cause analysis.
  • Provide clear guidance, best practices, and scalable ML solutions for teammates and cross‑functional collaborators.
  • Enable efficient, safe research‑to‑production cycles that continuously improve the product experience.
Qualifications
  • Strong background in deep learning and transformer‑based architectures.
  • Hands‑on experience training, fine‑tuning, or deploying large‑scale ML models in production.
  • Proficiency with at least one modern ML framework (e.g., PyTorch, JAX) and ability to learn others quickly.
  • Experience with distributed training and inference frameworks (DeepSpeed, FSDP, Megatron, ZeRO, Ray).
  • Strong software engineering fundamentals – write robust, maintainable, production‑grade systems.
  • Experience with GPU optimization, including memory efficiency, quantization, and mixed precision.
  • Comfort owning ambiguous, zero‑to‑one ML systems end‑to‑end.
  • Bias toward shipping, learning fast, and improving systems through iteration.
  • Must have experience with LLM inference frameworks such as vLLM, TensorRT‑LLM, or FasterTransformer.
  • Contributions to open‑source ML or systems libraries.
  • Background in scientific computing, compilers, or GPU kernels.
  • Experience with RLHF pipelines (PPO, DPO, ORPO).
  • Experience training or deploying multimodal or diffusion models.
  • Experience with large‑scale data processing (Apache Arrow, Spark, Ray).
Job Details
  • Job Type: Direct Hire
  • Pay Range: Market Rate
  • Location: Fully Remote
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Protingent • Washington

Hybrid
USD 180,000 - 240,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Protingent • Washington

Hybrid
USD 180,000 - 250,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

European Recruitment BV • United States

On-site
USD 180,000 - 230,000
Technical Lead, Machine Learning
Technical Lead, Machine Learning

Protingent • Washington

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Protingent • Washington

Hybrid
USD 120,000 - 210,000
Remote Principal ML Systems Architect
Remote Principal ML Systems Architect

Protingent • Washington

Hybrid
USD 180,000 - 240,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

On behalf of Next Deavor • New York (NY)

Hybrid
USD 200,000 - 250,000
Staff/ Principal Machine Learning (AI) Engineer
Staff/ Principal Machine Learning (AI) Engineer

Provectus • Michigan

On-site
USD 120,000 - 160,000
Health, dental, and vision insurance
401(k) with company match
Generous PTO
+1
Technical Lead - Machine Learning
Technical Lead - Machine Learning

USA Tech Recruit • San Francisco (CA)

On-site
USD 180,000 - 230,000
Applied AI Engineer
Applied AI Engineer

Protingent • Washington

Hybrid
USD 120,000 - 180,000