Principal Machine Learning Engineer

On behalf of Next Deavor

New York (NY)

Hybrid

USD 200,000 - 250,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Next Deavor in the New York City metro area is seeking a Principal Machine Learning Engineer to own the ML infrastructure behind real-time compliance enforcement systems. You will build training pipelines, evaluation workflows, and production-serving capabilities to train, measure, and deploy models with a strong focus on latency, reliability, and cost.

Hybrid work arrangement applies with 3 days onsite in NYC.

Qualifications

  • 8+ years of software engineering experience
  • 4+ years building ML or LLM infrastructure for production
  • Hands-on PyTorch, distributed training, LoRA, SFT
  • Experience with eval harnesses, regression gates, dataset pipelines
  • Production model serving with latency, reliability, and cost constraints

Responsibilities

  • Build and own training pipelines including data prep, reproducible fine-tuning runs, experiment tracking, and release automation
  • Develop evaluation infrastructure with automated eval runs, regression gates, dashboards, and dataset versioning
  • Own production model serving for low-latency inference, including batching, optimization, autoscaling, and cost management
  • Ship model updates safely using versioning, canarying, rollback, and drift monitoring
  • Create repeatable workflows to adapt models to new domains and changing customer needs
  • Convert expert labels and reviewer feedback into clean training and evaluation datasets
  • Help raise the team’s engineering bar for ML infrastructure as the organization grows

Skills

8+ years engineering
4+ years ML infrastructure
Python
CI/CD
Observability
Cloud infrastructure
LLM infrastructure
Team collaboration

Tools

PyTorch
LoRA
SFT
vLLM
TensorRT-LLM
Distributed training

Job description

Hybrid in the New York City Metro area (3 days onsite). In this Principal Machine Learning Engineering role, you will own the ML infrastructure behind real-time compliance enforcement systems. Expect a position centered on building the pipelines, evaluation workflow, and production serving needed to train, measure, and deploy models reliably, with a focus on latency, reliability, and cost.

Pay range: USD $200,000 - $250,000 per year.

What you’ll do
  • Build and own training pipelines including data preparation, reproducible fine-tuning runs, experiment tracking, and release automation.
  • Develop evaluation infrastructure with automated eval runs, regression gates, dashboards, and dataset versioning.
  • Own production model serving for low-latency inference, including batching, optimization, autoscaling, and cost management.
  • Ship model updates safely using versioning, canarying, rollback, and drift monitoring.
  • Create repeatable workflows to adapt models to new domains and changing customer needs.
  • Convert expert labels and reviewer feedback into clean training and evaluation datasets.
  • Help raise the team’s engineering bar for ML infrastructure as the organization grows.
What you’ll bring
  • 8+ years of software engineering experience, including 4+ years building ML or LLM infrastructure for production.
  • Hands-on experience with the modern LLM stack: PyTorch, distributed training, and fine-tuning at scale (e.g., LoRA, SFT) using inference engines such as vLLM or TensorRT-LLM.
  • Experience building eval harnesses, regression gates, or dataset pipelines, with strong understanding of precision, recall, and calibration.
  • Proven ownership of production model serving with real latency, reliability, and cost constraints.
  • Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability.
  • Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions.
  • Experience collaborating tightly with research partners and defining clear interfaces.
Technologies you’ll work with
  • PyTorch, LoRA, SFT, vLLM, TensorRT-LLM
  • Python, containers, CI/CD, cloud infrastructure, observability
Additional qualifications that may help
  • Experience productionizing small or specialized language models.
  • Experience with structured-output serving or constrained decoding in production.
  • Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety).
  • Experience deploying models into customer-controlled environments.

Work location: Hybrid remote in New York, NY 10001 (3 days onsite).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Architect for Real-Time Compliance (Hybrid NYC)
ML Infra Architect for Real-Time Compliance (Hybrid NYC)

On behalf of Next Deavor • New York (NY)

Hybrid
USD 200,000 - 250,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

Protingent • Washington

Hybrid
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Doist • Arlington (VA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Health insurance
Dental coverage
Vision coverage
+6
Senior ML Engineer
Senior ML Engineer

Next Ventures • New York (NY)

On-site
USD 130,000 - 160,000
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)

Doist • West Hollywood (CA)

Hybrid
USD 190,000 - 246,000
Machine Learning Engineer
Machine Learning Engineer

ACI Infotech • San Francisco (CA)

Hybrid
USD 140,000 - 160,000
Competitive salary
Health, dental, and vision insurance
401(k) with company match
+2
Staff Machine Learning Engineer
Staff Machine Learning Engineer

People In AI • New York (NY)

Hybrid
USD 180,000 - 240,000
Hybrid NYC
Machine Learning & AI Engineer
Machine Learning & AI Engineer

Atlas Search • United States

Hybrid
USD 150,000 - 210,000
Senior ML Engineer
Senior ML Engineer

Jobzhr • New York (NY)

On-site
USD 300,000 - 375,000
Direct founder collaboration
Autonomy in lean engineering team
Ownership of core systems
+1
Technical Lead - Machine Learning
Technical Lead - Machine Learning

USA Tech Recruit • San Francisco (CA)

On-site
USD 180,000 - 230,000