Principal Machine Learning Engineer

blazetalent

New York (NY)

Hybrid

USD 200,000 - 250,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Equity
Hybrid work

Job summary

blazetalent in New York, NY is seeking a Principal ML Engineer to own end-to-end ML infrastructure for production models. You will build training pipelines, evaluation tooling, model serving, and release processes on a small, senior team in a hybrid work setting.

The role emphasizes ownership, scalability, cost-aware deployments, and collaboration with research counterparts to push industrial-grade ML in a regulated or high-stakes domain.

Qualifications

  • 8+ years of software engineering experience, with 4+ years building ML/LLM infra in production.
  • Hands-on depth with modern LLM stack, distributed training, and scalable inference.
  • Experience building eval harnesses, regression gates, and dataset pipelines.
  • Proven track record owning model serving under latency, reliability, and cost constraints.
  • Strong Python, containers, CI/CD, cloud infra, and observability.

Responsibilities

  • Build and own training pipelines: data prep, reproducible fine-tuning, experiment tracking, and release automation.
  • Build evaluation infrastructure: automated eval runs, regression gates, dashboards, dataset versioning.
  • Own model serving in production: low-latency inference, batching, optimization, autoscaling, cost management.
  • Ship model updates safely with versioning, canarying, rollback, and drift monitoring.
  • Build repeatable workflows for adapting models to new domains and customer needs.
  • Convert expert labels and reviewer feedback into clean training and evaluation datasets.
  • Set the technical bar for ML infrastructure as the team grows.

Skills

Python
PyTorch
LLM stack
CI/CD
Cloud infra

Education

BS/MS in CS or related

Tools

vLLM
TensorRT-LLM
LoRA/SFT

Job description

Principal ML Engineer

New York, NY (Hybrid)

About the Company

We're building AI-native enforcement infrastructure for enterprise communication — technology that catches and fixes compliance issues in real time, before an AI-generated message ever reaches a customer or counterparty, across every channel where AI represents the business. Most existing tools only flag problems after the fact, once the risk is already out the door; we intervene before send. This is a new category, and we're the ones defining it.

We're backed by top-tier venture capital and built by a team with backgrounds at major tech and financial firms, led by a founder who has built and scaled AI companies before.

The Role

Specialized language models sit at the core of our enforcement layer, making real-time decisions about whether a communication is safe to send. These models need to be accurate, fast, and dependable, since they operate directly in the path of live traffic.

Our research team owns the underlying science — model behavior, training objectives, data strategy, and quality standards. You'll own the systems that turn that science into a reliable, production-grade product: the pipelines that train models reproducibly, the evaluation infrastructure that proves they work, and the serving stack that runs them at scale.

This is a hands-on, principal-level individual contributor role on a small, senior team. It's a systems and infrastructure role, not a research role — ideal for someone who loves making ML industrial-grade.

What You'll Do
  • Build and own training pipelines: data prep, reproducible fine-tuning runs, experiment tracking, and release automation
  • Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning
  • Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management
  • Ship model updates safely with versioning, canarying, rollback, and drift monitoring
  • Build repeatable workflows for adapting models to new domains and customer needs
  • Convert expert labels and reviewer feedback into clean training and evaluation datasets
  • Set the technical bar for ML infrastructure as the team grows
What We're Looking For
  • 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
  • Hands-on depth with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM
  • Experience building eval harnesses, regression gates, or dataset pipelines, with solid understanding of precision, recall, and calibration
  • Proven track record owning model serving under real latency, reliability, and cost constraints — not just in notebooks
  • Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
  • Comfort with high ownership on a small team: scoping your own work, shipping weekly, and making pragmatic build-vs-buy calls
  • Enjoyment of close collaboration with a research counterpart, with clear interfaces and no turf wars
Nice to Have
  • Experience productionizing small or specialized language models
  • Experience with structured-output serving or constrained decoding in production
  • Background in a regulated or high-stakes domain such as fintech, healthcare, legal, or trust and safety
  • Experience deploying models into customer-controlled environments
Compensation & Benefits
  • $200,000–$250,000 base salary, depending on experience
  • Performance bonus and meaningful early-stage equity
  • Health, dental, and vision coverage
  • Hybrid work from a New York office
Compensation

The base pay range for this role is $200,000 – $250,000 per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Fuel Talent • Seattle (WA)

Hybrid
USD 160,000 - 230,000
Senior Machine Learning Engineer (LLMs)
Senior Machine Learning Engineer (LLMs)

Albiware Inc. • Chicago (IL)

On-site
USD 140,000 - 210,000
Competitive salary
Generous PTO
Medical, dental, and vision coverage
+2
Principal Machine Learning Engineer I
Principal Machine Learning Engineer I

RELX • Raleigh (NC)

On-site
USD 136,000 - 253,000
Country-specific benefits
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Syndesus, Inc. • Austin (TX)

On-site
USD 140,000 - 160,000
Health, vision, dental insurance (100%
401k retirement plans
Disability insurance
+1
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

GoTo Meeting • Mountain View (CA)

On-site
USD 150,000 - 230,000
Health, dental, and vision care for you and your family
Top-tier 401(K) plan with company matching
Paid time off and paid holidays
+2
Machine Learning Engineer
Machine Learning Engineer

Golden Gate Recruiting • San Francisco (CA)

On-site
USD 155,000 - 230,000
Member of ML Technical Staff
Member of ML Technical Staff

Pragmatike • San Francisco (CA)

On-site
USD 200,000 - 350,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Syndesus, Inc. • Austin (TX)

On-site
USD 100,000 - 140,000
100% employer-paid health, vision, and dental insurance
Retirement plans (401(k))
Disability insurance
+1
Machine Learning Engineer (LLM)
Machine Learning Engineer (LLM)

DeepRec.ai • Boston (MA)

On-site
USD 170,000 - 200,000
Principal Machine Learning Engineer I
Principal Machine Learning Engineer I

LexisNexis • Raleigh (NC)

On-site
USD 136,000 - 253,000