Hybrid in the New York City Metro area (3 days onsite). In this Principal Machine Learning Engineering role, you will own the ML infrastructure behind real-time compliance enforcement systems. Expect a position centered on building the pipelines, evaluation workflow, and production serving needed to train, measure, and deploy models reliably, with a focus on latency, reliability, and cost.
Pay range: USD $200,000 - $250,000 per year.
What you’ll do
- Build and own training pipelines including data preparation, reproducible fine-tuning runs, experiment tracking, and release automation.
- Develop evaluation infrastructure with automated eval runs, regression gates, dashboards, and dataset versioning.
- Own production model serving for low-latency inference, including batching, optimization, autoscaling, and cost management.
- Ship model updates safely using versioning, canarying, rollback, and drift monitoring.
- Create repeatable workflows to adapt models to new domains and changing customer needs.
- Convert expert labels and reviewer feedback into clean training and evaluation datasets.
- Help raise the team’s engineering bar for ML infrastructure as the organization grows.
What you’ll bring
- 8+ years of software engineering experience, including 4+ years building ML or LLM infrastructure for production.
- Hands-on experience with the modern LLM stack: PyTorch, distributed training, and fine-tuning at scale (e.g., LoRA, SFT) using inference engines such as vLLM or TensorRT-LLM.
- Experience building eval harnesses, regression gates, or dataset pipelines, with strong understanding of precision, recall, and calibration.
- Proven ownership of production model serving with real latency, reliability, and cost constraints.
- Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability.
- Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions.
- Experience collaborating tightly with research partners and defining clear interfaces.
Technologies you’ll work with
- PyTorch, LoRA, SFT, vLLM, TensorRT-LLM
- Python, containers, CI/CD, cloud infrastructure, observability
Additional qualifications that may help
- Experience productionizing small or specialized language models.
- Experience with structured-output serving or constrained decoding in production.
- Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety).
- Experience deploying models into customer-controlled environments.
Work location: Hybrid remote in New York, NY 10001 (3 days onsite).