Get more replies from employers
Send a job-specific resume in minutes.
NLP PEOPLE is hiring a Lead Machine Learning Operations Engineer in New York. This role involves owning the operational excellence and reliability of ML systems, ensuring efficient monitoring and deployment, and leading cross-functional collaboration across teams.
The ideal candidate has over 5 years of experience in machine learning engineering or related fields and a strong track record of operating production systems. This position offers an attractive salary range of $157,000 to $235,000.
#WeAreParamount on a mission to unleash the power of content… you in?
We’ve got the brands, we’ve got the stars, we’ve got the power to achieve our mission to entertain the planet – now all we’re missing is… YOU! Becoming a part of Paramount means joining a team of passionate people who not only recognize the power of content but also enjoy a touch of fun and uniqueness. Together, we co-create moments that matter – both for our audiences and our employees – and aim to leave a positive mark on culture.
We’re hiring a Lead Machine Learning Operations Engineer to own the operational excellence, observability, reliability, and governance layer around our personalization and recommendation ML systems.
Our recommendation models retrain and deploy frequently. You will define how we detect model behavior changes, diagnose issues quickly, and prevent bad deployments from reaching customers.
This is a lead-level IC role: you’ll set technical direction and drive adoption across ML Engineering, DevOps, Platform Engineering, Data Engineering, and Product.
Sitting within ML Platform and Infrastructure, you’ll partner closely with ML engineers who own model development. You’re not expected to build infrastructure from scratch, but you’ll define what good looks like, evaluate tooling, and own the day-to-day operational layer.
Own ML production reliability strategy: define and lead the operational strategy for production ML systems, including monitoring, traceability, deployment safety, incident response, and post-deployment validation. Set the standards ML teams use to assess model health, performance, and trustworthiness in production.
Own model traceability and governance: ensure every production model has clear lineage and drive adoption of model registry and metadata tooling across ML teams.
Build end-to-end ML observability: design and implement monitoring across the full ML signal path – data arrival, feature freshness, distribution stability, candidate generation, ranking behavior, model metrics, serving latency, and SLA performance.
Define production health metrics: partner with ML, data, product, and business stakeholders to define post-deployment metrics covering model quality, system reliability, business guardrails, and degradation indicators.
Detect drift and degradation proactively: detect data drift, feature drift, model behavior changes, and silent failures before they impact customers via thresholding, alerting, anomaly detection, and release-over-release monitoring.
Lead diagnostic tooling and root-cause analysis: build dashboards, logs, and diagnostic workflows that progress quickly from “recommendations look off” to root cause, with context captured across candidates, features, scores, ranking decisions, and downstream outcomes.
Own ML deployment safety: define and operate automated gates that prevent bad models or bad data from being promoted to production. Partner with MLEs to establish validation checks, rollback criteria, canary strategies, shadow testing, and release health reviews.
Lead ML incident response: own incident response practices for ML systems, including rollback playbooks, hotfix strategies, severity definitions, tradeoff frameworks, communications, and post-mortems. Drive closure of systemic gaps after incidents rather than only resolving the immediate issue.
Partner across ML Platform, Data, and ML: partner with DevOps/Platform on infrastructure and observability needs; with Data Engineering on data quality, drift, and freshness; and with ML Engineering to embed operational requirements into development and deployment workflows.
Set standards and mentor others: act as the technical lead for ML operations, establishing reusable patterns, playbooks, and standards, and mentoring engineers on reliability, observability, and operational rigor.
Within the first 90 days, you will have assessed the current production ML operational landscape, identified the highest‑risk gaps, and established a prioritized roadmap for observability, traceability, deployment safety, and incident response.
Within six months, our production recommendation systems will have clearer model lineage, stronger monitoring coverage, better diagnostic workflows, and more reliable deployment gates.
Within one year, ML operations will be a repeatable, trusted function: teams will know what is running, why it was promoted, how it is performing, when something is wrong, and how to recover quickly.
The right candidate defines strategy, leads cross‑functional execution, and sets engineering standards while staying hands‑on enough to investigate production issues directly.
$157,000.00 – $235,000.00 (New York, California, Colorado, Washington state, and most other geographies)
Paramount is an equal opportunity employer (EOE) including disability/vet. Paramount is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ethnicity, ancestry, religion, creed, sex, national origin, sexual orientation, age, citizenship status, marital status, disability, gender identity, gender expression, and Veteran status.
If you are a qualified individual with a disability or a disabled veteran, you may request a reasonable accommodation if you are unable or limited in your ability to use or access our careers site as a result of your disability. You can request reasonable accommodations by calling 212.846.5500 or by sending an email to an appropriate contact. Only messages left for this purpose will be returned.
Paramount Streaming, a division within Paramount Global, is the home to the company’s direct‑to‑consumer services spanning free and paid in the form of Pluto TV and Paramount+.
Pluto TV is the global leader in free ad‑supported TV, delivering more than 1,400 global channels and an extensive library of streaming content, including live and original channels. Paramount+ is a digital subscription video‑on‑demand and live streaming service, combining live sports, breaking news, and a mountain of entertainment. Paramount+ features an expansive library of original series, hit shows and popular movies across every genre from world‑renowned brands and production studios, including Showtime.