MLOps Engineer

Fractal Analytics Inc

New York (NY)

On-site

USD 120,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, vision, life insurance
401(k) eligibility after 30 days
11 paid holidays
12 weeks parental leave
Free-time PTO

Job summary

Fractal Analytics Inc. in New York is seeking an experienced MLOps Engineer to operationalize machine learning solutions from data preprocessing to production deployment.

The role emphasizes building production-grade services, scalable pipelines, and robust observability across batch and real-time paths. The candidate will own data preprocessing, feature engineering, and model deployment pipelines using Databricks, PySpark, and MLflow, ensuring parity between training and serving environments.

Qualifications

  • Deep hands-on Python for data engineering and application development.
  • 2nd item placeholder

Responsibilities

  • Design and build FastAPI services exposing models with contracts, auth, input validation, and observability.

Skills

Python data engineering
FastAPI services
Messaging systems (Kafka/RabbitMQ/SQS)
Docker
Kubernetes
Databricks / MLFlow
CI/CD (Jenkins/GitHub Actions)
Spark / PySpark
ML tooling & libraries (scikit-learn,X

Tools

Docker
Kubernetes
Databricks
MLFlow
Airflow / Prefect
Spark
Jenkins / GitHub Actions

Job description

MLOps Engineer – Consultant

Senior consulting role to operationalize machine learning solutions in purchase and underwriting. Hands‑on engineering role focused on writing production code—designing services, hardening pipelines, and ensuring ML production solutions are deployed, monitored, and consistent across batch and real‑time paths.

Responsibilities
Model Serving

Design and build FastAPI services that expose models to downstream applications, including request/response contracts, authentication and authorization, input validation, error semantics, and structured logging, tracing, and metrics.

Implement queue‑based asynchronous serving for higher‑latency or higher‑throughput workloads, covering producers and consumers, worker concurrency, retries, back‑off, dead‑letter handling, back‑pressure, idempotency, and end‑to‑end traceability of a request across the pipeline.

Containerization and Deployment

Containerize services with Docker and deploy them to enable reliable scaling, rollout, and rollback.

Data Preprocessing, Feature Engineering, and Pipelines

Own the data preprocessing, transformation, and feature engineering code between raw sources and the model, refactoring notebook or script‑style logic into modular, tested, and reusable components.

Build reproducible training and batch inference pipelines on Databricks and PySpark from raw sources through curated feature and training datasets.

Manage model artifacts, versions, and promotion across environments to ensure reproducibility.

Feature Parity Across ML Lifecycle

Guarantee that the feature values a model sees at training time match those at batch scoring and real‑time serving through consistency in definitions, transformations, and edge‑case handling.

Establish parity checks and reconciliation between training data, batch outputs, and real‑time predictions as part of the pipeline.

Reliability and Observability

Implement monitoring for model performance, prediction drift, data quality, and pipeline health with actionable alerts routed to the relevant owners.

Diagnose production incidents in pipelines and services, identify root causes, and drive fixes through to closure.

Engineering Practices and Documentation

Apply strong software engineering fundamentals—testing, code review, CI/CD, semantic versioning, and dependency hygiene—to ML code.

Build and maintain shared libraries, utilities, and repository patterns usable by other ML use cases.

Create clear documentation of pipelines, frameworks, and operational runbooks to enable smooth ownership transfer.

Collaboration and Communication

Work closely with Data Science, Data Engineering, business partners, and IT teams to align on requirements, handoffs, and production readiness.

Produce clear documentation of pipelines, frameworks, and operational runbooks so ownership can transition seamlessly to internal teams.

Qualifications
  • Deep hands‑on Python for data engineering and application development; comfortable with SQL, PySpark, and shell scripting.
  • Production experience building services with FastAPI or a comparable Python web framework, including auth, validation, error handling, and observability.
  • Experience building queue‑based asynchronous processing systems using Kafka, RabbitMQ, SQS, Redis Streams, Celery, or equivalent, with operational concerns such as retries, idempotency, back‑pressure, and dead‑letter queues.
  • Strong Docker and general containerization skills; comfortable with Kubernetes concepts.
  • Hands‑on Databricks experience including MLFlow and distributed compute in Spark.
  • Working experience with common ML libraries (scikit‑learn, XGBoost, PyTorch, or similar) sufficient to partner with data scientists.
  • Strong grasp of the end‑to‑end ML lifecycle and experience building or migrating feature engineering code with training/batch/realtime parity focus.
  • Comfort reading and refactoring batch ML or data pipeline code while preserving intent and edge cases.
  • Experience with CI/CD tools (Jenkins, GitHub Actions, or equivalent), version control workflows, and orchestration (Airflow, Prefect, or equivalent).
  • Excellent written and verbal communication skills; able to drive alignment without middle manager mediation.
Nice to Have
  • Prior experience in Group Insurance Domain or Life Insurance Underwriting Domain.
  • Experience operationalizing LLM‑based systems— inference serving, evaluation, cost, and latency controls.
Compensation

Salary range: $120,000 to $140,000 yearly. Potential for discretionary bonus based on performance.

Benefits
  • Health, dental, vision, life insurance, and disability plans for full‑time and hourly employees working over 30 hours per week, effective from day one.
  • Eligibility for 401(k) plan after 30 days of employment.
  • 11 paid holidays and 12 weeks of parental leave.
  • Free‑time PTO policy allowing flexible sick time or vacation.
Equal Employment Opportunity Statement

Fractal provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer
Staff ML Engineer

Gainbridge • Zionsville (IN)

On-site
USD 190,000 - 215,000
Comprehensive health, dental, and vision insurance plans
401(k) plan with company matching contributions
FOUNDING MACHINE LEARNING ENGINEER
FOUNDING MACHINE LEARNING ENGINEER

Shepherd Insurance Agency • San Francisco (CA)

On-site
USD 180,000 - 220,000
Lead Software Platform Engineer, MLOps
Lead Software Platform Engineer, MLOps

TetraScience • Cambridge (MA)

On-site
USD 200,000 - 270,000
Employer-paid benefits
Unlimited PTO
401K
+3
Lead Software Platform Engineer, MLOps
Lead Software Platform Engineer, MLOps

TetraScience, Inc. • Cambridge (MA)

Hybrid
USD 200,000 - 270,000
Employer-paid benefits
Unlimited PTO
401K
+3
Senior ML Ops Engineer
Senior ML Ops Engineer

United States Digital Space LLC • New York (NY)

On-site
USD 180,000 - 240,000
Equity
Fully paid health coverage
Dental and vision
+7
MLOps Lead
MLOps Lead

Fractal • New York (NY)

On-site
USD 140,000 - 205,000
401(k) Plan
Paid holidays
Parental Leave
Senior Software Engineer AI
Senior Software Engineer AI

BlackLine Systems Inc (U.S.) • Pleasanton (TX)

On-site
USD 145,000 - 182,000
Senior ML Software Engineer
Senior ML Software Engineer

Greater Giving, Inc. • Atlanta (GA)

On-site
USD 100,000 - 130,000
Medical coverage and wellbeing support
Base salary plus bonuses
Paid time off and holidays
Senior Machine Learning (ML) Engineer (AI Insurtech)
Senior Machine Learning (ML) Engineer (AI Insurtech)

Cerebras • New York (NY)

On-site
USD 190,000 - 225,000
Annual bonus plan
Company equity (RSUs)
Medical, dental, vision insurance
+2
Machine Learning Engineer
Machine Learning Engineer

CRC Group • Charlotte (NC)

On-site
USD 140,000 - 180,000