Senior AI Engineer – Live Operations

Jobtailor

New Jersey

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior ML engineer to advance AI model deployment, monitoring, and safety in production. You will build automated evaluation loops, real-time guardrails, and scalable backend infra across cloud and on-prem environments.

Ideal candidates have 4+ years deploying large-scale ML/DL systems, strong Python skills, and hands-on experience with Airflow, Kubeflow, MLflow, Kafka, and Spark. Join a team pushing safe, reliable Generative AI at scale.

Qualifications

  • Bachelor's degree or four+ years of work experience.
  • Four+ years of relevant experience demonstrated in practice or military.
  • Hands-on experience deploying and monitoring ML/DL models in large-scale production.
  • Experience with MLOps frameworks and automated pipelines (Airflow, Kubeflow, MLflow).

Responsibilities

  • Apply advanced expertise to maintain AI model stability, safety, and efficiency.
  • Monitor production model behavior to detect data drift and degradation.
  • Implement automated evaluation loops using user feedback.
  • Deploy real-time safety filters and guardrails to prevent hallucinations.
  • Develop CI/CD pipelines to retrain, validate, and release models.
  • Resolve production incidents for live models as senior escalation contact.
  • Collaborate with platform architects and data science teams to scale backend infra.

Skills

Machine Learning Deployment
MLOps Frameworks
Python Programming
Cloud Platforms
Safety Guardrails

Education

Bachelor's degree or equivalent

Tools

Airflow
Kubeflow
MLflow
Kafka
Spark

Job description

  • Apply advanced technical expertise in data science, system orchestration, and software engineering to maintain the stability, safety, and efficiency of deployed AI models
  • Monitor production model behavior to identify data drift and model degradation, ensuring high predictive accuracy over time
  • Implement automated continuous evaluation loops that capture user feedback to dynamically assess performance
  • Deploy real-time safety filters, system guardrails, and input/output moderators to mitigate model hallucinations
  • Integrate robust fallback systems to preserve end-user experience
  • Develop automated continuous integration and deployment pipelines to retrain, validate, and release updated models
  • Resolve production model incidents as a senior escalation contact, diagnosing issues with live models
  • Collaborate with platform architects and core data science teams to scale backend architectures across cloud, on-premises, and edge network infrastructures

Requirements

  • Bachelor's degree or four or more years of work experience
  • Four or more years of relevant experience required, demonstrated through work experience and/or military experience
  • Extensive hands-on experience deploying, monitoring, and debugging complex machine learning or deep learning models in large-scale production environments
  • Practical familiarity with MLOps frameworks and automated pipeline tools, including Airflow, Kubeflow, or MLflow
  • Advanced programming expertise in Python, Java, or C++, with a strong grasp of data structures and cloud systems such as AWS, Google Cloud, or Azure
  • Deep experience developing safety guardrail middleware, system monitoring dashboards, and telemetry systems for Generative AI applications
  • Knowledge of distributed compute architectures, database design, and real-time streaming technologies such as Kafka or Spark.

Core Competencies

Demonstrates advanced technical expertise in data science, system orchestration, and software engineering, with a focus on deploying and monitoring machine learning models in production environments. Proficient in implementing MLOps frameworks and developing safety guardrails for AI applications.

Highest-signal resume keywords

  • Machine Learning Model Deployment
  • MLOps Frameworks
  • Python Programming
  • Cloud Systems (AWS, Google Cloud, Azure)
  • Safety Guardrail Development

ATS Optimization Keywords

Hard Skills

  • Data Science
  • Software Engineering
  • Model Monitoring
  • Automated Pipeline Development
  • Debugging Complex Models
  • Data Structures
  • Distributed Compute Architectures
  • Database Design
  • Real-time Streaming Technologies
  • Telemetry Systems

Industry Keywords

  • AI Models
  • Data Drift
  • Model Degradation
  • Continuous Integration
  • Continuous Deployment

Tools & Technologies

  • Airflow
  • Kubeflow
  • MLflow
  • Kafka
  • Spark
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Associate Director, AI Engineering – Live Operations
Associate Director, AI Engineering – Live Operations

Jobtailor • New Jersey

On-site
USD 180,000 - 240,000
Staff Machine Learning Engineer, Document & Vision Intelligence
Staff Machine Learning Engineer, Document & Vision Intelligence

Jobtailor • California (MO)

On-site
USD 150,000 - 230,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • United States

On-site
USD 180,000 - 240,000
AI Engineering Technical Lead
AI Engineering Technical Lead

Jobtailor • Burbank (CA)

On-site
USD 180,000 - 240,000
Senior MLOps Architect: Scalable Cloud ML Platforms
Senior MLOps Architect: Scalable Cloud ML Platforms

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 210,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • New York (NY)

On-site
USD 180,000 - 260,000
Distinguished Data Scientist
Distinguished Data Scientist

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Senior Data Engineer – AI Infrastructure Integration, High Performance Compute
Senior Data Engineer – AI Infrastructure Integration, High Performance Compute

Jobtailor • New Jersey

On-site
USD 180,000 - 240,000
Founding AI Engineer – Senior/Staff
Founding AI Engineer – Senior/Staff

Jobtailor • New York (NY)

On-site
USD 120,000 - 170,000
Director of Machine Learning
Director of Machine Learning

Jobtailor • Washington

On-site
USD 180,000 - 260,000