Machine Learning Operations Engineer

Mosai

Nashville (TN)

On-site

USD 140,000 - 170,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

null

Job summary

Mosai is seeking an experienced Machine Learning Ops (MLOps) Engineer to architect, develop, and maintain the full lifecycle of data and model pipelines powering training, inference, evaluation, and analytics workflows. The role emphasizes reliability, scalability, and observability of ML systems in production, with focus on auditing and consolidating existing pipelines.

The ideal candidate will be proficient in Python and Jupyter, deeply familiar with Snowflake and cloud environments (Azure and

Qualifications

  • Bachelor’s degree in Computer Science, Engineering or equivalent work experience.
  • 5–7 years of combined experience in Data Engineering, MLOps, Machine Learning Engineering, or related fields.
  • Experience operationalizing traditional ML models as well as LLM-based and MCP-orchestrated systems.
  • Strong working knowledge of Azure and AWS cloud platforms, including compute orchestration, networking, and security best practices.
  • Experience with CI/CD tools, Docker, infrastructure-as-code, and ML pipeline frameworks.
  • Strong ability to diagnose and resolve pipeline failures, data anomalies, and complex system issues.
  • Advanced proficiency in Python, Jupyter, and common ML/analytics frameworks.
  • Hands-on experience with Snowflake or similar cloud data warehousing enviro
  • Excellent problem-solving skills, attention to detail, and a proactive, self-directed work ethic.
  • Strong communication skills and comfort working in fast-paced, cross-functional environments.

Responsibilities

  • Design, build, and maintain scalable data pipelines supporting model training, inference, batch processing, and real-time analytics workflows.
  • Audit, refactor, and consolidate existing ML pipelines and deployment processes to eliminate technical debt, redundant workflows, and undocumented manual steps.
  • Monitor and deploy production ML pipelines to identify anomalies, performance degradations, or failures.
  • Execute rapid troubleshooting and root-cause analysis followed by timely remediation, validation, and full regression testing prior to redeployment.
  • Collaborate with Data Science, Engineering, and Product teams to operationalize ML models—including LLM-based and MCP-orchestrated systems.
  • Develop CI/CD workflows, model deployment strategies, and automated testing frameworks to support reliable, repeatable releases.
  • Implement and maintain observability tooling (logging, monitoring, alerting) to ensure high availability and traceability of ML systems.
  • Manage and optimize cloud infrastructure across Azure and AWS for compute, storage, orchestration, and security needs.
  • Create and maintain documentation, runbooks, and best practices for model operations and system maintenance.

Skills

Python
Jupyter
Cross-functional collaboration
Problem-solving

Education

Bachelor's degree in CS or related field

Tools

Docker
CI/CD tools
Snowflake
Azure
AWS
Infrastructure as Code

Job description

About Mosai

Mosai™ is the intelligent care coordination platform that brings together the fragmented pieces of healthcare into a clear, connected picture. Like a mosaic, our platform unites data, people, and processes so providers can make better decisions, coordinate care in real time, and deliver improved outcomes. With Mosai, home-based care organizations can thrive in value-based care while giving every patient the right care, in the right place, at the right time.

Learn more at https://www.mosai.com/

Position Summary

We are seeking an experienced Machine Learning Ops (MLOps) Engineer to architect, develop, and maintain the full lifecycle of data and model pipelines that power training, inference, evaluation, and analytics workflows. This role is responsible for ensuring the reliability, scalability, and observability of all machine learning systems in production, including traditional ML models and modern LLM-based/MCP-orchestrated architectures. A key focus of this role in the near term is auditing and consolidating our existing pipelines and deployment processes. The ideal candidate is highly skilled in Python, Jupyter, Snowflake, and both Azure and AWS cloud environments, and thrives in environments requiring continuous monitoring, rapid issue diagnosis, and rigorous validation before deployment.

Job Duties
  • Design, build, and maintain scalable data pipelines supporting model training, inference, batch processing, and real-time analytics workflows.
  • Audit, refactor, and consolidate existing ML pipelines and deployment processes to eliminate technical debt, redundant workflows, and undocumented manual steps.
  • Audit, refactor, and consolidate existing ML pipelines and deployment processes to eliminate technical debt, redundant workflows, and undocumented manual steps.
  • Monitor and deploy and deploy production ML pipelines to identify anomalies, performance degradations, or failures related to data quality, logic defects, or infrastructure issues.
  • Execute rapid troubleshooting and root-cause analysis followed by timely remediation, validation, and full regression testing prior to redeployment.
  • Collaborate with Data Science, Engineering, and Product teams to operationalize machine learning models—including LLM-based and MCP-orchestrated systems—ensuring seamless integration into production environments.
  • Develop CI/CD workflows, model deployment strategies, and automated testing frameworks to support reliable, repeatable releases.
  • Implement and maintain observability tooling (logging, monitoring, alerting) to ensure high availability and traceability of ML systems.
  • Manage and optimize cloud infrastructure across Azure and AWS for compute, storage, orchestration, and security needs.
  • Create and maintain documentation, runbooks, and best practices for model operations and system maintenance.
  • Perform all other job-related duties as assigned.
Minimum Requirements
  • Bachelor’s Degree in Computer Science, Engineering or equivalent work experience.
  • 5–7 years of combined experience in Data Engineering, MLOps, Machine Learning Engineering, or related fields.
  • Demonstrated experience operationalizing traditional ML models as well as LLM-based and MCP-orchestrated systems.
  • Strong working knowledge of both Azure and AWS cloud platforms, including compute orchestration, networking, and security best practices.
  • Experience with CI/CD tools, containerization (Docker), infrastructure-as-code, and ML pipeline frameworks.
  • Strong ability to diagnose and resolve pipeline failures, data anomalies, and complex system issues.
  • Advanced proficiency in Python, Jupyter, and common ML/analytics frameworks.
  • Hands-on experience with Snowflake or similar cloud data warehousing enviro
  • Excellent problem-solving skills, attention to detail, and a proactive, self-directed work ethic.
  • Strong communication skills and comfort working in fast-paced, cross-functional environments.
Work Environment
  • This role is preferred to be based in Nashville or Jacksonville, near Mosai’s offices.
Physical Demands of Our Work Environment
  • This position uses a computer and other office equipment as needed to perform duties. The in-office noise level in the work environment is typical of that of an office. Frequent interruptions may be encountered throughout the workday.
  • The employee is required to either stand or sit, talk and hear frequently required to use repetitive keying or hand motions.
  • The physical demands are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

Mosai is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, sexual orientation, gender identity, veteran status, and disability, or other legally protected status, If you are unable to submit an application because of a incompatible assistive technology or disability, please contact us at careers@mosai.com. We will make every effort to respond to your request for disability assistance as soon as possible.

Mosai is an E-verify employer. Your eligibility to work in the United States will be verified through the E-verify system if you apply and are selected for a position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Operations Engineer
Machine Learning Operations Engineer

Mosai • Jacksonville (FL)

On-site
USD 140,000 - 190,000
Relocation assistance
MLOps Engineer: Build Reliable AI Pipelines & Deployments
MLOps Engineer: Build Reliable AI Pipelines & Deployments

Mosai • Jacksonville (FL)

On-site
USD 140,000 - 190,000
Relocation assistance
Machine Learning ML Data Engineer
Machine Learning ML Data Engineer

Moser Consulting • Indianapolis (IN)

On-site
USD 120,000 - 155,000
Training Opportunities
Fully Invested 401K Plan
PPO and HDHP Medical Plans
+4
MLOps Engineer: Deployments, Pipelines & Observability
MLOps Engineer: Deployments, Pipelines & Observability

Mosai • Nashville (TN)

On-site
USD 140,000 - 170,000
null
Machine Learning Operations Engineer
Machine Learning Operations Engineer

Speria • Atlanta (GA)

On-site
USD 120,000 - 180,000
ML Operations Engineer
ML Operations Engineer

NextGen Healthcare • Georgia

On-site
USD 80,000 - 120,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
ML Ops Architect
ML Ops Architect

Tiger Analytics • Dallas (TX)

On-site
USD 120,000 - 150,000
Career development opportunities
Collaborative work environment
Challenging projects
Sales & Customer Success Trainer
Sales & Customer Success Trainer

Jobless • Nashville (TN), Northern (KY)

Hybrid
USD 80,000 - 120,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites