Senior AI/ML Operations Engineer

Jobgether

United States

On-site

USD 150,000 - 230,000

Full time

3 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Work-from-anywhere flexibility
Unlimited paid time off
Comprehensive health coverage
Equity grants for all employees
One-time home office setup allowance
Monthly cell phone allowance

Job summary

Partner Company in the United States seeks a Senior AI/ML Operations Engineer to own the infrastructure, pipelines, and reliability behind production AI systems.

You will work across classical ML, generative AI, and agentic applications, building scalable deployment patterns and governance practices while collaborating with cross-functional teams in a healthcare data environment.

Qualifications

  • 5+ years of professional experience in AI/ML engineering, MLOps, or related discipline.
  • Experience with Databricks and Snowflake governance and access-control frameworks.

Responsibilities

  • Deploy, promote, and maintain ML and GenAI models and pipelines across environments using CI/CD infrastructure.
  • Develop reusable deployment patterns and tooling to reduce time-to-launch for AI/ML use cases.
  • Build and maintain data pipelines for classical ML and GenAI workloads.
  • Operate within Databricks and Snowflake governance frameworks to support secure data, code, and models promotion.
  • Diagnose and resolve production issues across data pipelines, infrastructure, and model-serving systems.
  • Automate and monitor production ML inference and feature-engineering workflows with alerting and reliability efforts.
  • Manage model lifecycles through MLflow and Unity Catalog, including experiment tracking and versioning.

Skills

Python
SQL
Databricks
Snowflake
MLflow
CI/CD
RAG
Vector search
Agentic systems
Model-serving

Tools

Unity Catalog
Kubernetes

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI/ML Operations Engineer based in United States.


This senior individual contributor role is responsible for the infrastructure, pipelines, and operational reliability behind production AI and machine learning systems.


You will work across classical ML, generative AI, and agentic applications, helping move solutions from prototypes into dependable production environments.


The role combines platform engineering, MLOps, model lifecycle management, data pipelines, and AI infrastructure.


You will build scalable deployment patterns, operate RAG and agentic systems, and strengthen observability, evaluation, and governance practices.


Working closely with engineering, data science, security, DevOps, and business teams, you will solve complex operational challenges and enable new AI/ML use cases.


The position offers significant hands-on ownership in a technology-driven environment where emerging AI capabilities are applied to real-world healthcare data challenges.


You will also mentor junior engineers and contribute to customer-specific implementations when needed.


Accountabilities


  • Deploy, promote, and maintain ML and GenAI models, pipelines, and code across environments using established CI/CD infrastructure.

  • Develop reusable deployment patterns, tooling, and operational practices that reduce the effort required to launch new AI/ML use cases and client-specific solutions.

  • Build and maintain data pipelines supporting classical ML and GenAI workloads, covering ingestion, feature engineering, processing and serving.

  • Operate within Databricks and Snowflake governance frameworks, including access controls and environment boundaries, to support secure and compliant promotion of data, code, and models.

  • Independently diagnose and resolve production issues across data pipelines, infrastructure, model-serving systems, and related AI/ML components.

  • Automate and monitor production ML inference and feature-engineering workflows, including alerting, incident response, and operational reliability.

  • Manage model lifecycles through platforms such as MLflow and Unity Catalog, including experiment tracking, registration, versioning, and controlled promotion across environments.

  • Build and maintain infrastructure for retrieval-augmented generation (RAG), including vector search indexing and retrieval pipelines.

  • Deploy, host, and maintain MCP servers and tool integrations supporting agentic applications.

  • Build evaluation infrastructure for AI systems and collaborate with AI engineering teams on evaluation methodologies and quality standards.

  • Implement observability for agents and models through logging, tracing, monitoring, and analysis of production behavior.

  • Partner with business stakeholders to understand data and feature requirements and coordinate infrastructure needs with Software Engineering, Data Engineering, Data Science, Security, and DevOps teams.

  • Mentor junior engineers on platform practices, operational standards and reliable AI/ML engineering.

  • Contribute to customer-specific implementations when required, including semantic layer configuration and domain-specific analytics initiatives.


Requirements


  • 5+ years of professional experience in AI/ML engineering, MLOps, or a closely related discipline.

  • Extensive hands-on AI/ML experience with meaningful depth in either generative AI/agentic systems or classical machine learning, alongside working exposure to the other area.

  • Experience with AI/GenAI technologies such as agentic frameworks, RAG systems, vector search, MCP or comparable tool-integration protocols, model-serving or gateway layers, and AI evaluation design.

  • Alternatively, strong classical ML experience covering common algorithm families such as XGBoost, gradient boosting, random forests and neural networks, along with feature engineering, training pipelines and production deployment.

  • Deep hands-on experience with Databricks and/or Snowflake, including AI/ML pipeline development, governance and access-control frameworks such as Unity Catalog, and integration with CI/CD infrastructure.

  • Strong experience with model lifecycle and registry platforms such as MLflow, including experiment tracking, model registration, versioning and promotion across environments.

  • Strong proficiency in Python and SQL.

  • Demonstrated ability to independently investigate, troubleshoot and resolve production infrastructure and operational issues.

  • Excellent communication skills and the ability to collaborate directly with both technical and non-technical stakeholders.

  • A proactive approach to learning new technologies, tools and frameworks and applying them effectively to real-world projects.

  • Experience in healthcare, health insurance or regulated data environments is a plus.

  • Experience building or operating multi-agent systems is a plus.

  • Experience with Mosaic AI Gateway or comparable model-serving and gateway platforms is a plus.


Benefits


  • Compensation based on experience, skills and location, including base salary plus eligibility for performance bonuses and equity grants.

  • Unlimited paid time off.

  • Work-from-anywhere flexibility.

  • Comprehensive health coverage with multiple plan options.

  • Equity grants for all employees.

  • Growth-focused environment with opportunities for professional development.

  • One-time home office setup allowance.

  • Monthly cell phone allowance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Applied AI Solutions Engineer
Senior Applied AI Solutions Engineer

Jobgether • United States

Hybrid
USD 200,000 - 350,000
Competitive compensation
Flexible working arrangements
Comprehensive benefits
Senior Manager, AI Engineering
Senior Manager, AI Engineering

Scorpion Therapeutics • San Diego (CA)

On-site
USD 150,000 - 188,000
Discretionary bonus
Equity awards eligibility
Medical/dental/vision; life/disability
Senior AI Engineer
Senior AI Engineer

Weekday (YC W21) • San Francisco (CA)

On-site
USD 200,000 - 275,000
[Job-31573] AI Engineer Master
[Job-31573] AI Engineer Master

ciandt • United States

Hybrid
USD 180,000 - 270,000
Health insurance
Dental insurance
Life insurance
+8
Principal, AI Engineer, Solution Delivery
Principal, AI Engineer, Solution Delivery

Tata Consultancy Services • New York (NY)

On-site
USD 120,000 - 130,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
401K Plan
+1
Lead Software Platform Engineer, MLOps
Lead Software Platform Engineer, MLOps

TetraScience • Cambridge (MA)

On-site
USD 200,000 - 270,000
Employer-paid benefits
Unlimited PTO
401K
+3
Lead Software Platform Engineer, MLOps
Lead Software Platform Engineer, MLOps

TetraScience, Inc. • Cambridge (MA)

Hybrid
USD 200,000 - 270,000
Employer-paid benefits
Unlimited PTO
401K
+3
Senior AI Engineer | Onsite
Senior AI Engineer | Onsite

Worky • United States

On-site
USD 40,000 - 140,000
Medical benefits
Vision benefits
Dental benefits
+4
AI/ML Ops Architect
AI/ML Ops Architect

Tata Consultancy Services • Atlanta (GA)

On-site
USD 130,000 - 150,000
Medical Coverage
Dental & Vision
Parental Leaves
+1
Senior MLOps Engineer / Databricks / AWS Bedrock / Remote
Senior MLOps Engineer / Databricks / AWS Bedrock / Remote

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 90,000 - 130,000
Medical, Dental, and Vision Insurance
401(k) Program
Paid Vacation & Holidays
+4