Mlops Support Engineer

CloudFactory

Medellín

Presencial

COP 40.000.000 - 85.000.000

Jornada completa

14 días+
Generador de candidaturas

Convierte este puesto en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

CloudFactory is seeking a proactive MLOps Support Engineer to ensure AI/ML systems remain stable in production. You will handle Tier 1 and Tier 2 incidents, triage issues, and coordinate with Tier 3 for engineering escalations.

Role focuses on maintaining model performance, pipeline health, and on-call readiness across Colombia time zones. A strong background in DevOps, SQL, Python, and cloud platforms is essential.

Formación

  • Experience in operations, DevOps, SRE, or platform support roles.
  • Strong troubleshooting skills in production environments.
  • Proficiency in SQL and scripting (Python, Bash) for ML workflows.
  • Familiarity with Cloud-hosted systems (AWS, GCP, Azure).
  • Knowledge of MLOps tooling and data platforms.

Responsabilidades

  • Provide Tier 1 / Tier 2 operational support for AI/ML solutions.
  • Identify failed jobs, degraded pipelines, or performance issues.
  • Triage incidents and coordinate escalation to Tier 3 Engineering.
  • Monitor data pipeline health, model execution, and performance metrics.
  • Work with Engineering to improve observability and runbooks.

Conocimientos

Operations experience
DevOps
SRE
Python
SQL
Cloud platforms
Kubernetes

Educación

Computer science background

Herramientas

Grafana
New Relic
Databricks
MLFlow

Descripción del empleo

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale.

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

About the role:

The MLOps Support Engineer is an operations-first role, focused on ensuring AI/ML systems remain stable, observable, and supportable in production environments. This is not a data science or feature development role.

The primary objective is to maintain continuous performance of ML models and associated pipelines with minimal disruption to both internal and client-facing services. You will provide Tier 1 and Tier 2 support, escalating to Tier 3 Engineering as needed.

What you’ll do:
  • Provide Tier 1 / Tier 2 operational support for AI/ML solutions.
  • Identify failed jobs, degraded pipelines, or performance anomalies.
  • Triage incidents, investigate issues, and coordinate escalation to Tier 3 Engineering.
  • Participate in on-call rotas once established.
  • Validate that pipelines and jobs complete successfully.
  • Monitor data pipeline health, model execution, and basic performance metrics.
  • Identify operational issues before they impact customers
  • Respond or alert customers when there has been an outage or issue with one of their models.
  • Support incident management, rollback, and recovery activities.
  • Use and maintain runbooks and operational documentation.
  • Work with Engineering to improve supportability and observability.
  • Contribute to knowledge sharing to reduce single points of failure.
  • Work within defined SLAs and support processes as the service matures
  • Build quarterly business reviews to provide updates on the health of the ML Models.
  • Evaluate champion/challenger models to see if a new model should be promoted.
  • Monitor for model drift and performance degradation, while validating that updates (new champion models or added data) do not introduce bias.
Requirements
  • Experience in operations, DevOps, SRE, or platform support roles.
  • Strong troubleshooting skills in production environments.
  • Proficiency in SQL and scripting (Python, Bash) for developing and automating ML workflows.
  • Familiarity with Cloud-hosted systems (AWS, GCP, Azure) for cloud-based ML services.
  • Git: Solid understanding of version control, particularly in collaborative development environments.
  • Comfortable working from runbooks and structured processes.
  • Exposure to AI/ML systems in production.
  • Familiarity with monitoring and observability tools (Grafana, PowerBI, New Relic).
  • Knowledge of MLOps tooling and data platforms (ML FLow, Databricks)
  • Experience supporting customer-facing platforms.
  • Knowledge of containerization (Kubernetes) is a plus.
  • Experience of LLM Prompt Engineering and troubleshooting
  • Early career in MLOps or ML Engineering.
  • Someone who is eager to learn about complex predictive models.
  • Background in computer science, informatics, or related fields
  • Passion for Machine Learning and AI: An eager learner who is excited about working with cutting-edge ML technologies and is passionate about optimizing and maintaining ML models in production environments.
  • Early Career in MLOps or ML Engineering: Ideally, Junior ML Engineer with a strong desire to grow in the field of MLOps and AI operations.
  • A Collaborative Mindset: You thrive in a team setting and are ready to contribute to model improvement, A/B testing, and iterative development.
  • Attention to Detail: A focus on model performance, bias prevention, and ensuring optimal model behavior as new data and models are introduced.
Additional information:

Nepal

  • This role provides MLOps coverage from 07:45 – 16:45* NPT for US-based customers.You will be required to work on a shift rota to cover 8 hour time blocks during this time period and potentially outside of them if a model has issues.
  • Rotational On-Call work will also be required.

Colombia

  • This role provides MLOps coverage from 9am to 9pm* Colombia. You will be required to work on a shift rota to cover 8 hour time blocks during this time period and potentially outside of them if a model has issues.
  • Rotational On-Call work will also be required.

*note that these hours are subject to change upon review.

At CloudFactory, we believe that work should be more than just a job, it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

MLOps Support Engineer: Reliability & Incident Response
MLOps Support Engineer: Reliability & Incident Response

INGEPSY • Medellín

Presencial
COP 40.000.000 - 85.000.000
Machine Learning Engineer
Machine Learning Engineer

Capgemini Engineering • Colombia

Presencial
COP 182.501.734 - 255.502.427
Stable Employment
Learning & Development
Language Training
+3
Machine Learning Engineer
Machine Learning Engineer

Capgemini Engineering • Colombia

Presencial
COP 218.762.533 - 328.143.799
Stable Employment
Learning & Development
Language Training
+3
SR MLOps Developer (Python), Colombia
SR MLOps Developer (Python), Colombia

CI&T • Norte

Presencial
COP 30.000 - 70.000
Maternity and Parental leaves
Mobile services subsidy
Sick pay
+4
ML Tech Lead (GenAI)
ML Tech Lead (GenAI)

Provectus • Cali

Presencial
COP 306.713.184 - 383.391.481
Mlops Technical Manager
Mlops Technical Manager

The Parser, Llc • Antioquia

Híbrido
COP 313.886.000 - 481.292.000
Growth opportunities
Multicultural team
Hybrid work model
+1
ML Ops Engineer ID38029 – $3,000 Sign-On Bonus
ML Ops Engineer ID38029 – $3,000 Sign-On Bonus

AgileEngine • Perímetro Urbano Cúcuta

Presencial
COP 283.918.069 - 365.037.517
Professional growth opportunities
Competitive USD-based compensation
Flexible working hours
+1
ML Solutions Architect (GenAI)
ML Solutions Architect (GenAI)

Provectus • Cali

Presencial
COP 120.000.000 - 180.000.000
ML Tech Lead (wih GenAI)
ML Tech Lead (wih GenAI)

Provectus • Medellín

Presencial
COP 301.625.004 - 452.437.507
Sr Machine Learning Engineer (Remote - LATAM Based)
Sr Machine Learning Engineer (Remote - LATAM Based)

Arionkoder • Bogotá

A distancia
COP 374.895.000 - 562.342.000
Competitive USD salary
20 business days of vacation
Family Leave
+2