Senior Machine Learning Engineer

HP

Tlaquepaque

Presencial

MXN 1.200.000 - 1.600.000

Jornada completa

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Transforma esta oferta en una entrevista: un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

HP is seeking a Senior MLOps Engineer to design, build, and operate scalable infrastructure for ML models and large language models. You will enable end-to-end deployment from experimentation to production, exposing services through secure endpoints across AWS and Databricks.

The role requires building robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and production-grade reliability.

Formación

  • Experience with cloud platforms (AWS) and ML infrastructure.
  • Strong background in MLOps, CI/CD pipelines, and model deployment.
  • Familiarity with observability, security, and governance for ML systems.

Responsabilidades

  • Design scalable MLOps platform using AWS and Databricks.
  • Build CI/CD pipelines for models, code, and artifacts.
  • Deploy secure, low-latency inference endpoints for models/LLMs.
  • Implement multi-region, autoscaling, and rollback strategies.
  • Ensure observability with dashboards, metrics, and tracing.
  • Collaborate with data scientists, ML engineers, and security teams.

Herramientas

AWS
Databricks
SageMaker
Kubernetes
Docker
APIGateway
EKS

Descripción del empleo

We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale.

We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability.

Key Responsibilities
MLOps Platform and Architecture
  • Design and implement a scalable MLOps platform using AWS and Databricks.
  • Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models.
  • Build standardized workflows that move models from development and validation into staging and production.
  • Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure.
  • Establish clear separation between development, testing, staging, and production environments.
  • Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives.
CI/CD and Model Deployment
  • Build automated CI/CD pipelines for model code, inference services, infrastructure, configuration, and model artifacts.
  • Implement automated testing across the deployment lifecycle, including:
    • Unit testing
    • Integration testing
    • Model validation
    • Data contract validation
    • API and endpoint testing
    • Security testing
    • Performance and load testing
    • Regression testing
  • Automate model packaging, containerization, versioning, approval, promotion, and deployment.
  • Support deployment strategies such as blue-green deployments, canary releases, shadow deployments, and controlled traffic shifting.
  • Implement reliable rollback and roll-forward mechanisms for application code, infrastructure, model versions, prompts, and configuration.
  • Ensure deployments are reproducible, auditable, and recoverable.
Model and LLM Serving
  • Design and operate secure, scalable, low-latency inference endpoints.
  • Deploy models using appropriate services and patterns across AWS and Databricks, such as:
    • Databricks Model Serving
    • MLflow Model Registry
    • Amazon SageMaker
    • Amazon ECS or EKS
    • AWS Lambda, where appropriate
    • API Gateway
    • Application Load Balancers
  • Build synchronous, asynchronous, batch, and streaming inference capabilities.
  • Design serving architectures for LLM-powered applications, including:
    • Hosted foundation models
    • Open-source models
    • Fine-tuned models
    • Retrieval-augmented generation
    • Embeddings services
    • Vector search
    • Prompt and response orchestration
    • Tool-calling and agentic workflows
  • Optimize inference performance, scalability, GPU utilization, concurrency, throughput, latency, and cost.
  • Implement autoscaling, request throttling, queuing, caching, timeout handling, and graceful degradation.
Reliability, Recovery, and Business Continuity
  • Build recoverable model-serving endpoints with clearly defined recovery time and recovery point objectives.
  • Implement automated health checks, failover mechanisms, retry policies, circuit breakers, and service recovery procedures.
  • Design backup and recovery processes for:
    • Model artifacts
    • Model registry metadata
    • Feature definitions
    • Deployment configurations
    • Infrastructure state
    • Prompts and application configuration
    • Vector indexes and knowledge-base assets
  • Create disaster recovery procedures and regularly test restoration and failover capabilities.
  • Ensure production services can recover from failed deployments, infrastructure outages, model errors, and upstream dependency failures.
  • Develop operational runbooks and incident response procedures.
Monitoring and Observability
  • Implement end-to-end observability for infrastructure, applications, models, data, and user interactions.
  • Monitor:
    • Availability
    • Request volume
    • Latency
    • Error rates
    • Resource utilization
    • Model performance
    • Data quality
    • Data drift
    • Concept drift
    • Prediction distributions
    • LLM response quality
    • Hallucination and safety indicators
    • Token consumption
    • Cost per request
  • Establish dashboards, alerts, service-level indicators, and service-level objectives.
  • Integrate monitoring with incident management and on-call processes.
  • Enable traceability from user requests through model inference, retrieval, orchestration, and downstream services.
  • Support root-cause analysis by maintaining structured logs, metrics, traces, model lineage, and deployment history.
Security and Governance
  • Implement security controls for model-serving environments, APIs, data access, and deployment pipelines.
  • Apply least-privilege access using AWS IAM, Databricks permissions, service principals, and role-based access control.
  • Secure secrets, credentials, API keys, certificates, and tokens using approved secrets-management solutions.
  • Implement encryption in transit and at rest.
  • Design private networking, endpoint controls, firewall rules, and secure connectivity patterns.
  • Support authentication, authorization, rate limiting, and tenant isolation for AI services.
  • Ensure models and LLM applications comply with organizational requirements for privacy, security, auditability, and responsible AI.
  • Maintain model lineage, approval records, version history, and deployment audit trails.
  • Implement controls for sensitive data, personally identifiable information, prompt injection, unsafe outputs, and unauthorized model access.
Infrastructure as Code and Automation
  • Build and
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior MLOps Engineer - Scale AI with Secure Deployments
Senior MLOps Engineer - Scale AI with Secure Deployments

HP • Tlaquepaque

Presencial
MXN 1.200.000 - 1.600.000
MLOps Azure DevOps Engineer
MLOps Azure DevOps Engineer

Apex Systems • Estado de México

Presencial
MXN 80.000 - 120.000
MLOps Engineer
MLOps Engineer

Acute Talent • México

Presencial
MXN 600.000 - 1.000.000
SR Machine Learning Engineer
SR Machine Learning Engineer

Pacificacontinental • Ciudad de México

Presencial
MXN 1.108.000 - 1.663.000
Senior Staff Engineer - Data Science
Senior Staff Engineer - Data Science

Nagarro • Región Centro

Presencial
MXN 900.000 - 1.300.000
Data & Machine Learning Engineer
Data & Machine Learning Engineer

IDT • México

Presencial
MXN 1.048.035 - 1.572.052
Sr Lead Machine Learning Ops
Sr Lead Machine Learning Ops

Openbank México • Ciudad de México

Presencial
MXN 60.000 - 100.000
Sr. AI Engineer
Sr. AI Engineer

Johnson Controls, Inc. • San Pedro

Presencial
MXN 600.000 - 900.000
Senior AI Engineer
Senior AI Engineer

Crew • Americas

Presencial
MXN 1.907.000 - 2.601.000
FBS Sr. MLOps Engineer
FBS Sr. MLOps Engineer

Capgemini • Ciudad de México

A distancia
MXN 1.281.000 - 1.831.000
Competitive salary and performance-based bonuses
Comprehensive benefits package
Flexible work arrangements
+4