Senior MLOps Engineer (#5860)

N-iX

Georgia

Hybrid

USD 140,000 - 180,000

Full time

30 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Flexible remote work
Mentorship program
Tech talks and trainings

Job summary

N-iX is seeking a Senior MLOps Engineer to lead data & AI platform delivery for a major telecommunications client. You will build and operationalize end-to-end MLOps pipelines using SageMaker, Bedrock, and MLflow, while implementing GPU-optimized training and real-time inference in a hybrid cloud setting.

Collaborating with Data Engineering, AI Architects, and Cloud Teams, you will drive CI/CD for ML, enforce security with FPE, and establish governance and disaster recovery playbooks.

Qualifications

  • 4+ years hands-on experience in MLOps, DataOps, or Platform Engineering focusing on Amazon SageMaker (Pipelines, Feature Store, Model Registry, Endpoints).
  • Proven experience deploying and operating Generative AI/LLM models and Amazon Bedrock services.
  • Hands-on expertise with MLflow for experiment tracking, model registry, and lifecycle management.
  • Strong experience with GPU optimization and orchestration (NVIDIA A100/L40S, AWS EC2 GPU) for training and low-latency inference.
  • Proficient in building CI/CD for ML (GitLab CI/CD, GitHub Actions) and IaC (Terraform or AWS CDK).
  • Practical knowledge of LLMOps / AgentOps tools (AgentCore, prompt evaluations, Bedrock Guardrails) and RAG pipelines.
  • Security/privacy/tokenization (FPE) in ML pipelines.
  • Proficient in Python, PySpark, Docker, and Kubernetes/EKS for containerized ML workloads.

Responsibilities

  • Build and operationalize end-to-end MLOps pipelines using SageMaker Pipelines and MLflow.
  • Design and manage production SageMaker endpoints and Bedrock integrations with cost controls.
  • Implement AgentOps/LLMOps frameworks for multi-agent orchestration and RAG pipelines.
  • Operationalize real-time STT/TTS voicebot pipelines on hybrid/cloud GPU node pools.
  • Optimize GPU node pools for Azerbaijani SLM/SLM training and scalable inference.
  • Establish CI/CD for ML and IaC to enforce secure promotion workflows.
  • Integrate data de-identification and tokenization like FPE in ML pipelines.
  • Set up telemetry, monitoring, and cost anomaly alerts for AI/ML workloads.

Skills

Analytical thinking
Communication
Ownership

Tools

Amazon SageMaker
Amazon Bedrock
MLflow
GitLab CI/CD
Terraform
AWS CDK
Docker
Kubernetes
Python
PySpark

Job description

Work type:

Office/Remote

Technical Level:

Senior

Job Category:

Software Development

Client Overview:
Our client is an Azerbaijani telecommunications company, the largest mobile network operator in Azerbaijan. The main products are: Fixed telephony, Mobile telephony, Internet services, Wireless broadband, and Value-added services.

Project Objectives:
The primary goal is to accelerate the client’s Data & AI initiatives via a secure, hybrid cloud foundation on AWS while systematically modernizing the IT estate as part of the cloud migration.

Key Project Objectives include:

  • Cloud Foundation & Landing Zone: Deploy target hybrid network architectures, establishing a secure Landing Zone and hybrid Data/AI platforms on AWS.
  • Security, Compliance & Governance: Operationalize on-prem tokenization (achieving zero raw PII in the cloud), resolve policy blockers to include AWS in the ISMS, and establish a Cloud Center of Excellence (CCoE) to govern Cloud adoption.
  • AI Chatbot & Voicebot Design & Implementation: Develop and operationalize a flagship Customer Care Chatbot and Voicebot as the first hybrid-setup consumer.
Responsibilities:
  • Build, operationalize, and automate end-to-end MLOps pipelines using Amazon SageMaker Pipelines and MLflow for experiment tracking, model versioning, and registry lifecycle management.
  • Design, deploy, and manage production SageMaker inference endpoints (real-time, serverless, and batch) and Amazon Bedrock API integrations for LLM/SLM deployment with cost controls and latency optimization (Bedrock API Gatekeeper).
  • Implement AgentOps / LLMOps frameworks (AgentCore, Bedrock Guardrails, Promptfoo) to manage multi-agent orchestration, prompt evaluation, safety guardrails, and RAG retrieval pipelines.
  • Operationalize real-time STT / TTS (Speech-to-Text / Text-to-Speech) voicebot pipelines and low-latency speech inference on hybrid/cloud GPU node pools for the flagship Customer Care Voicebot.
  • Optimize specialized GPU node pools (NVIDIA A100/L40S / EC2 GPU instance types) for Azerbaijani SLM/SLM model training, fine-tuning, and scalable inference workloads.
  • Establish automated CI/CD for Machine Learning using GitLab CI/CD pipelines and Infrastructure-as-Code (Terraform or AWS CDK) to enforce security-gated MLOps promotion workflows (from SageMaker Canvas/Sandbox to production).
  • Integrate data de-identification, Format Preserving Encryption (FPE), and tokenization wrappers into ML data pipelines to ensure zero raw PII enters AWS cloud environments during model training and inference.
  • Set up telemetry, performance monitoring, model drift detection, and cost anomaly alerting for AI/ML workloads using Amazon CloudWatch, Splunk, and FinOps spend control frameworks.
  • Collaborate with Data Engineering, AI Architects, and Cloud Teams to integrate vector storage/retrieval (RAG), Apache Spark/EMR-on-EKS runtimes, and local tokenization databases.
  • Author technical MLOps runbooks, model deployment procedures, governance documentation, and disaster recovery playbooks.
Requirements:
  • 4+ years of hands-on experience in MLOps, DataOps, or Platform Engineering with a primary focus on enterprise Amazon SageMaker (Pipelines, Feature Store, Model Registry, Endpoints).
  • Proven experience deploying and operating Generative AI, LLM/SLM models, and Amazon Bedrock services alongside agentic frameworks and RAG pipelines.
  • Hands‑on expertise with MLflow for experiment tracking, model registry, and lifecycle management.
  • Solid experience in GPU optimization and orchestration (NVIDIA A100/L40S, AWS EC2 GPU instances) for model training, fine‑tuning, and low‑latency real‑time inference (STT/TTS voice pipelines).
  • Proficient in building CI/CD for Machine Learning (GitLab CI/CD, GitHub Actions) and Infrastructure-as-Code (Terraform or AWS CDK).
  • Practical knowledge of LLMOps / AgentOps tools and methodologies (AgentCore, prompt evaluations, Bedrock Guardrails, vector databases for RAG).
  • Strong understanding of data security, privacy, and tokenization (FPE, handling sensitive/PII data within ML pipelines).
  • Proficient in Python, PySpark, Docker, and Kubernetes/EKS fundamentals for containerized ML workloads.
Nice‑to‑Have Skills:
  • AWS Certified Machine Learning – Specialty certification.
  • AWS Certified Solutions Architect – Associate/Professional or AWS Certified DevOps Engineer – Professional.
  • Experience in telecom domain AI/ML applications, low‑latency real‑time voice/chat processing (ASR/TTS), or hybrid cloud data sovereignty architectures.
  • Experience with EMR-on-EKS, Starburst/Athena, or Apache Iceberg data lake integrations.
Soft Skills & Team Fit:
  • Strong critical thinking, problem‑solving, and analytical skills.
  • Excellent communication and collaboration skills to work closely with cross‑functional teams (Data Engineering, AI/GenAI Engineers, Security, Cloud/Infrastructure).
  • Results‑oriented, proactive mindset with strong ownership of deliverables within an Agile / Scrum framework.
  • Upper‑Intermediate+ English level (written and spoken).
What we propose:
  • Opportunity to lead critical, high‑impact Data & AI platform delivery for a major telecommunications operator.
  • Hands‑on work with modern MLOps and GenAI stack (Amazon SageMaker, Amazon Bedrock, MLflow, AgentCore, STT/TTS voicebot pipelines).
  • Flexible remote work options with structured, predictable collaboration within a well‑balanced team.
We offer*:
  • Flexible working format - remote, office‑based or flexible
  • A competitive salary and good compensation package
  • Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)
  • Active tech communities with regular knowledge sharing

Project: Global biopharmaceutical company

Project: Global biopharmaceutical company

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Senior MLOps Engineer
Senior MLOps Engineer

AppRecode, Inc. • Town of Middletown (NY)

On-site
USD 120,000 - 160,000
Senior MLOps Engineer - USA Onsite (Scottsdale, AZ)
Senior MLOps Engineer - USA Onsite (Scottsdale, AZ)

S27a • Scottsdale (AZ)

On-site
USD 170,000 - 250,000
Unlimited PTO
Competitive parental leave
Annual bonus program
+1
Senior DataOps Engineer (#5858)
Senior DataOps Engineer (#5858)

N-iX • Georgia

Hybrid
USD 120,000 - 180,000
MLOps Engineer
MLOps Engineer

Codinix Consulting Services • California (MO)

On-site
USD 120,000 - 150,000
MLOps Engineer - Scalable ML Pipelines & CI/CD
MLOps Engineer - Scalable ML Pipelines & CI/CD

Codinix Consulting Services • California (MO)

On-site
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Architect
MLOps Architect

Kapitus • Virginia (MN)

On-site
USD 180,000 - 240,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
Senior MLOps Engineer
Senior MLOps Engineer

C the Signs • United States

On-site
USD 120,000 - 160,000
Competitive salary and benefits
Flexible working arrangements
Continuous learning opportunities