Senior MLOps Engineer

Keysight Technologies SAles Spain SL.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equal Opportunity Employer

Job summary

Keysight Technologies is expanding its engineering team with an MLOps Engineer specializing in AWS to deploy and scale ML solutions across our manufacturing analytics platforms.

You will collaborate with ML engineers to automate pipelines, monitor model performance, and manage infrastructure for production-ready ML workstreams on AWS.

This mid-level role emphasizes DevOps in ML contexts with reliability, security, and cost efficiency in regulated industrial environments.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related technical field.
  • 3–5 years of experience in MLOps, DevOps, or cloud engineering roles with production ML deployments.
  • Deep expertise in AWS services for ML and data workflows including SageMaker, Bedrock, EMR, Lambda, S3, ECR, and orchestration tools.

Responsibilities

  • Design, implement, and maintain end-to-end MLOps pipelines on AWS, including CI/CD automation for training, validation, deployment and retraining.
  • Operationalize AWS Bedrock workflows and ensure scalable, reliable data processing pipelines.
  • Deploy and monitor models with common ML libraries and drift detection.
  • Manage infrastructure as code with Terraform or CloudFormation to provision AWS resources.
  • Implement comprehensive monitoring with CloudWatch, X-Ray, and Prometheus/Grafana integrations.
  • Collaborate in Agile teams to conduct A/B testing, version models, and automate rollbacks.
  • Enforce security and compliance best practices, including IAM, VPCs, encryption, and audit logging.
  • Troubleshoot production issues and drive continuous improvements in ML operations.

Skills

AWS
MLOps
DevOps
Python
CI/CD
SageMaker
CloudWatch

Education

BSc/MSc in CS/Engineering

Tools

Terraform
CloudFormation
OpenSearch
Pinecone
Airflow
Prometheus
Grafana
ECR
SageMaker Endpoints

Job description

Overview

We are expanding our engineering team with a dedicated MLOps Engineer specializing in AWS to support the deployment, scaling, and operationalization of machine learning solutions across our manufacturing and semiconductor analytics platforms. This role will serve as a critical bridge between our Machine Learning Engineers—focused on Generative AI and classical ML—and production environments, ensuring seamless, reliable, and efficient ML workflows.

You will collaborate closely with the Senior Machine Learning Engineer (GenAI Platform) and the Machine Learning Engineer (Classical ML and Predictive Analytics) to automate pipelines, monitor model performance, and manage infrastructure for high-stakes applications like test plan generation, anomaly detection, predictive maintenance, and market intelligence. In our AWS-centric ecosystem, you will leverage best-in-class tools to enable rapid iteration while maintaining compliance, security, and cost efficiency in regulated industrial settings.

This position is perfect for a mid-level professional with a passion for DevOps in ML contexts, who excels at turning complex models into robust, production-ready systems.

Responsibilities
  • Design, implement, and maintain end-to-end MLOps pipelines on AWS, including CI/CD automation for model training, validation, deployment, and retraining, using services like SageMaker, CodePipeline, CodeBuild, and Step Functions.
  • Support the Generative AI platform by operationalizing AWS Bedrock workflows, including RAG pipelines, vector databases (e.g., via OpenSearch or Pinecone integrations), Lambda functions, and agentic systems—ensuring scalability for large-scale data processing like historical test plans and news article summarization.
  • Enable classical ML initiatives by deploying and monitoring models built with XGBoost, Scikit-learn, and NLP architectures (e.g., RNNs/LSTMs) on AWS infrastructure, incorporating drift detection for anomaly tracking in sensor data and competitor pricing monitoring.
  • Manage infrastructure as code (IaC) using Terraform or CloudFormation to provision and optimize AWS resources, such as EC2 instances, S3 buckets, EMR for Apache Spark-based processing (supporting our PMA product), and ECS/EKS for containerized deployments.
  • Implement comprehensive monitoring, logging, and alerting systems with CloudWatch, X-Ray, and third-party tools (e.g., Prometheus/Grafana integrations) to track model performance, detect anomalies, handle concept drift, and ensure high availability for customer-facing tools like Q&A chatbots and predictive maintenance advisors.
  • Collaborate in an Agile environment with ML engineers, data scientists, and SRE teams to conduct A/B testing, version models, automate rollbacks, and optimize costs through auto-scaling and spot instances.
  • Enforce security and compliance best practices, including IAM roles, VPC configurations, data encryption, and audit logging, to safeguard sensitive manufacturing data and meet industry standards.
  • Troubleshoot production issues, perform root-cause analysis, and drive continuous improvements in ML operations, staying ahead of AWS innovations to enhance platform reliability and efficiency.
Qualifications
  • Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related technical field.
  • 3–5 years of experience in MLOps, DevOps, or cloud engineering roles, with a proven track record of deploying and managing ML models in production environments.
  • Deep expertise in AWS services for ML and data workflows, including SageMaker (real-time endpoints, inference components, multi-instance/multi-variant deployments), Bedrock (provisioned throughput, cross-Region inference profiles for scaling & resilience), EMR (for Spark-based PMA workloads), Lambda, S3, ECR, and orchestration tools like Step Functions or Airflow.
  • Proven experience with Amazon Elastic Container Registry (ECR): building, scanning for vulnerabilities, tagging, versioning, and pushing custom Docker images for inference containers (including Bring-Your-Own-Container patterns for custom ML frameworks, vLLM, or deep learning environments); managing ECR lifecycle policies, replication across regions, and secure access via IAM roles.
  • Strong proficiency in EC2-based ML deployments and infrastructure: selecting optimal instance types (e.g., ml.g family for GPU-heavy GenAI inference, g5/g6 for newer accelerators), configuring Auto Scaling Groups, managing spot instances for cost optimization, and handling EC2 fleets for custom hosting when SageMaker/Bedrock abstractions are insufficient.
  • Expertise in load balancing & scaling for ML inference: configuring and troubleshooting Application Load Balancers (ALB) or Network Load Balancers (NLB) integrated with SageMaker endpoints or ECS/EKS tasks; implementing SageMaker's built-in routing strategies (e.g., least outstanding requests for latency optimization); setting up auto-scaling policies (target tracking on CPU utilization, invocations per instance, or custom CloudWatch metrics); using cross-Region inference profiles in Bedrock for burst handling and global resilience; and ensuring high availability through multi-AZ deployments with minimum instance counts ≥2.
  • Demonstrated ability to resolve common deployment issues in production ML environments, including: cold‑start latency in serverless/containerized inference, container pull failures from ECR, IAM permission misconfigurations causing access denied errors, model artifact corruption or version mismatches post‑deployment, endpoint update failures without downtime (using blue/green or canary strategies), drift/throttling in high‑concurrency scenarios (e.g., 429 errors in Bedrock), unhealthy instance recovery, and debugging via CloudWatch Logs, X‑Ray traces, and SageMaker Model Monitor alerts.
  • Proficiency in IaC tools such as Terraform or CloudFormation to provision and optimize AWS resources (e.g., ECR repositories, EC2 fleets, ALBs, SageMaker endpoints, and auto‑scaling configurations) in a repeatable, auditable manner.
  • Strong scripting and programming skills in Python (with libraries like Boto3), along with experience in CI/CD pipelines using Jenkins, GitHub Actions, or AWS CodePipeline — with specific focus on automated ECR image builds, model artifact promotion, and safe endpoint updates.
  • Familiarity with monitoring and observability stacks (e.g., CloudWatch, ELK Stack) and ML-specific tools for versioning (e.g., MLflow) and experiment tracking.
  • Experience in Agile methodologies, with hands‑on participation in sprints, code reviews, and cross‑functional problem‑solving.
  • Solid understanding of ML concepts, including model drift, bias detection, and serving patterns, to effectively support both GenAI and classical ML teams.

Keysight is an Equal Opportunity Employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer
Senior MLOps Engineer

keysight technologies singapore (sales) pte. ltd. • Singapore

On-site
SGD 90,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

Keysight Technologies SAles Spain SL. • Singapore

On-site
SGD 120,000 - 180,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Keysight Technologies • Singapore

On-site
SGD 120,000 - 190,000
Machine Learning Engineer
Machine Learning Engineer

KEYSIGHT TECHNOLOGIES SINGAPORE (SALES) PTE. LTD. • Singapore

On-site
SGD 120,000 - 190,000
Senior AI Platform Engineer (MLOps & Data Science Infrastructure
Senior AI Platform Engineer (MLOps & Data Science Infrastructure

PEOPLESEARCH PTE. LTD. • Singapore

Hybrid
SGD 180,000 - 240,000
Senior MLOps Engineer - AWS AI Pipelines & Scale
Senior MLOps Engineer - AWS AI Pipelines & Scale

Keysight Technologies SAles Spain SL. • Singapore

On-site
SGD 120,000 - 180,000
Equal Opportunity Employer
Lead ML Engineer — GenAI, Anomaly Detection & MLOps
Lead ML Engineer — GenAI, Anomaly Detection & MLOps

Keysight Technologies • Singapore

On-site
SGD 120,000 - 190,000
Project Manager, MLOps
Project Manager, MLOps

hyundai motor group innovation center in singapore pte. ltd. • Singapore

On-site
SGD 180,000 - 240,000
Senior AI/Machine Learning Engineer
Senior AI/Machine Learning Engineer

Good co India • Singapore

On-site
SGD 120,000 - 170,000
Senior Devops Engineer
Senior Devops Engineer

luxoft information technology (singapore) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000