Production Support Engineer

ASM Tech Solutions

United States

Hybrid

USD 90,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ASM Tech Solutions is seeking an experienced L2 Production Support Engineer to monitor and troubleshoot AI platforms, cloud infrastructure, and data pipelines in a hybrid US-based setup. The role emphasizes deployment, automation, and proactive incident management.

You will work with DevOps, MLOps, data engineering, and platform teams to improve CI/CD and operational reliability, with visibility into observability and incident trends.

Qualifications

  • 3–5 years in Level 2 application, platform, DevOps or production support roles.
  • Experience troubleshooting cloud-native apps, distributed systems, containers, and Kubernetes- based environments.
  • Strong knowledge of CI/CD pipelines and deployment automation (GitLab or similar).
  • Experience with monitoring, logging and observability tools.
  • Familiar with ITIL incident, problem, change and release management processes.

Responsibilities

  • Monitor, troubleshoot, and resolve Level 2 production incidents across AI platforms and cloud infrastructure.
  • Provide hands-on support for deployment and orchestration of AI/ML workloads.
  • Build and maintain monitoring, alerting, and observability for infra and data pipelines.
  • Perform root cause analysis and drive permanent fixes for production issues.
  • Collaborate with DevOps, MLOps, data engineering, and platform teams on CI/CD and automation.
  • Develop automation and self-service tooling to reduce manual workload.
  • Monitor batch processes, data ingestion, model training, and deployment failures.
  • Apply ITIL-based incident and change management to production operations.
  • Analyze support-ticket trends and implement operational improvements including AI-driven automation.

Skills

UNIX/Linux
SQL
Scripting (Python/Shell)
CI/CD
Monitoring/Observability

Tools

GitLab
Docker & Kubernetes
Splunk
Grafana / AppDynamics / Prometheus
Terraform / Ansible

Job description

As of now regular shift during US hours. Hybrid work ( 3 days from Office is Mandatory and 2 days remote)

Location: Lake Mary or 240 NY Office

Key Responsibilities
  • Monitor, troubleshoot, and resolve Level 2 production incidents across AI platforms, cloud infrastructure, data pipelines, model-serving environments, and associated services.
  • Provide hands-on support for deployment, orchestration, and operational management of AI/ML workloads across cloud-native environments.
  • Build and maintain monitoring, alerting, and observability capabilities for infrastructure, applications, data pipelines, model operations, and distributed compute workloads.
  • Perform root cause analysis for production issues, implement permanent fixes, and drive problem-management activities to improve service reliability.
  • Collaborate with DevOps, MLOps, data engineering, platform engineering, and application teams to maintain and enhance CI/CD and deployment automation.
  • Develop and support automation, self-service tooling, recovery mechanisms, and self-healing controls to reduce manual operational effort.
  • Monitor and troubleshoot batch processes, workflow orchestration, data ingestion, model training, and model deployment failures.
  • Apply ITIL-based incident, problem, change, and release management processes to support stable production operations.
  • Analyse support-ticket and incident trends; recommend and implement operational improvements, including AI-driven automation where appropriate.
Qualifications & Skills Mandatory:
  • 3-5 years of experience in Level 2 application, platform, DevOps, or production support roles.
  • Strong hands-on experience with UNIX/Linux, SQL, and shell or Python scripting.
  • Experience troubleshooting cloud-native applications, distributed systems, containers, and Kubernetes-based environments.
  • Working knowledge of CI/CD pipelines, deployment automation, and source-control platforms such as GitLab.
  • Experience with monitoring, logging, and observability tools such as Splunk, Grafana, AppDynamics, Prometheus, or similar tools.
  • Understanding of incident management, root cause analysis, problem management, and ITIL support processes.
  • Strong analytical and problem-solving skills, with a client-service mindset.
  • Ability to troubleshoot data-pipeline, workflow, API, and production deployment issues.
Good-to-Have:
  • Exposure to MLOps practices, including model deployment, model monitoring, feature/data pipelines, and AI workload orchestration.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience with infrastructure automation and configuration-management tools such as Ansible, Terraform, or Ansible Tower.
  • Familiarity with workflow and job-scheduling tools, such as Contro-M or equivalent enterprise schedulers.
  • Knowledge of Docker, Kubernetes, and distributed compute/data-processing technologies.
  • Exposure to Kafka, MQ, or other messaging and event-streaming platforms.
  • Experience implementing self-healing, auto-remediation, resilience, backup, or disaster-recovery mechanisms.
  • Familiarity with AI-driven operational tooling, ticket-trend analysis, or automation agents.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support Engineer
Production Support Engineer

ASM Tech Solutions • Town of Florida (NY)

Hybrid
USD 90,000 - 120,000
L2 Production/Platform Support Engineer
L2 Production/Platform Support Engineer

ASM Tech Solutions • Town of Florida (NY)

Hybrid
USD 95,000 - 125,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Kansas Ag Connection • Kansas City (KS)

On-site
USD 140,000 - 190,000
AI/ML Production Support Engineer — Hybrid (US Hours)
AI/ML Production Support Engineer — Hybrid (US Hours)

ASM Tech Solutions • United States

Hybrid
USD 90,000 - 150,000
Hybrid AI Production Support Engineer (Cloud/ML Ops)
Hybrid AI Production Support Engineer (Cloud/ML Ops)

ASM Tech Solutions • Town of Florida (NY)

Hybrid
USD 90,000 - 120,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Dairy Farmers of America • Kansas City (KS)

On-site
USD 98,000 - 139,000
Principal AI Application Operations Engineer
Principal AI Application Operations Engineer

GE Energy (Finland) Oy • Town of Niskayuna (NY)

On-site
USD 180,000 - 230,000
Mid Level AI/ML Engineer/Remote EST/CST
Mid Level AI/ML Engineer/Remote EST/CST

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 110,000 - 160,000
Medical Insurance
Dental Insurance
Vision Insurance
+3
AI/ML Platform L2 Support Engineer - Hybrid
AI/ML Platform L2 Support Engineer - Hybrid

ASM Tech Solutions • Town of Florida (NY)

Hybrid
USD 95,000 - 125,000
Senior Engineer – AI Monitoring & Control Capabilities
Senior Engineer – AI Monitoring & Control Capabilities

Bank of America • Plano (TX)

On-site
USD 170,000 - 230,000