Machine Learning Engineer

Aziro Technologies LLC

United States

Remote

USD 120,000 - 170,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Aziro Technologies LLC is seeking a Machine Learning Engineer with a strong Site Reliability Engineering mindset to join our remote team. The role focuses on deploying and maintaining ML applications across Windows and Linux on-premises environments, with Kubernetes clusters and Python-based tooling.

You will collaborate with data scientists and engineers, implement monitoring with DataDog, and drive reliability through CI/CD pipelines, automation, and best-practice security and performance.

Qualifications

  • Windows and Linux production environments experience.
  • On-premises servers management and Kubernetes clusters.
  • Strong Python programming skills.
  • Understanding of ML concepts and workflows.
  • Experience with ML model deployment and lifecycle management.
  • Familiarity with monitoring and debugging tools such as DataDog.
  • Experience with CI/CD pipelines for ML apps.
  • Familiarity with AWS cloud platforms.
  • Background in Site Reliability Engineering or DevOps.
  • Strong problem-solving and attention to detail.
  • Excellent communication and collaboration skills.

Responsibilities

  • Maintain and support ML applications on Windows and Linux in on‑premises environments.
  • Manage and troubleshoot Kubernetes clusters hosting ML workloads.
  • Collaborate with data scientists and engineers to deploy ML models reliably.
  • Implement and maintain monitoring and alerting using DataDog.
  • Debug production issues with Python and monitoring tools.
  • Automate operational tasks to improve reliability and scalability.
  • Ensure security, performance, and availability for ML apps.
  • Document system architecture and deployment processes.

Skills

Python programming
DevOps mindset
Problem-solving
Collaboration
CI/CD understanding

Tools

Kubernetes
Docker
DataDog
Git CI/CD
AWS

Job description

Job Title: Machine Learning Engineer (SRE Focus)
Location: Remote
Job Summary:

We are seeking a skilled Machine Learning Engineer with a strong Site Reliability Engineering (SRE) mindset to join our team. The ideal candidate will have hands-on experience maintaining applications on both Windows and Linux environments, managing on-premises servers, and working with Kubernetes clusters. This role requires solid Python programming skills, a good understanding of machine learning concepts, and practical knowledge of ML model deployment, monitoring, and debugging.

Key Responsibilities:
  • Maintain and support machine learning applications running on Windows and Linux servers in on-premises environments.
  • Manage and troubleshoot Kubernetes clusters hosting ML workloads.
  • Collaborate with data scientists and engineers to deploy machine learning models reliably and efficiently.
  • Implement and maintain monitoring and alerting solutions using DataDog to ensure system health and performance.
  • Debug and resolve issues in production environments using Python and monitoring tools.
  • Automate operational tasks to improve system reliability and scalability.
  • Ensure best practices in security, performance, and availability for ML applications.
  • Document system architecture, deployment processes, and troubleshooting guides.
Required Qualifications:
  • Proven experience working with Windows and Linux operating systems in production environments.
  • Hands-on experience managing on-premises servers and Kubernetes clusters and Docker containers
  • Strong proficiency in Python programming.
  • Solid understanding of machine learning concepts and workflows.
  • Experience with machine learning model deployment and lifecycle management.
  • Familiarity with monitoring and debugging tools, e.g. DataDog.
  • Ability to troubleshoot complex issues in distributed systems.
  • Experience with CI/CD pipelines for ML applications.
  • Familiarity with AWS cloud platforms
  • Background in Site Reliability Engineering or DevOps practices.
  • Strong problem-solving skills and attention to detail.
  • Excellent communication and collaboration skills
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Machine Learning Systems & Reliability Engineer (Moveworks)
Staff Machine Learning Systems & Reliability Engineer (Moveworks)

ServiceNow • Mountain View (CA)

On-site
USD 250,000 - 320,000
Generous family leave
Annual learning stipend
Flexible PTO
+2
Machine Learning Engineer
Machine Learning Engineer

Jobaaj Com • United States

Remote
USD 90,000 - 135,000
Remote ML Engineer - SRE & Kubernetes Reliability
Remote ML Engineer - SRE & Kubernetes Reliability

Aziro Technologies LLC • United States

Remote
USD 120,000 - 170,000
Site Reliability Engineer with ML platform - Only W2
Site Reliability Engineer with ML platform - Only W2

Saransh Inc • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

5 Star Recruitment • Newark (NJ)

On-site
USD 120,000 - 160,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Bollinger Shipyards, Inc. • Metairie (LA)

On-site
USD 140,000 - 210,000
Remote Machine Learning Engineer
Remote Machine Learning Engineer

Angenex • Jersey City (NJ)

Remote
USD 90,000 - 120,000
Senior Machine Learning Engineer (DevOps/SRE)
Senior Machine Learning Engineer (DevOps/SRE)

Roku • Austin (TX)

On-site
USD 120,000 - 150,000
Remote ML Engineer & AI Platform Architect
Remote ML Engineer & AI Platform Architect

Two95 International Inc. • Herndon (VA)

On-site
USD 100,000 - 130,000
Remote Machine Learning Engineer
Remote Machine Learning Engineer

Spotter Labs Inc • Lemont (IL)

Remote
USD 90,000 - 150,000
Fully remote work
Flexible working environment
Real-world AI projects