devops engineer for machine learning platforms

HireHi

Polska

On-site

PLN 180,000 - 300,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Flexible work schedule
Medical insurance
Financial assistance

Job summary

Innowise ищет специалиста по MLOps-инфраструктуре для проектирования, развёртывания и поддержки устойчивых систем для ML-платформ и дата-пайплайнов. Работа с Python, SQL и облачными технологиями обеспечит надежность и масштабируемость ваших сервисов.

Кандидат должен владеть Docker, Kubernetes, GitLab/GitHub/Jenkins, MLflow/Kubeflow/ClearML и методами IaC; приветствуется опыт работы с Databricks и аналитикой данных. Гибкий график, работа с клиентами по всему миру, медицинское страхование.

Qualifications

  • Опыт управления серверной или облачной инфраструктурой.
  • Опыт работы с инструментами IaC; Docker и Kubernetes в продакшн.
  • Опыт проектирования и поддержки CI/CD в продакшне.
  • Опыт развёртывания ML-платформ (MLflow/Kubeflow/ClearML) и сопутствующих сервисов.
  • Сильная Linux-система инжиниринга; Python и Bash для автоматизации.
  • Опыт работы с Databricks и PostgreSQL, NFS, S3-совместимыми хранилищами.

Responsibilities

  • Проектирование, разворачивание и поддержка инфраструктуры для ML-платформ и данных.
  • Реализация Application Lifecycle Management для ML, автоматизация обучения, версионирования и развёртывания.
  • Обеспечение надежности, масштабируемости и HA инфраструктуры MLOps.
  • Проектирование распределённых вычислительных сред на bare metal, VM и облаке.
  • Контейнеризация ML-сервисов Docker и Kubernetes; оркестрация релизов.
  • Автоматизация Provisioning, lifecycle кластера и конфигураций via IaC.
  • Построение и поддержка CI/CD через GitLab CI, GitHub Actions и Jenkins.
  • Обеспечение наблюдаемости через логи, метрики и дашборды; RCA по инцидентам.

Skills

Python
Bash
Linux
Docker
Kubernetes
CI/CD
SQL
Databricks
MLflow
Kubeflow
ClearML
Observability
Go/Rust scripting

Tools

Terraform
Ansible
GitLab CI
GitHub Actions
Jenkins
Databricks
Kubeflow

Job description

Описание:

Innowise delivers technology projects for clients worldwide and has implemented more than 1600 cases for clients from the USA, Denmark, Germany, and other countries.


Задачи:

Design, deploy, and support resilient infrastructure for machine learning platforms and data pipelines using Python and SQL; Implement Application Lifecycle Management for machine learning, automating training, versioning, and deployment processes; Ensure reliability, scalability, and high availability of MLOps infrastructure and backend services; Design and manage distributed compute environments across bare metal, virtual machines, and private or public cloud; Containerize ML services and applications using Docker and Kubernetes, orchestrating production rollouts; Automate infrastructure provisioning, cluster lifecycle, and configuration using Infrastructure as Code; Build, integrate, and maintain CI/CD pipelines with GitLab CI, GitHub Actions, and Jenkins; Implement observability through logging, metrics, and dashboards for overall cluster health; Diagnose bottlenecks, resolve node and network failures, and conduct Root Cause Analysis as part of proactive incident management.


Требования:

Proven experience managing server or cloud infrastructure, Infrastructure as Code tools, Docker, and Kubernetes for at least 250 person-days; At least 300 person-days designing and maintaining CI/CD solutions in production environments; At least 200 person-days deploying and using MLflow, Kubeflow, ClearML, or similar platforms; At least 150 person-days in root cause analysis and stabilizing critical systems; Strong Linux systems engineering background with RHEL, Rocky, or SLES; Proficiency in Python and Bash for automation; Experience with PostgreSQL, NFS, and S3-compatible object stores; Experience with Data Analytics and Data Analysis using Databricks; Strong analytical mindset and an implementation-oriented approach; Ability to translate business requirements into scalable technical solutions; Excellent cross-functional collaboration with Data Scientists, Developers, and IT operations; Nice to have: deep experience with HPC schedulers and integrations, bridging HPC schedulers with cloud-native platforms and configuring dynamic scaling, Go or Rust scripting, SQL and PowerShell, MPI workloads and GPU scheduling, parallel file systems, configuration management with Ansible or Puppet.


Условия:


  • Flexible work schedule;

  • Experience working with clients worldwide;

  • Financial assistance;

  • Medical insurance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Platform Engineer for Scalable AI Infrastructure
MLOps Platform Engineer for Scalable AI Infrastructure

HireHi • Polska

On-site
PLN 180,000 - 300,000
Flexible work schedule
Medical insurance
Financial assistance
machine learning engineer for geospatial intelligence
machine learning engineer for geospatial intelligence

Enfint • Warszawa

On-site
PLN 240,000 - 360,000
Sport Card (MultiSport)
Medical care
Wellbeing support
+2
platform engineer AI/ML infrastructure
platform engineer AI/ML infrastructure

Enfint • Warszawa

Hybrid
PLN 180,000 - 280,000
platform engineer in fintech
platform engineer in fintech

Enfint • Warszawa

On-site
PLN 200,000 - 320,000
Profit sharing
Passion Sponsorship program
Integration events
+6
platform engineer for digital intelligence
platform engineer for digital intelligence

Enfint • Warszawa

On-site
PLN 180,000 - 240,000
devops engineer in talent consultancy
devops engineer in talent consultancy

Enfint • Warszawa

On-site
PLN 150,000 - 210,000
data engineer for cloud-native data platforms
data engineer for cloud-native data platforms

Enfint • Warszawa

On-site
PLN 180,000 - 300,000
devops engineer for GenAI platforms
devops engineer for GenAI platforms

HireHi • Polska

Hybrid
PLN 180,000 - 280,000
Flexible hours
Training & certifications
Team-building budget
+3
machine learning engineer
machine learning engineer

Enfint • Warszawa

On-site
PLN 210,000 - 300,000
mlops engineer in insurance services
mlops engineer in insurance services

HireHi • Polska

Hybrid
PLN 207,000 - 234,000