ML Platform Software Engineer - Timisoara, Romania

SimplyHired

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Arriver System Software S.R.L. is seeking a Staff Software Engineer to join the ML Platform team.

You will design, build, and optimize large-scale ML infrastructure across on-prem NVIDIA DGX clusters and AWS Cloud, enabling training pipelines and production-grade ML solutions. You will write high-quality, maintainable code in multiple languages, lead complex projects, and collaborate with data science and engineering teams to deliver scalable data workflows and reliable platform services across

Qualifications

  • Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Engineering or related work experience.
  • Master's degree in Engineering, Information Systems, Computer Science, or related field and 3+ years of Software Engineering or related work experience.
  • PhD in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
  • 2+ years of work experience with programming languages such as C, C++, Java, Python, etc.

Responsibilities

  • Architect and develop core components of the ML platform and data infrastructure for training, inference, and large-scale data processing.
  • Design and implement scalable solutions for GPU clusters and distributed data pipelines on-prem and in AWS.
  • Lead project workstreams, ensuring timely delivery and alignment with platform roadmap; operate independently and drive outcomes end-to-end.
  • Build and optimize data pipelines for ingestion, transformation, storage, and retrieval supporting ML workflows.
  • Write clean, efficient, and maintainable code (services, operators, automation, tooling) in multiple programming languages.
  • Collaborate with data science and engineering teams to integrate ML and data workflows seamlessly.
  • Implement CI/CD for ML and data workflows using Argo Workflows, ArgoCD, GitHub Actions; champion testability and reproducibility.
  • Maintain observability (Prometheus, Grafana) and logging (AWS CloudWatch, ELK/Opensearch); drive SLOs, tracing, and cost-awareness.
  • Operate AWS services (EKS, EC2, VPC, IAM, S3, EFS, Batch) across hybrid environments; contribute to security and compliance controls.
  • Continuously improve platform reliability, performance (GPU utilization, throughput), and developer experience; stay current on modern MLOps/Data engineering practices.

Skills

C
C++
Java
Python

Education

Bachelor's degree
Master's degree
PhD

Job description

Company:

Arriver System Software S.R.L.

Job Area:

Engineering Group, Engineering Group > Software Engineering

General Summary:

We are seeking an exceptional Staff Software Engineer to join our ML Platform team. This role is ideal for a highly technical individual contributor who combines deep hands‑on expertise with the ability to lead and drive complex projects independently. You will design, build, and optimize large‑scale ML and data infrastructure across on‑premises NVIDIA DGX clusters and AWS Cloud, enabling advanced training pipelines, robust data workflows, and production‑ready ML solutions.

Beyond infrastructure, this position requires strong software engineering fundamentals. You will be expected to write high‑quality, production‑grade code and master common programming languages. The ideal candidate is passionate about building scalable systems and can balance platform engineering with solid software development practices.

Minimum Qualifications:
  • Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Engineering or related work experience.
  • Master's degree in Engineering, Information Systems, Computer Science, or related field and 3+ years of Software Engineering or related work experience.
  • PhD in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
  • 2+ years of work experience with Programming Language such as C, C++, Java, Python, etc.
Key Responsibilities
  • Architect and develop core components of the ML platform and data infrastructure for training, inference, and large-scale data processing.

  • Design and implement scalable solutions for GPU clusters and distributed data pipelines on-prem and in AWS.

  • Lead project workstreams , ensuring timely delivery and alignment with platform roadmap; operate independently and drive outcomes end‑to‑end.

  • Build and optimize data pipelines for ingestion, transformation, storage, and retrieval supporting ML workflows (batch and streaming).

  • Write clean, efficient, and maintainable code (services, operators, automation, tooling) in multiple programming languages.

  • Collaborate with data science and engineering teams to integrate ML and data workflows seamlessly (feature stores, model registries, artifact stores).

  • Implement CI/CD for ML and data workflows using Argo Workflows, ArgoCD, GitHub Actions ; champion testability and reproducibility.

  • Maintain observability (Prometheus, Grafana) and logging (AWS CloudWatch, ELK/Opensearch); drive SLOs, tracing, and cost‑awareness.

  • Operate AWS services (EKS, EC2, VPC, IAM, S3, EFS, Batch) across hybrid environments; contribute to security and compliance controls.

  • Continuously improve platform reliability, performance (GPU utilization, throughput), and developer experience; stay current on modern MLOps/Data engineering practices.

Minimum Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Platform Engineer: Scale GPU Pipelines
Staff ML Platform Engineer: Scale GPU Pipelines

SimplyHired • United States

Remote
USD 180,000 - 240,000
ML Ops Engineer (EMEA Remote)
ML Ops Engineer (EMEA Remote)

Pragmatike • Town of Italy (NY)

On-site
USD 120,000 - 160,000
Senior MLOps Engineer - DSX Enablement
Senior MLOps Engineer - DSX Enablement

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Senior MLOps Engineer - DSX Enablement
Senior MLOps Engineer - DSX Enablement

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 357,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
ML Infrastructure Engineer
ML Infrastructure Engineer

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
Senior Software Engineer
Senior Software Engineer

Compunnel, Inc. • Westbrook (ME)

On-site
USD 120,000 - 160,000
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior Developer
Senior Developer

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000
Analytics Platform Engineer (Pharma) - Plainsboro Township, New Jersey
Analytics Platform Engineer (Pharma) - Plainsboro Township, New Jersey

MissionHires • Plainsboro Township (NJ)

On-site
USD 120,000 - 150,000