Stand out for this role — generate a tailored resume and cover letter in about a minute.
Arriver System Software S.R.L. is seeking a Staff Software Engineer to join the ML Platform team.
You will design, build, and optimize large-scale ML infrastructure across on-prem NVIDIA DGX clusters and AWS Cloud, enabling training pipelines and production-grade ML solutions. You will write high-quality, maintainable code in multiple languages, lead complex projects, and collaborate with data science and engineering teams to deliver scalable data workflows and reliable platform services across
Arriver System Software S.R.L.
Engineering Group, Engineering Group > Software Engineering
We are seeking an exceptional Staff Software Engineer to join our ML Platform team. This role is ideal for a highly technical individual contributor who combines deep hands‑on expertise with the ability to lead and drive complex projects independently. You will design, build, and optimize large‑scale ML and data infrastructure across on‑premises NVIDIA DGX clusters and AWS Cloud, enabling advanced training pipelines, robust data workflows, and production‑ready ML solutions.
Beyond infrastructure, this position requires strong software engineering fundamentals. You will be expected to write high‑quality, production‑grade code and master common programming languages. The ideal candidate is passionate about building scalable systems and can balance platform engineering with solid software development practices.
Architect and develop core components of the ML platform and data infrastructure for training, inference, and large-scale data processing.
Design and implement scalable solutions for GPU clusters and distributed data pipelines on-prem and in AWS.
Lead project workstreams , ensuring timely delivery and alignment with platform roadmap; operate independently and drive outcomes end‑to‑end.
Build and optimize data pipelines for ingestion, transformation, storage, and retrieval supporting ML workflows (batch and streaming).
Write clean, efficient, and maintainable code (services, operators, automation, tooling) in multiple programming languages.
Collaborate with data science and engineering teams to integrate ML and data workflows seamlessly (feature stores, model registries, artifact stores).
Implement CI/CD for ML and data workflows using Argo Workflows, ArgoCD, GitHub Actions ; champion testability and reproducibility.
Maintain observability (Prometheus, Grafana) and logging (AWS CloudWatch, ELK/Opensearch); drive SLOs, tracing, and cost‑awareness.
Operate AWS services (EKS, EC2, VPC, IAM, S3, EFS, Batch) across hybrid environments; contribute to security and compliance controls.
Continuously improve platform reliability, performance (GPU utilization, throughput), and developer experience; stay current on modern MLOps/Data engineering practices.
Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.