AI Platform Engineer

Reply

Torino

On-site

EUR 50,000 - 75,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Reply in Torino, Italy is seeking an AI Platform Engineer to own and scale the core Model Factory platform. You will orchestrate GPU workloads, ensure reproducibility of experiments, and drive governance across data, models and evaluation runs.

Join a team mentoring researchers and engineers, work with Kubernetes, SkyPilot, Dagster, NVIDIA NeMo, and Unity Catalog, and contribute to cost attribution and multi- tenancy from day one.

Qualifications

  • Degree in Computer Engineering, Computer Science or related field; fundamentals in distributed systems.
  • 3-8 years of professional experience; strong Python and infrastructure as code; scalable ML platforms experience.
  • Pragmatic, ownership-driven with clear written communication and automation bias.
  • Nice to have SkyPilot, Dagster or NVIDIA NeMo; HPC schedulers; Unity Catalog; on-prem AI environments.

Responsibilities

  • Maintain and scale the Model Factory platform core and orchestrate GPU workloads.
  • Implement versioning and governance for data, models and evaluation runs.
  • Collaborate with AI researchers, model design engineers, project managers, cloud and security specialists.

Skills

Kubernetes
Python
Infrastructure as code
Distributed training
ML platforms at scale

Education

Degree in Computer Engineering or Computer Science

Tools

SkyPilot
Dagster
NVIDIA NeMo
Unity Catalog
GitOps CI/CD

Job description

Are you an AI Platform Engineer expert in the platform that turns a researcher's idea into a running job on a cluster?

Join Reply!

WHO WE ARE

Reply is a network of highly specialized companies which support leading industrial groups in defining and developing business models using new technology and communication paradigms, such as Big Data, Cloud Computing, Digital Communication, Internet of Things, Mobile and Social Networking. Reply focuses on Consultancy, System Integration and Application Management, covering three areas of competence: Processes, Applications and Technologies.

WHAT YOU WILL FIND
  • Core activities. You will keep the core of the Model Factory platform standing, and you will make it grow. You will orchestrate GPU workloads and distributed training that hold up under load, build versioning and governance across data, models and evaluation runs so results are still reproducible months later, and shorten the distance between a researcher's idea and a running job on the cluster. Every model we train, distil or optimise — for our own portfolio and for client engagements — runs on what you build.
  • Tech & tools stack. You will work SkyPilot, Dagster, NVIDIA NeMo, Kubernetes, MLflow and Unity Catalog, the core of the platform. Around them: the NVIDIA stack (CUDA, NCCL), job scheduling across national (Italian) and European HPC clusters and cloud (such as Scaleway or Nebius), infrastructure as code and GitOps CI/CD, vLLM for inference serving, and Prometheus and Grafana observability that tells us what a run cost and why it failed.
  • Mentoring, collaboration, and continuous growth. You will work alongside the AI research engineers who train the models, the model design engineers who specify them, the project managers who commit the dates, and cloud and security specialists — plus client platform teams when the environment lands on their side. Your first customers are internal: if the platform is slow, opaque or fragile, they feel it long before any client does.
WHAT WE OFFER
  • An offer tailored to your experience. This position is open to people with varying levels of expertise and seniority. The compensation package will be determined based on your professional background, technical skills, expertise, and the level of responsibility associated with the role. The collective labour agreement (CCNL) applied is the Italian Metalworking Industry Agreement. The job classification will be assessed starting from level B2, and the gross annual salary (RAL) will range from €50,000 to €75,000. Previous experience with cutting-edge technologies such as GPU cluster orchestration, distributed training on national and European HPC infrastructure, SkyPilot, Dagster and NVIDIA NeMo, LLM inference optimisation, and data, model and evaluation governance in MLflow and Unity Catalog will be considered a strong asset.
  • A structured career path. Our Career Path offers opportunities to grow both as a technical specialist and as a future leader. Your ambitions, the skills you develop, and the results you achieve will shape your professional journey.
  • Continuous learning. Technology evolves rapidly - and so do we. You will join an environment that encourages curiosity, continuous learning, and the exploration of new ideas and emerging technologies.
  • The benefits of being a Replyer. You will also have access to the benefits and initiatives dedicated to our community.
WHAT WE NEED
  • Academic background. Degree in Computer Engineering, Computer Science or a related technical field. Solid fundamentals in distributed systems, networking matter more to us than any specific coursework.
  • Technical & strategic skills.3-8 years of professional experience.Kubernetes and distributed training that genuinely work in your hands: you can debug a training job that hangs, size a GPU allocation, and tell a scheduling problem from a networking one. Strong Python, and infrastructure as code as a default rather than an afterthought. Having run ML platforms at scale is required — we care that the fundamentals are real and that you design for reproducibility, multi-tenancy and cost attribution from day one. The position is open to varying levels of expertise and seniority.
  • Soft skills. You should be pragmatic and ownership-driven, comfortable being the person everyone else depends on. Clear written communication, a bias for automation over heroics, and the patience to make someone else's workflow actually work.
  • Nice to have. Experience with SkyPilot, Dagster or NVIDIA NeMo in production; HPC schedulers (Slurm) and InfiniBand/RDMA networking; LLM inference optimisation (quantisation, continuous batching, KV-cache); Unity Catalog or Databricks governance; open-source contributions to the ML infrastructure ecosystem; sovereign or on-premise AI environments.
WHAT ARE THE NEXT STEPS

The first step of our recruiting process will be the meetings with the technical referents and then a face to face interview with the HR team. We care about an equal recruiting process.

Reply is committed to embracing diversity and creating an inclusive work environment by valuing the uniqueness of people regardless of age, gender, sexual orientation, religion, nationality, or disabilities as protected by Italian Law (L.68/99).

Furthermore, Reply is committed to ensuring a fair and accessible selection process: to help you during the recruitment process, please let us know of any kind of support you may need.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer
AI Research Engineer

Reply • Torino

Hybrid
EUR 55,000 - 85,000
Competitive salary
Structured career path
Continuous learning
+1
AI Project Manager
AI Project Manager

Reply • Milano

On-site
EUR 37,000 - 65,000
Structured career path
Continuous learning
Community benefits
AI Strategist & Solution Designer
AI Strategist & Solution Designer

Reply • Milano

On-site
EUR 37,000 - 65,000
MIT xPRO course access
AI Delivery Manager
AI Delivery Manager

Reply • Milano

On-site
EUR 45,000 - 85,000
Structured career path
Continuous learning
AI Engineer
AI Engineer

Reply • Turbigo

On-site
EUR 40,000 - 60,000
Hands-on experience
Flexible working hours
Free coffee
AI-Powered Software Engineer
AI-Powered Software Engineer

Reply • Roma

On-site
EUR 31,000 - 39,000
Competitive salary
Career path
Continuous learning
+1
Ai, Data And Emerging Tech Consultant
Ai, Data And Emerging Tech Consultant

Reply • Lazio

On-site
EUR 26,000 - 39,000
Structured growth path
Professional development program
Community benefits
Cyber Security Expert
Cyber Security Expert

Reply • Italy

On-site
EUR 33,200 - 65,000
Join Us: Technical Project Manager - Data Engineering & Infrastructure
Join Us: Technical Project Manager - Data Engineering & Infrastructure

Target Reply • Milano

Hybrid
EUR 31,000 - 38,000
Hybrid work policy
Mentoring & learning
Company benefits
+1
Join Us: Java Software Developer to Build Digital Solutions
Join Us: Java Software Developer to Build Digital Solutions

Target Reply • Milano

On-site
EUR 25,000 - 36,000
Mentoring & training
Hybrid policy (1–2 days on site)