Senior AI Infrastructure Engineer Amsterdam

Together Computer Inc

Amsterdam

Hybrid

EUR 120,000 - 150,000

Full time

19 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Together AI in Amsterdam is hiring a Senior AI Infrastructure Engineer to help build the next-generation AI cloud, enabling self-serve ML workloads and massive GPU-accelerated compute across data centers worldwide.

You will design and operate high-performance backend services, manage hardware orchestration, and contribute to open-source Together AI platform with a strong focus on reliability, scalability, and security. Hybrid work: two days weekly in Amsterdam.

Qualifications

  • 5+ years of professional software development experience and proficiency in at least one backend programming language (Golang desired)
  • 5+ years experience writing high-performance, well-tested, production quality code
  • Demonstrated experience with building and operating high-performance and/or globally distributed micro-service architectures across one or more cloud providers (AWS, Azure, GCP)
  • Excellent communication skills - able to write clear design docs and work effectively with both technical and non-technical team members
  • Deep experience with Kubernetes internals a big plus, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches thereon or Kubernetes itself
  • Deep experience with VMs/hypervisors a big plus, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV
  • Deep experience with DC networking tech + solutions a big plus, such as VLAN, VXLAN, VPN, VPC, OVS/OVN
  • Experience with Cluster API or similar a big plus
  • Experience working on high-performance compute, networking, and/or storage a big plus
  • Experience virtualizing GPUs and/or Infiniband a big plus
  • Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale
  • Experience with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD)
  • Experience building IaaS or PaaS systems at scale a plus
  • Experience with DPUs/SmartNICs a plus
  • GPU programming, NCCL, CUDA knowledge a plus

Responsibilities

  • Design, build, and maintain performant, secure, and highly-available backend services/operators that run in our data centers and automate hardware management, such as Infiniband partitioning, in-DC parallel storage provisioning, and VM provisioning
  • Design and build out the IaaS software layer for a new GB200 data center with thousands of GPUs
  • Work on a global multi-exabyte high-performance object store, serving massive datasets for pretraining
  • Build advanced observability stacks for our customers with automated node lifecycle management for fault-tolerant distributed pretraining
  • Perform architecture and research work for decentralized AI workloads
  • Work on the core, open-source Together AI platform
  • Create services, tools, and developer documentation
  • Create testing frameworks for robustness and fault-tolerance

Skills

Golang
Backend development
Distributed systems
Cloud providers
Networking
CI/CD
GPU/DPUs
Observability

Tools

Kubernetes
Terraform
Ansible
Prometheus
Grafana
GitHub Actions
ArgoCD

Job description

Senior Software Engineer Together Cloud Infrastructure
About the Role

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure.

As a Senior AI Infrastructure Engineer, you will play a key role in building the next generation AI cloud platform - a highly available, global, blazing-fast cloud infrastructure that virtualizes cutting-edge ML hardware (GB200s/GB300s, BlueField DPUs) and enables state-of-the-art ML practitioners with self-serve AI cloud services, such as on-demand + managed Kubernetes and Slurm clusters. This platform serves both our internal SaaS products (inference, fine-tuning) and our external cloud customers, spanning dozens of data centers across the world.

Hybrid working two days a week in the Amsterdam office.

Responsibilities
  • Design, build, and maintain performant, secure, and highly-available backend services/operators that run in our data centers and automate hardware management, such as Infiniband partitioning, in-DC parallel storage provisioning, and VM provisioning.
  • Design and build out the IaaS software layer for a new GB200 data center with thousands of GPUs.
  • Work on a global multi-exabyte high-performance object store, serving massive datasets for pretraining.
  • Build advanced observability stacks for our customers with automated node lifecycle management for fault-tolerant distributed pretraining.
  • Perform architecture and research work for decentralized AI workloads
  • Work on the core, open-source Together AI platform
  • Create services, tools, and developer documentation
  • Create testing frameworks for robustness and fault-tolerance

To be successful, you'll need to be deeply technical and possess excellent communication, collaboration, and diplomacy skills. You have strong fundamental software development skills. In addition, you have strong systems knowledge and troubleshooting abilities.

Requirements
  • 5+ years of professional software development experience and proficiency in at least one backend programming language (Golang desired)
  • 5+ years experience writing high-performance, well-tested, production quality code
  • Demonstrated experience with building and operating high-performance and/or globally distributed micro-service architectures across one or more cloud providers (AWS, Azure, GCP)
  • Excellent communication skills - able to write clear design docs and work effectively with both technical and non-technical team members
  • Deep experience with Kubernetes internals a big plus, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches thereon or Kubernetes itself
  • Deep experience with VMs/hypervisors a big plus, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV
  • Deep experience with DC networking tech + solutions a big plus, such as VLAN, VXLAN, VPN, VPC, OVS/OVN
  • Experience with Cluster API or similar a big plus
  • Experience working on high-performance compute, networking, and/or storage a big plus
  • Experience virtualizing GPUs and/or Infiniband a big plus
  • Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale
  • Experience with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD)
  • Experience building IaaS or PaaS systems at scale a plus
  • Experience with DPUs/SmartNICs a plus
  • GPU programming, NCCL, CUDA knowledge a plus
About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer Together Cloud Infrastructure
Senior Software Engineer Together Cloud Infrastructure

Together AI • Amsterdam

On-site
EUR 70,000 - 90,000
Lead/Manager Together Cloud Infrastructure
Lead/Manager Together Cloud Infrastructure

Together AI • Netherlands

Hybrid
EUR 80,000 - 100,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Together • Amsterdam

Hybrid
EUR 90,000 - 130,000
Lead/Manager Together Cloud Infrastructure
Lead/Manager Together Cloud Infrastructure

Together AI • Amsterdam

On-site
EUR 80,000 - 120,000
Lead/Manager Together Cloud Infrastructure
Lead/Manager Together Cloud Infrastructure

Together AI • Amsterdam

Hybrid
EUR 80,000 - 100,000
Junior/Senior/Staff Software Engineer, Inference / Compute Infrastructure Engineering Together AI Amsterdam
Junior/Senior/Staff Software Engineer, Inference / Compute Infrastructure Engineering Together AI Amsterdam

Neura Market • Amsterdam

Hybrid
EUR 120,000 - 180,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Together AI • Amsterdam

On-site
EUR 95,000 - 150,000
Senior AI Infra Engineer - Global GPU Cloud (Hybrid)
Senior AI Infra Engineer - Global GPU Cloud (Hybrid)

Together Computer Inc • Amsterdam

Hybrid
EUR 120,000 - 150,000
Junior/Senior/Staff Software Engineer, Inference / Compute Infrastructure Engineering
Junior/Senior/Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • Amsterdam

On-site
EUR 90,000 - 130,000
Senior Network Engineer (Amsterdam)
Senior Network Engineer (Amsterdam)

Together AI • Netherlands

On-site
EUR 90,000 - 130,000