AI Infrastructure & Operations Engineer

KYLG AB

Sollentuna kommun

On-site

SEK 700,000 - 950,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary
Flexible working environment
Opportunity to work with cutting-edge

Job summary

KYLG AB in Sweden seeks an Infrastructure & Operations Engineer to design, deploy, and maintain AI server and GPU infrastructure, including NVIDIA GPU clusters, CUDA, Docker, and Kubernetes in a data center environment.

You will install, rack, cable, commission, and manage AI computing environments, deploy and optimize model training and inference workloads, and monitor system performance, power, cooling, and hardware health while supporting AI researchers and software engineers.

Qualifications

  • Strong Linux system administration required.
  • Experience deploying and maintaining NVIDIA GPU servers.
  • Experience with data center infrastructure, including power distribution, rack installation, cabling, cooling, and hardware commissioning.
  • Experience with Docker, Kubernetes, and the AI software stack.
  • Experience with Python and AI frameworks (PyTorch/TensorFlow).
  • Familiar with LLM deployment and enterprise networking.

Responsibilities

  • Design, deploy, and maintain AI server and GPU infrastructure.
  • Install, rack, cable, commission, and maintain AI servers in data center environments.
  • Plan and execute AI infrastructure expansion, upgrades, and hardware lifecycle management.
  • Deploy and manage Linux servers.
  • Install, configure, and maintain NVIDIA CUDA, Docker, Kubernetes, and the AI software stack.
  • Deploy and optimize AI model training and inference environments.
  • Build and maintain AI agent platforms.
  • Develop AI workflows and automation pipelines.
  • Maintain AI computing clusters and GPU infrastructure.
  • Monitor system performance and optimize GPU utilization.
  • Monitor power distribution, cooling systems, and hardware health.
  • Troubleshoot AI server hardware, GPU, networking, storage, OS and software issues.
  • Support AI researchers and software engineers.

Skills

Linux administration
AI hardware deployment
Data center infra
NVIDIA GPU servers
Docker
Kubernetes
Python
CUDA
PyTorch
TensorFlow
Git
Networking

Tools

Docker
Kubernetes
CUDA toolkit
NVIDIA DGX
Git

Job description

Sammanfattning

Join an expanding AI and Data Mining company as an Infrastructure & Operations Engineer, responsible for building and maintaining AI laboratory infrastructure. This role involves working with NVIDIA GPU clusters and enterprise AI platforms, focusing on deployment, operation, and optimization of AI computing environments. The position is based in a data center environment, requiring strong Linux and AI hardware experience, and offers opportunities to work with cutting-edge technologies.

  • Competitive salary
  • Flexible working environment
  • Opportunity to work with cutting-edge AI technologies
  • Career development in AI infrastructure and intelligent systems
AI Laboratory Infrastructure & Operations Engineer
About Us

We are an AI and Data Mining company developing. Our AI infrastructure includes NVIDIA GPU clusters, enterprise AI platforms, large-scale model training and inference environments, and intelligent AI workflow automation.

We are expanding our AI Laboratory and are looking for an experienced engineer to build and operate our AI computing environment.

Responsibilities
  • Design, deploy, and maintain AI server and GPU infrastructure
  • Install, rack, cable, commission, and maintain AI servers in data center environments
  • Plan and execute AI infrastructure expansion, upgrades, and hardware lifecycle management
  • Deploy and manage Linux servers
  • Install, configure, and maintain NVIDIA CUDA, Docker, Kubernetes, and the AI software stack
  • Deploy and optimize AI model training and inference environments
  • Build and maintain AI agent platforms
  • Develop AI workflows and automation pipelines
  • Maintain AI computing clusters and GPU infrastructure
  • Monitor system performance and optimize GPU utilization
  • Monitor power distribution, cooling systems, and overall hardware health
  • Troubleshoot AI server hardware, GPU, networking, storage, operating system, and software issues
  • Support AI researchers and software engineers
Requirements
Required
  • Strong Linux system administration experience
  • Experience with AI server hardware deployment, installation, maintenance, and troubleshooting
  • Experience with data center infrastructure, including power distribution, rack installation, cabling, cooling, and hardware commissioning
  • Experience deploying and maintaining NVIDIA GPU servers
  • Experience with Docker
  • Experience with Python
  • Experience with CUDA
  • Experience with AI frameworks (PyTorch / TensorFlow)
  • Familiar with LLM deployment
  • Experience with Git
  • Experience with enterprise networking (10/25/100Gb Ethernet or InfiniBand)
  • Experience assembling, upgrading and repairing enterprise servers
Preferred
  • Kubernetes
  • Slurm
  • Ansible
  • GPU cluster management
  • AI Agent Platforms
  • LangChain
  • MCP
  • Vector Database
  • RAG
  • Workflow Automation
  • CI/CD
Nice to Have
  • Experience with one or more:
  • Multi-GPU training
  • Distributed AI computing
  • NVIDIA DGX
  • AI Data Center
  • Enterprise AI Infrastructure
  • InfiniBand or other high-speed networking
  • Storage systems (NAS / SAN)
  • OpenAI API
  • DeepSeek
  • Qwen
  • Llama
We Offer
  • Competitive salary
  • Flexible working environment
  • Opportunity to work with cutting-edge AI technologies
  • Career development in AI infrastructure and intelligent systems
Language

English required

Chinese is a strong advantage(Chinese (Mandarin) proficiency is highly preferred, as the role involves collaboration with our Chinese engineering teams and technical documentation.)

AI Infrastructure & Operations Engineer
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

Motorica • Stockholms kommun

On-site
SEK 600,000 - 800,000
Stock Options program
Retirement Plan
Health Benefits (5000 SEK/year)
+4
Senior AI Enablement Engineer
Senior AI Enablement Engineer

Gears of Leo AB • Stockholms kommun

Hybrid
SEK 700,000 - 900,000
Hybrid work policy
4 weeks of Workation
30 annual vacation days
+4
Senior AI & Data Scientist
Senior AI & Data Scientist

Tyche Consulting AB • Solna kommun

Hybrid
SEK 70,000 - 90,000
Attractive bonus package
Flexible working environment
Structured professional development opportunities
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Tenth Revolution Group • Stockholms kommun

Hybrid
SEK 900,000 - 1,300,000
Mac and phone provided
Learning budget
Hybrid work model
+3
AI Enablement Engineer
AI Enablement Engineer

LeoVegas Mobile Gaming Group • Stockholms kommun

Hybrid
SEK 900,000 - 1,300,000
Hybrid work policy
Workation 4 weeks
30 vacation days
+6
AI GPU Infra & Operations Engineer
AI GPU Infra & Operations Engineer

KYLG AB • Sollentuna kommun

On-site
SEK 700,000 - 950,000
Competitive salary
Flexible working environment
Opportunity to work with cutting-edge
Senior AI Engineer
Senior AI Engineer

Capgemini • Stockholms kommun

On-site
SEK 900,000 - 1,300,000
Senior AI Engineer role (Senior to Principal)
Senior AI Engineer role (Senior to Principal)

Devoteam • Stockholms län

Hybrid
SEK 550,000 - 750,000
Wellness grants for sports activities (5,000 SEK/year)
Training and certification support
Social events
+1
Senior AI Engineer
Senior AI Engineer

Codon Consulting • Stockholms kommun

On-site
SEK 750,000 - 900,000
Knowledge sharing
Team events
Office near Mariatorget
+1
Data engineer
Data engineer

Verisure Innovation AB • Malmö kommun

On-site
SEK 45,000 - 65,000