Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

Uvation

Rondônia

Presencial

BRL 350 000 - 520 000

Tempo integral

14 dias+
Gerador de candidaturas

Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.

Ultrapassa os filtros ATS

Resumo da oferta

Uvation is seeking a Senior Linux Infrastructure Engineer in Rondônia to design, deploy, operate, and troubleshoot large-scale Linux environments for traditional workloads and AI/ML use cases. The role emphasizes hands-on BMaaS, GPU infrastructure, and enterprise storage with HPC-scale performance.

The ideal candidate will lead hardware-level deployment, manage GPU clusters, and ensure fault-tolerant operations in demanding data-center settings, collaborating with storage and networking teams to

Qualificações

  • Extensive hands-on experience with bare metal infrastructure and BMaaS.
  • Proven track record delivering enterprise Linux platforms at scale.
  • Experience with high-performance storage and AI factory platforms.

Responsabilidades

  • Design, deploy, and operate large-scale Linux infrastructure.
  • Manage BMaaS platforms and GPU clusters.
  • Administer Ceph storage and high-performance AI storage systems.
  • Ensure reliability, uptime, and operational excellence.
  • Collaborate with data-center operations and networking teams.

Conhecimentos

Linux administration
BMaaS
GPU infrastructure
Ceph storage
HPC networking
NVIDIA GPUs
Scripting (Bash/Python)
Enterprise storage

Ferramentas

Ceph
NVIDIA GPUs
KVM/VMware
Prometheus/Grafana

Descrição da oferta de emprego

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms.

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
  • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
  • Strong understanding of server hardware, including:
    • BIOS/UEFI
    • RAID controllers
    • Firmware management
    • iLO/iDRAC/IPMI
    • NICs and SmartNICs
    • HBA cards
    • Hardware diagnostics and troubleshooting
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
  • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
  • Understanding of NVIDIA GPU technologies including:
    • A100, H100, H200, B200, or equivalent GPU platforms
    • NVIDIA DGX and OEM GPU servers
    • GPU provisioning and lifecycle management
    • GPU monitoring and performance optimization
  • Knowledge of AI Factory architecture and infrastructure requirements
  • Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
  • Understanding of:
    • GPU resource allocation and scheduling
    • Multi-GPU systems
    • GPU networking requirements
    • High-bandwidth, low-latency infrastructure design
  • Familiarity with NVIDIA ecosystem technologies such as:
    • CUDA
    • NCCL
    • GPUDirect Storage
    • NVIDIA Fabric Manager
    • NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
  • Advanced Linux storage administration:
    • LVM
    • XFS, EXT4
    • NFS
    • iSCSI
    • Fibre Channel SAN
    • Multipath I/O
  • Strong hands-on experience with Ceph, including:
    • Cluster architecture
    • MON, OSD, MDS
    • RBD, CephFS, RGW
    • Capacity planning
    • Performance tuning
    • Failure recovery
  • Experience with high-performance AI storage platforms such as:
    • WEKA
    • VAST Data
    • Dell PowerScale
    • Pure Storage FlashBlade
    • NetApp
  • Understanding of:
    • NVMe-over-Fabrics (NVMe-oF)
    • RDMA
    • GPUDirect Storage
    • Parallel file systems
    • AI data pipelines
Networking & Infrastructure
  • Strong networking knowledge:
    • Bonding
    • VLANs
    • Routing
    • MTU optimization
    • DNS
    • DHCP
  • Experience with high-performance data center networking:
    • 100G/200G/400G Ethernet
    • RoCE
    • RDMA
    • Spine-Leaf architectures
  • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
  • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
  • Experience with high availability, clustering, and disaster recovery
  • Strong troubleshooting skills across:
    • Linux operating systems
    • Hardware platforms
    • GPU infrastructure
    • Networking
    • Enterprise storage
  • Experience supporting mission-critical production environments
  • Bash and Python scripting for automation and operational efficiency
  • Experience creating operational documentation, runbooks, and infrastructure standards
Nice to Have
  • Kubernetes infrastructure (especially AI/ML and GPU integration)
  • KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
  • Ansible automation
  • NVIDIA Base Command Manager
  • Slurm or HPC workload schedulers
  • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
  • Data Center Infrastructure Management (DCIM) tools
  • IPAM solutions
  • AWS, Azure, or hybrid cloud exposure
We Are Not Looking For
  • Candidates whose experience is primarily CI/CD pipeline engineering
  • Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
  • Cloud-only administrators with limited bare metal, storage, or hardware experience
  • Professionals whose primary expertise is software development rather than infrastructure engineering
Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.

Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Cloud Specialist
Cloud Specialist

NeoSpace AI • Brasil

Presencial
BRL 180 000 - 240 000
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Platform Science, Inc. • Londrina

Presencial
BRL 180 000 - 300 000
Senior Staff Engineer - DevOps
Senior Staff Engineer - DevOps

Nagarro • Rio de Janeiro

Presencial
BRL 180 000 - 260 000
Server Infrastructure Administrator
Server Infrastructure Administrator

Keywords Studios • São Paulo

Híbrido
BRL 90 000 - 120 000
Infrastructure Engineer (Azure / Kubernetes / Docker / Terraform)
Infrastructure Engineer (Azure / Kubernetes / Docker / Terraform)

Quil • Brasil

Teletrabalho
BRL 120 000 - 160 000
Flexible remote work environment
Opportunity to work on AI infrastructure
AI Engineering Lead
AI Engineering Lead

Espire Infolabs (Singapore) Pte Ltd • Região Norte

Presencial
BRL 300 000 - 600 000
Software Engineering - Distributed Agentic AI Systems
Software Engineering - Distributed Agentic AI Systems

HP • Porto Alegre

Presencial
BRL 300 000 - 550 000
Infrastructure Engineer (Brazil)
Infrastructure Engineer (Brazil)

Articul8 AI • Brasil

Presencial
BRL 120 000 - 180 000
Platform Engineer Id90030
Platform Engineer Id90030

Agileengine • Belo Horizonte

Híbrido
BRL 622 000 - 985 000
Professional growth
Competitive compensation
A selection of exciting projects
+1
Senior AI Compute Infrastructure Engineer
Senior AI Compute Infrastructure Engineer

Kraken • Brasil

Presencial
BRL 120 000 - 150 000