Infrastructure Engineer - AI, Kubernetes & Edge Systems

UMATR

Austin (TX)

On-site

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

UMATR is hiring an Infrastructure Engineer to own platform deployment and scaling across Kubernetes, GPU infrastructure and cloud environments. As an early engineer, you’ll shape architecture and tooling with high ownership from day one.

You’ll manage end-to-end deployment, build Kubernetes on-prem deployments, implement GitOps, and own CI/CD pipelines and AI inference workloads on customer hardware.

Qualifications

  • Strong Kubernetes expertise with troubleshooting production environments.
  • Hands-on with Kubernetes on-prem deployments (k3s, RKE2, MicroK8s or similar).
  • Production experience with Terraform and Docker in CI/CD.
  • Experience running AI/ML workloads on GPU infrastructure.
  • Own CI/CD pipelines and software delivery improvements.
  • Strong Linux/Bash skills and Python for applications.

Responsibilities

  • Own end-to-end deployment lifecycle for customer environments from bare hardware to running systems.
  • Build and maintain Kubernetes deployments across on-premise environments.
  • Develop GitOps workflows and automated reconciliation.
  • Create IaC-based cloud infrastructure for dev and staging.
  • Own CI/CD pipelines, builds, tests, deployments and releases.
  • Deploy and optimise AI inference workloads on GPU hardware.
  • Tune model serving infrastructure for performance and latency.
  • Improve observability via monitoring, tracing and diagnostics.
  • Secure systems for secrets, permissions and industrial networks.
  • Produce documentation, runbooks and automated tooling.
  • Identify opportunities to reduce manual work and improve efficiency.

Skills

Kubernetes
CI/CD
Linux
Python
GPU workloads
Observability
GitOps

Tools

Terraform
Docker
Flux/Argo CD
OpenTelemetry

Job description

We're partnering with an AI company transforming industrial environments through intelligent, on-premise AI systems. Their platform processes live operational data on customer hardware, delivering real-time insights for critical industries.

They're looking for an Infrastructure Engineer to own platform deployment, reliability and scaling across Kubernetes, GPU infrastructure and cloud environments. As an early engineer, you'll have significant ownership in shaping the architecture, tooling and future direction of the platform.

What You'll Do

  • Own the end-to-end deployment lifecycle for customer environments, from bare hardware through to fully operational systems.
  • Build and maintain Kubernetes-based deployments across single-node and multi-node on-premise environments.
  • Develop reliable GitOps workflows, managing deployments through version-controlled infrastructure and automated reconciliation.
  • Build and maintain cloud infrastructure for development and staging environments using Infrastructure as Code.
  • Own CI/CD pipelines, including automated builds, testing workflows, deployment processes and release management.
  • Deploy and optimise AI inference workloads running on customer GPU hardware.
  • Configure and tune model serving infrastructure, balancing performance, latency and hardware constraints.
  • Improve observability across applications, infrastructure and AI workloads through monitoring, tracing and diagnostics.
  • Build secure systems for managing secrets, permissions and industrial network environments.
  • Create documentation, runbooks and automation that make deployments repeatable and supportable.
  • Identify opportunities to remove manual processes and improve engineering efficiency.

Who We’re Looking For

  • Strong experience with Kubernetes, including troubleshooting production environments, networking, storage and cluster operations.
  • Hands-on experience deploying Kubernetes outside managed cloud environments (e.g. k3s, RKE2, MicroK8s or similar).
  • Production experience with Infrastructure as Code tools such as Terraform.
  • Strong Docker experience, including multi-service applications, image optimisation and registry workflows.
  • Experience owning CI/CD pipelines and improving software delivery processes.
  • Experience running AI/ML workloads on GPU infrastructure, including model serving and performance optimisation.
  • Understanding of GPU resources, including VRAM management, batching, quantisation and inference optimisation.
  • Strong Linux and Bash skills, with the ability to work with Python-based applications.
  • A first-principles approach to debugging complex technical issues.
  • Strong ownership mindset and ability to operate independently in a fast-moving environment.

Nice To Have

  • Experience with GitOps tools such as Flux or Argo CD.
  • Background working with edge computing, air-gapped environments or customer-hosted infrastructure.
  • Experience with industrial environments, manufacturing systems, PLCs or industrial protocols.
  • Knowledge of GPU scheduling, multi-model inference or GPU sharing technologies.
  • Experience managing databases such as PostgreSQL or time-series databases.
  • Familiarity with observability tools such as OpenTelemetry or AI tracing platforms.

What Is In It For You?

  • Opportunity to join an early-stage AI company solving complex real-world industrial challenges.
  • High ownership and autonomy from day one, working directly with engineering leadership.
  • Build infrastructure where the challenges go beyond traditional cloud environments.
  • Work on cutting-edge AI systems running directly on physical hardware.
  • Shape engineering practices, architecture and deployment strategy as the company scales.
  • Exposure to complex problems across infrastructure, AI, hardware and software engineering.

Interested?

If you’re an Infrastructure Engineer who enjoys solving difficult engineering problems and wants to build the next generation of AI-powered industrial systems, we’d love to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior / Lead / Principal Platform Engineer
Senior / Lead / Principal Platform Engineer

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Opportunities for career growth in a high-growth AI company
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Long-term career growth opportunities
Data Center Infrastructure Software Engineer
Data Center Infrastructure Software Engineer

Doist • Bellevue (WA)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
+3
AI Training Infrastructure Engineer
AI Training Infrastructure Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Infrastructure Engineer
Infrastructure Engineer

People In AI • United States

On-site
USD 240,000 - 280,000
Healthcare
401(k) match
Generous PTO
+4
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Colossus Technologies Group • Boston (MA)

Hybrid
USD 180,000 - 220,000
Health & wellness benefit
Competitive equity package
Hybrid work flexibility
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000