Senior AI Infrastructure Engineer — Kubernetes & GPU

Seekr

Austin (TX)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity ownership
Unlimited PTO
14 holidays
Hybrid work
401(k) match
Health insurance
Parental leave

Job summary

Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large‑scale training, serving, evaluation, and deployment of foundation models and autonomous AI agents.

You will work across distributed systems, Kubernetes, GPU infrastructure, high‑performance inference, and enterprise AI platforms to build secure, scalable, and highly reliable systems capable of serving

Qualifications

  • 8–12+ years of professional software engineering experience building distributed systems.
  • Architects systems, drives technical direction cross‑functionally
  • 4 year or higher degree or additional relevant experience, in addition to years of work experience
  • Demonstrated success designing and operating production Kubernetes environments supporting cloud‑native applications and distributed services.
  • Strong software engineering skills using Python and one or more modern programming languages such as Go, Rust, or C++.
  • Proven ability to design, build, and operate production AI or machine learning infrastructure.
  • Expertise developing and optimizing large‑scale AI inference platforms, including GPU utilization, distributed inference, batching, caching, quantization, and accelerator performance.
  • Familiarity with modern AI serving technologies such as vLLM, SGLang, TensorRT‑LLM, Triton Inference Server, Ray Serve, or similar platforms.
  • Knowledge of distributed computing, networking, storage systems, cloud‑native architectures, and infrastructure automation using technologies such as Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, and Infrastructure‑as‑Code tools.
  • Experience developing enterprise AI platforms, autonomous agents, or multi‑agent systems, including orchestration, tool execution, governance, observability, and evaluation.
  • Familiarity with event‑driven architectures, distributed messaging systems, and public cloud platforms including AWS, Azure, Oracle Cloud Infrastructure, or Google Cloud Platform.
  • Demonstrated technical leadership, including driving architectural decisions, mentoring engineers, and leading complex technical initiatives across cross‑functional teams.
  • Demonstrated ability to analyze, profile, and optimize AI systems for performance, scalability, reliability, and cost across distributed compute environments.

Responsibilities

  • Design, develop, deploy, and maintain production AI infrastructure supporting model training, fine‑tuning, inference, evaluation, and agentic AI workloads.
  • Design and operate scalable Kubernetes‑based infrastructure supporting GPU‑accelerated workloads across cloud, on‑premises, hybrid, and edge environments.
  • Architect and optimize high‑performance inference platforms capable of serving models ranging from resource‑constrained edge deployments to trillion‑parameter foundation models, with a focus on latency, throughput, scalability, reliability, and cost efficiency.
  • Build and maintain distributed systems that enable reliable scheduling, orchestration, deployment, monitoring, and lifecycle management of AI workloads.
  • Develop enterprise platforms supporting autonomous and multi‑agent AI systems, including secure tool execution, orchestration, memory, evaluation, governance, and observability.
  • Design, implement, and automate AI infrastructure using Infrastructure‑as‑Code, GitOps, CI/CD pipelines, and modern software engineering practices.
  • Evaluate and integrate emerging AI infrastructure technologies, model serving frameworks, hardware accelerators, and cloud‑native platforms to improve platform performance, scalability, and reliability.
  • Collaborate with engineering, research, product, and cross‑functional teams to deliver secure, scalable, and production‑ready AI platforms.
  • Lead technical design discussions, perform architecture reviews, mentor engineers, and establish engineering standards and best practices across the AI Infrastructure organization.
  • Participate in production support activities, including troubleshooting complex distributed systems, performance tuning, incident response, and continuous operational improvement.

Skills

Distributed systems
Kubernetes
GPU infrastructure
Python
Go
Rust
C++
CI/CD
Infrastructure-as-Code
Cloud platforms
Leadership
Telemetry / Observability

Education

4 year or higher degree

Tools

Kubernetes
Docker
Helm
Argo CD
Terraform
OpenTelemetry
Prometheus
Grafana

Job description

Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large‑scale training, serving, evaluation, and deployment of foundation models and autonomous AI agents.

You will work across distributed systems, Kubernetes, GPU infrastructure, high‑performance inference, and enterprise AI platforms to build secure, scalable, and highly reliable systems capable of serving

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer: Scalable GPU & Kubernetes
Senior AI Infrastructure Engineer: Scalable GPU & Kubernetes

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+3
Lead AI Infrastructure Engineer: Kubernetes & GPU
Lead AI Infrastructure Engineer: Kubernetes & GPU

Seekr • Washington

Hybrid
USD 180,000 - 230,000
Equity ownership
Unlimited PTO
Hybrid work (Reston, VA & Austin, TX)
Senior AI Infra Engineer: Kubernetes, GPUs & Scale
Senior AI Infra Engineer: Kubernetes, GPUs & Scale

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Unlimited PTO
14 paid company holidays
Hybrid work offices: Reston, VA &
+3
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2
Senior AI Infra Architect: Scalable GPU & Edge Platforms
Senior AI Infra Architect: Scalable GPU & Edge Platforms

Seekr • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Equity ownership – RSUs
Unlimited PTO
Flexible hybrid work environment
+3
Senior AI Infra Engineer for Scalable Enterprise Platforms
Senior AI Infra Engineer for Scalable Enterprise Platforms

Seekr • Austin (TX)

Hybrid
USD 180,000 - 260,000
Equity Ownership
Unlimited PTO and holidays
Hybrid work environment (Reston, VA &\
+2
Senior AI Infra Architect — Scale Enterprise AI
Senior AI Infra Architect — Scale Enterprise AI

Seekr • Washington

Hybrid
USD 170,000 - 260,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+2
Senior AI Platform Engineer: Kubernetes & GPU Infra
Senior AI Platform Engineer: Kubernetes & GPU Infra

EY • Charlotte (NC)

On-site
USD 126,000 - 230,000
Medical and dental coverage
Hybrid work model
Paid time off
Senior AI Infra Engineer - GPU & Kubernetes
Senior AI Infra Engineer - GPU & Kubernetes

SB Telecom America Corp. • Sunnyvale (CA)

On-site
USD 150,000 - 250,000
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1