Senior AI Infrastructure Engineer: Scalable GPU & Kubernetes

Seekr

Reston (VA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity RSUs
Unlimited PTO
Hybrid work environment
401(k) with company match
Health insurance
Parental leave

Job summary

Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large‑scale training, serving, evaluation, and deployment of foundation models and autonomous AI agents.

You will work across distributed systems, Kubernetes, GPU infrastructure, high‑performance inference, and enterprise AI platforms to build secure, scalable, and highly reliable systems capable of serving

Qualifications

  • 8–12+ years of professional software engineering experience in distributed systems or cloud infrastructure.
  • Experience designing and operating production Kubernetes environments.
  • Strong software engineering skills in Python and Go/Rust/C/C++.
  • Proven ability to design, build, and operate production AI or machine learning infrastructure.
  • Familiarity with cloud platforms such as AWS, Azure, OCI, or GCP.

Responsibilities

  • Design, build, deploy, and maintain production AI infrastructure for training, fine-tuning, inference, evaluation, and agentic workloads.
  • Operate scalable Kubernetes-based infrastructure for GPU workloads across cloud, on-prem, hybrid, and edge environments.
  • Architect and optimize high-performance inference platforms from edge to trillion-parameter models with focus on latency, throughput, scalability, reliability, and cost.
  • Build and maintain distributed systems enabling scheduling, orchestration, deployment, monitoring, and lifecycle management of AI workloads.
  • Develop enterprise platforms supporting autonomous and multi-agent AI systems, including governance and observability.
  • Automate AI infrastructure using IaC, GitOps, CI/CD, and modern software practices.
  • Evaluate and integrate new AI infra tech to improve performance and reliability.
  • Collaborate with engineering, research, product, and cross-functional teams to deliver secure, scalable platforms.
  • Lead technical design discussions, conduct architecture reviews, mentor engineers, and establish standards.

Skills

Distributed systems
Kubernetes
GPU infrastructure
Python
Go
Rust
C++
Cloud platforms
CI/CD
Tech leadership
Performance tuning
Architecture design

Education

Bachelor's degree or higher

Tools

Terraform
Argo CD
Prometheus
Grafana
OpenTelemetry

Job description

Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large‑scale training, serving, evaluation, and deployment of foundation models and autonomous AI agents.

You will work across distributed systems, Kubernetes, GPU infrastructure, high‑performance inference, and enterprise AI platforms to build secure, scalable, and highly reliable systems capable of serving

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Kubernetes & GPU
Senior AI Infrastructure Engineer — Kubernetes & GPU

Seekr • Austin (TX)

Hybrid
USD 180,000 - 240,000
Equity ownership
Unlimited PTO
14 holidays
+4
Senior AI Infra Architect: Scalable GPU & Kubernetes
Senior AI Infra Architect: Scalable GPU & Kubernetes

Seekr • Austin (TX)

Hybrid
USD 190,000 - 270,000
Equity Ownership
Unlimited PTO
Hybrid work environment
+2
Lead AI Infrastructure Engineer: Kubernetes & GPU
Lead AI Infrastructure Engineer: Kubernetes & GPU

Seekr • Washington

Hybrid
USD 180,000 - 230,000
Equity ownership
Unlimited PTO
Hybrid work (Reston, VA & Austin, TX)
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2
Senior AI Infra Architect — Scale Enterprise AI
Senior AI Infra Architect — Scale Enterprise AI

Seekr • Washington

Hybrid
USD 170,000 - 260,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+2
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer - Kubernetes & Scale
Senior AI Infrastructure Engineer - Kubernetes & Scale

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
AI Infrastructure Engineer: Build Scalable ML Platforms
AI Infrastructure Engineer: Build Scalable ML Platforms

KLA-Belgium • Milpitas (CA)

On-site
USD 136,000 - 232,000
Medical benefits
401(k) matching
Employee stock purchase program
+1