AI Infrastructure Engineer

The Josef Group

Chantilly (VA)

On-site

USD 200,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Josef Group is seeking an AI Infrastructure Engineer in Chantilly, VA to deploy and optimize self-hosted LLM inference servers and containerize AI workloads. You will manage production Kubernetes environments with GPU scheduling and build robust AI serving infrastructure with gateways, TLS, and rate limiting.

You will also optimize GPU memory use, quantization, batching, and capacity planning, while developing CI/CD, observability, and incident response processes for production apps.

Qualifications

  • Hands-on experience deploying and serving Large Language Models (LLMs) in production.
  • Strong experience with Docker and production Kubernetes environments, including GPU scheduling.
  • Deep understanding of self-hosted AI infrastructure, including model formats, quantization, GPU memory management, batching, and inference optimization.
  • Experience supporting production applications with networking, reverse proxies, load balancing, authentication, and TLS.
  • Proficiency with Linux administration and Python and/or Bash scripting.
  • Ownership mindset with the ability to operate and improve production AI infrastructure.

Responsibilities

  • Deploy and optimize self-hosted LLM inference servers (vLLM, Ollama, and similar).
  • Containerize AI workloads using Docker and orchestrate production environments with Kubernetes, including GPU scheduling.
  • Build and maintain AI serving infrastructure, including gateways, load balancing, authentication, TLS, and rate limiting.
  • Optimize GPU utilization, memory management, quantization, batching, and capacity planning to balance performance and cost.
  • Develop and maintain CI/CD pipelines, observability, monitoring, and incident response processes.

Skills

LLM deployment
Linux administration
Python/Bash scripting
Production systems ownership
GPU scheduling optimization

Tools

Docker
Kubernetes
GPU drivers/tools
Terraform/Helm

Job description

AI Infrastructure Engineer
Top Secret or TS/SCI is required to start
$200K to $250K
Chantilly, VA
What You'll Do
  • Deploy and optimize self-hosted LLM inference servers (vLLM, Ollama, and similar).
  • Containerize AI workloads using Docker and orchestrate production environments with Kubernetes, including GPU scheduling.
  • Build and maintain AI serving infrastructure, including gateways, load balancing, authentication, TLS, and rate limiting.
  • Optimize GPU utilization, memory management, quantization, batching, and capacity planning to balance performance and cost.
  • Develop and maintain CI/CD pipelines, observability, monitoring, and incident response processes.
What You'll Bring (Required)
  • Hands‑on experience deploying and serving Large Language Models (LLMs) in production.
  • Strong experience with Docker and production Kubernetes environments, including GPU scheduling.
  • Deep understanding of self‑hosted AI infrastructure, including model formats, quantization, GPU memory management, batching, and inference optimization.
  • Experience supporting production applications with networking, reverse proxies, load balancing, authentication, and TLS.
  • Proficiency with Linux administration and Python and/or Bash scripting.
  • Ownership mindset with the ability to operate and improve production AI infrastructure.
Nice to Have
  • Experience with CUDA, NVIDIA drivers, GPU Operators, or other GPU infrastructure technologies.
  • Experience with Infrastructure as Code (Terraform, Helm).
  • Familiarity with observability and monitoring tools such as Prometheus and Grafana.
  • Experience building Retrieval‑Augmented Generation (RAG) pipelines and working with vector databases (pgvector, Qdrant, Weaviate).
  • Experience with LLM gateway tools such as LiteLLM.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Full Stack Software Engineer (AI Infrastructure)
Full Stack Software Engineer (AI Infrastructure)

Bytoa • Laurel (MD)

On-site
USD 200,000 - 220,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation

Jobtailor • Redwood City (CA)

On-site
USD 180,000 - 260,000
AI Implementation Engineer
AI Implementation Engineer

Jobtailor • New Jersey

On-site
USD 150,000 - 210,000
Software Engineer – Inference (SWE1) [D.26.0174]
Software Engineer – Inference (SWE1) [D.26.0174]

Dover Networks LLC • Maryland

On-site
USD 163,000 - 178,000
401(k) contributions
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Junior Software Engineer – Inference with Security Clearance
Junior Software Engineer – Inference with Security Clearance

Neural Solutions • Columbia (MD)

On-site
USD 138,000 - 163,000
AI Cloud Security and Infrastructure Engineer
AI Cloud Security and Infrastructure Engineer

Troutman Pepper • Atlanta (GA)

On-site
USD 130,000 - 150,000