AI Infrastructure Engineer - Inference Platform

Hoonify Technologies Inc.

Albuquerque (NM)

On-site

USD 120,000 - 190,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hoonify Technologies Inc. seeks an AI Infrastructure Engineer to build, deploy, and operate the LLM serving infrastructure of our AI/ML platform. You will optimize production inference systems, tune runtimes, and scale across NVIDIA and AMD GPUs.

You will work with senior engineers, ship well-tested changes, and grow expertise in GPU-backed workloads, observability, and continuous delivery, with opportunities for broader ownership as you prove your impact.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Applied Math, or Data Science.
  • 3 years of relevant work experience or equivalent combination of education and relevant experience.
  • Professional software engineering experience touching ML systems, GPU workloads, or high-performance backend services.
  • Working knowledge of Kubernetes in production contexts, including writing and debugging manifests.
  • Hands-on experience serving or deploying LLMs (vLLM, SGLang, TensorRT-LLM, or similar).

Responsibilities

  • Deploy, tune, and optimize high-performance LLM inference pipelines on GPU infrastructure.
  • Analyze, profile, and optimize model serving workloads across frameworks and hardware.
  • Build and operate scalable, production-grade API services for model inference.
  • Develop benchmarking harnesses, monitoring infrastructure, and automation tooling.
  • Scale inference workloads across multi-GPU, multi-node environments.
  • Collaborate with engineering and product teams to align infrastructure with services.
  • Write runbooks, design notes, and documentation for changes.
  • Participate in code/design reviews and incorporate feedback from seniors.

Skills

Kubernetes
Python
Linux
ML systems
CI/CD
Git workflows
LLM deployments
Go/Rust/C++

Education

Bachelor's degree in CS/CE/Applied Math/Data Science

Tools

vLLM
SGLang
TensorRT-LLM
Git
Grafana

Job description

We are seeking an AI Infrastructure Engineer to help build, deploy, and operate the LLM serving infrastructure underpinning our AI/ML platform. This role focuses on implementation, automation, and optimization of production inference systems, tuning model serving runtimes for performance and scale across NVIDIA and AMD GPU fleets, working under the technical direction of senior engineering leadership and the platform's established architectural patterns.

You'll ship well-engineered, well-tested infrastructure changes and grow your depth in GPU-backed workloads, distributed model serving, observability, and continuous delivery. You'll work directly with senior engineers on real production systems, receive code and design review on everything you ship, and have a clear path to expanded scope and ownership as your experience deepens.

What You'll Do
  • Deploy, tune, and optimize high-performance LLM inference pipelines on GPU infrastructure, improving throughput, latency, and cost efficiency within established design patterns.
  • Analyze, profile, and optimize model serving workloads across inference frameworks such as vLLM, SGLang, and TensorRT-LLM, and across different model families and hardware architectures.
  • Build and operate scalable, production-grade API services for model inference, including request routing, multi-tenant isolation, usage metering, and observability.
  • Develop benchmarking harnesses, monitoring infrastructure, and automation tooling that make serving performance measurable and reproducible.
  • Scale inference workloads across multi-GPU, multi-node environments spanning NVIDIA and AMD accelerators.
  • Evaluate, prototype, and integrate model fine-tuning workflows and frameworks.
  • Collaborate closely with engineering and product teams to align infrastructure capabilities with customer-facing services.
  • Investigate and resolve issues across the stack, including container, node, network, and accelerator-level problems, escalating appropriately when scope exceeds the role.
  • Write clear documentation, including runbooks, internal references, and design notes for the changes you ship.
  • Participate in code and design reviews, both as author and reviewer, and incorporate feedback from senior engineers into your work.
Required Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, Applied Math, or Data
  • Science, plus three (3) years relevant work experience or equivalent combination of education and relevant experience.
  • Professional software engineering experience, with at least some of it touching ML systems, GPU workloads, or high-performance backend services.
  • Working knowledge of Kubernetes in a production context, including writing and
  • debugging manifests, understanding core resource types, and operating production workloads.
  • Hands-on experience serving or deploying LLMs — you've run vLLM, SGLang, TGI, TensorRT-LLM, or similar.
  • Comfort working in a Linux environment and with standard developer tooling, including Git-based workflows.
  • Familiarity with CI/CD systems and the basic mechanics of automated build, test, and deployment pipelines.
  • Strong proficiency in Python with familiarity at least one programming or scripting language used for infrastructure work (Go, Rust, C++ or Bash).
Preferred Qualifications
  • Experience building or operating retrieval-augmented generation (RAG) pipelines, including vector databases, embedding models, and retrieval serving at scale.
  • AMD/ROCm experience.
  • Experience cleaning and curating datasets for LLM training and fine tuning.
  • Fine-tuning experience of any depth: LoRA/QLoRA, full fine-tunes, dataset curation, or evaluation design.
  • Experience with usage metering, billing systems, or multi-tenant API platforms.
  • Experience instrumenting services and consuming observability data, including writing Prometheus queries, building Grafana dashboards, or working with distributed traces.
  • Experience with HPC batch schedulers and MPI based workloads.
  • Experience with alternative compute architectures for inference (RISC-V, FPGA, ASIC).
Why Join Hoonify

You'll have a direct line to leadership and genuine influence over the company's growth trajectory. This is a rare opportunity to build a cutting-edge multi-cloud computational platform at a company doing meaningful work in AI — with the autonomy and resources to make it your own.

About Our Team

Hoonify delivers secure, sovereign AI infrastructure designed for the next generation of inference workloads. Powered by TurbOS , our platform enables organizations and NeoCloud/data center operators to transform CPU/GPU infrastructure into production-ready AI environments—supporting local LLMs, agentic copilots, RAG, and embeddings. We empower teams with robust model lifecycle management, multi-tenant controls, usage metering, and fully auditable operations.

Hoonify is an equal opportunity employer. We welcome applicants from all backgrounds and are committed to building a diverse and inclusive team.

Must be eligible to obtain and maintain a US government security clearance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Platform Engineer
AI Inference Platform Engineer

Hoonify Technologies Inc. • Albuquerque (NM)

On-site
USD 120,000 - 190,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel Corporation • Austin (TX)

Hybrid
USD 171,000 - 315,000
Competitive pay
Stock bonuses
Health & retirement
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Senior Inference Platform Engineer - Data Center
Senior Inference Platform Engineer - Data Center

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Inference Engineer
Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Office Bellevue
Competitive compensation
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership