AI Inference Platform Engineer

Hoonify Technologies Inc.

Albuquerque (NM)

On-site

USD 120,000 - 190,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hoonify Technologies Inc. seeks an AI Infrastructure Engineer to build, deploy, and operate the LLM serving infrastructure of our AI/ML platform. You will optimize production inference systems, tune runtimes, and scale across NVIDIA and AMD GPUs.

You will work with senior engineers, ship well-tested changes, and grow expertise in GPU-backed workloads, observability, and continuous delivery, with opportunities for broader ownership as you prove your impact.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Applied Math, or Data Science.
  • 3 years of relevant work experience or equivalent combination of education and relevant experience.
  • Professional software engineering experience touching ML systems, GPU workloads, or high-performance backend services.
  • Working knowledge of Kubernetes in production contexts, including writing and debugging manifests.
  • Hands-on experience serving or deploying LLMs (vLLM, SGLang, TensorRT-LLM, or similar).

Responsibilities

  • Deploy, tune, and optimize high-performance LLM inference pipelines on GPU infrastructure.
  • Analyze, profile, and optimize model serving workloads across frameworks and hardware.
  • Build and operate scalable, production-grade API services for model inference.
  • Develop benchmarking harnesses, monitoring infrastructure, and automation tooling.
  • Scale inference workloads across multi-GPU, multi-node environments.
  • Collaborate with engineering and product teams to align infrastructure with services.
  • Write runbooks, design notes, and documentation for changes.
  • Participate in code/design reviews and incorporate feedback from seniors.

Skills

Kubernetes
Python
Linux
ML systems
CI/CD
Git workflows
LLM deployments
Go/Rust/C++

Education

Bachelor's degree in CS/CE/Applied Math/Data Science

Tools

vLLM
SGLang
TensorRT-LLM
Git
Grafana

Job description

Hoonify Technologies Inc. seeks an AI Infrastructure Engineer to build, deploy, and operate the LLM serving infrastructure of our AI/ML platform. You will optimize production inference systems, tune runtimes, and scale across NVIDIA and AMD GPUs.

You will work with senior engineers, ship well-tested changes, and grow expertise in GPU-backed workloads, observability, and continuous delivery, with opportunities for broader ownership as you prove your impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer - Inference Platform
AI Infrastructure Engineer - Inference Platform

Hoonify Technologies Inc. • Albuquerque (NM)

On-site
USD 120,000 - 190,000
Principal AI Inference Platform Engineer - GPU/K8s
Principal AI Inference Platform Engineer - GPU/K8s

Lila Sciences • Cambridge (ME)

On-site
USD 192,000 - 272,000
Medical, dental, and vision
Life and disability insurance
Flexible time off
+4
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
LLMOps Platform Engineer for GPU AI Inference
LLMOps Platform Engineer for GPU AI Inference

Cloud Analytics Technologies, LLC • Jersey City (NJ)

On-site
USD 130,000 - 160,000
AI Infra Architect — GPU HPC & Cloud/On‑Prem
AI Infra Architect — GPU HPC & Cloud/On‑Prem

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Staff AI Inference Engineer — Remote/Hybrid
Staff AI Inference Engineer — Remote/Hybrid

Zoom • Seattle (WA), Northern (KY)

Hybrid
USD 152,000 - 332,000
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000