Principal Inference Software Engineer

Optomi

Plano (TX)

On-site

USD 170,000 - 250,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Optomi, in partnership with a large telecommunications company, seeks a Principal Inference Software Engineer to join their Plano, TX team. You will build, deploy, and scale production AI inference infrastructure across Kubernetes, vLLM, Python, GPUs, and cloud/on-prem environments.

In this role you will optimize model serving, manage GPU resources, and develop automation to support both R&D and production workloads. A strong background in distributed systems and reliability is essential.

Qualifications

  • Strong software engineering and Python skills.
  • Production experience with LLM inference/model serving, ideally vLLM.
  • Experience with GPU infrastructure, containers, and CI/CD.
  • Understanding of scaling, automation, observability, and reliability for distributed systems.
  • Experience with cloud and/or on-prem infrastructure preferred.

Responsibilities

  • Build and operate LLM inference infrastructure at scale.
  • Deploy models with vLLM and Kubernetes.
  • Develop Python automation for deployment, scaling, updates, rollbacks, and health monitoring.
  • Manage GPU capacity, performance, reliability, and production workloads.
  • Build tools to simplify model deployment for engineers.
  • Support R&D and production inference environments.
  • Troubleshoot and optimize model-serving performance.

Skills

Python
Software engineering
LLM inference
Model serving
vLLM
GPU infrastructure
Containers
CI/CD
Observability
Distributed systems
Cloud/On-prem

Tools

Kubernetes
Docker
CI/CD tooling

Job description

Optomi, in partnership with a large telecommunications company, is seeking a Principal Inference Software Engineer to join their team in the Plano, TX Area. The ideal candidate will have experience building, deploy, and scale production AI inference infrastructure. You’ll work across Kubernetes, vLLM, Python, GPUs, and cloud/on-prem environments to make model serving reliable, scalable, and highly automated.

What The Right Professional Will Enjoy
  • Building and operating LLM inference infrastructure at scale.
  • Deploying and managing models using vLLM and Kubernetes.
  • Developing Python automation for model deployment, scaling, updates, rollbacks, and health monitoring.
  • Managing GPU capacity, performance, reliability, and production workloads.
  • Building tools that allow engineering teams to deploy and access models without managing underlying infrastructure.
  • Supporting both R&D and production inference environments.
  • Troubleshooting and optimizing model-serving performance and reliability.
Apply Today If Your Background Includes
  • Strong software engineering and Python skills.
  • Production experience with LLM inference/model serving, ideally vLLM.
  • Experience with GPU infrastructure, containers, and CI/CD.
  • Understanding of scaling, automation, observability, and reliability for distributed systems.
  • Experience with cloud and/or on-prem infrastructure preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Inference Platform Engineer
Principal AI Inference Platform Engineer

Optomi • Plano (TX)

On-site
USD 170,000 - 250,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development