Software Engineer, AI Infrastructure – LVM Inference & Evaluation

Jobtailor

Redwood City (CA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Jobtailor is building scalable AI infrastructure to support real‑time computer vision and multimodal inference, from data collection to deployment. The platform focuses on efficient model serving, benchmarking, and continuous improvement across large video and sensor data streams.

The team collaborates with researchers, product engineers, and infrastructure teams to optimize batching, caching, quantization, and GPU utilization for production readiness.

Qualifications

  • 2+ years of industry experience building infrastructure, distributed systems, machine learning platforms, or production AI systems.
  • BS/MS in Computer Science or a related technical field, or equivalent practical experience.
  • Strong programming background, especially in Python, with solid software engineering fundamentals.
  • Experience designing and building scalable machine learning infrastructure for training, inference, evaluation, and deployment.
  • Hands‑on experience running deep learning models in production, ideally including LLMs, LVMs, vision‑language models, or multimodal models.
  • Strong understanding of inference optimization, including batching, caching, quantization, parallelism, memory optimization, GPU utilization, and latency reduction.
  • Experience with model-serving frameworks such as vLLM, Triton Inference Server, or similar technologies.
  • Experience building evaluation frameworks, test harnesses, benchmarks, regression tests, or model‑quality measurement systems.
  • Strong background in machine learning and deep learning.
  • Experience designing data engines or pipelines for training and evaluation data.
  • Familiarity with LLMs, LVMs, RAG pipelines, embedding models, or multimodal models in production applications.
  • Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU‑based workloads.
  • Nice to have: CUDA, NCCL, PyTorch, TensorRT, ONNX, video understanding, model compression, speculative decoding, distillation, pruning, prompt evaluation, vector databases, re‑rankers, search infrastructure, or internal ML platforms.
  • Strong collaboration and communication skills.
  • Proactive problem‑solving ability, ownership mindset, and adaptability.

Responsibilities

  • Design, build, and maintain AI infrastructure for real‑time computer vision, LLM, LVM, and multimodal workloads.
  • Develop evaluation harnesses and benchmarking systems for model quality and system performance.
  • Collaborate with researchers and engineers to productionize advances and optimize serving architectures.
  • Improve data engines, observability, monitoring, and debugging tools for production AI systems.

Skills

Python Programming
Distributed Systems
ML Infrastructure
Model Serving
GPU Utilization
Cloud Infrastructure
CUDA
PyTorch
TensorRT
ONNX
Containers
Orchestration
Distributed Systems
Video Understanding

Education

BS/MS in Computer Science or related field

Tools

VLLM
Triton Inference Server
CUDA
PyTorch
TensorRT
ONNX
Containers
Orchestration

Job description

Design, build, and maintain AI infrastructure for real‑time computer vision, LLM, LVM, and multimodal inference workloads. Build scalable systems for state‑of‑the‑art models across large volumes of video and sensor data. Optimize inference for latency, throughput, GPU utilization, reliability, and cost. Develop evaluation harnesses and benchmarking systems for model quality, system performance, regressions, and production readiness. Build infrastructure for continuous model evaluation, experimentation, and deployment. Partner with research scientists to productionize advances in computer vision, LLMs, LVMs, RAG, and multimodal AI. Improve model-serving architecture through batching, caching, routing, quantization, model parallelism, and hardware utilization. Develop data engines and feedback loops for training data collection, model behavior evaluation, and continuous AI improvement. Create observability, monitoring, and debugging tools for production AI systems. Define best practices for deploying, evaluating, and operating AI systems in enterprise environments. Collaborate with research scientists, product engineering, infrastructure teams, and stakeholders.

Requirements
  • 2+ years of industry experience building infrastructure, distributed systems, machine learning platforms, or production AI systems
  • BS/MS in Computer Science or a related technical field, or equivalent practical experience
  • Strong programming background, especially in Python, with solid software engineering fundamentals
  • Experience designing and building scalable machine learning infrastructure for training, inference, evaluation, and deployment
  • Hands‑on experience running deep learning models in production, ideally including LLMs, LVMs, vision‑language models, or multimodal models
  • Strong understanding of inference optimization, including batching, caching, quantization, parallelism, memory optimization, GPU utilization, and latency reduction
  • Experience with model-serving frameworks such as vLLM, Triton Inference Server, or similar technologies
  • Experience building evaluation frameworks, test harnesses, benchmarks, regression tests, or model‑quality measurement systems
  • Strong background in machine learning and deep learning
  • Experience designing data engines or pipelines for training and evaluation data
  • Familiarity with LLMs, LVMs, RAG pipelines, embedding models, or multimodal models in production applications
  • Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU‑based workloads
  • Nice to have: large‑scale GPU infrastructure, CUDA, NCCL, PyTorch, TensorRT, ONNX, video understanding, model compression, speculative decoding, distillation, pruning, prompt evaluation, vector databases, re‑rankers, search infrastructure, or internal ML platforms
  • Strong collaboration and communication skills
  • Proactive problem‑solving ability, ownership mindset, and adaptability
Core Competencies

Demonstrates expertise in building and optimizing AI infrastructure for real‑time computer vision and multimodal inference workloads, with a strong focus on model evaluation, deployment, and performance optimization. Proficient in collaborating with cross‑functional teams to enhance model‑serving architectures and ensure production readiness.

Highest‑signal resume keywords
  • Python Programming
  • Machine Learning Infrastructure
  • Inference Optimization
  • Model‑Serving Frameworks
  • Cloud Infrastructure
ATS Optimization Keywords
Hard Skills
  • AI Infrastructure Design
  • Deep Learning Models
  • Scalable Systems Development
  • Evaluation Frameworks
  • Data Engine Design
  • Batching and Caching
  • Quantization Techniques
  • GPU Utilization
  • Model Quality Measurement
  • Production AI Systems
Soft Skills
  • Collaboration
  • Communication
  • Problem‑Solving
  • Ownership Mindset
  • Adaptability
Industry Keywords
  • Computer Vision
  • LLMs
  • LVMs
  • Multimodal AI
  • RAG Pipelines
  • Embedding Models
  • Model Compression
  • Speculative Decoding
  • Pruning
  • Vector Databases
Tools & Technologies
  • VLLM
  • Triton Inference Server
  • CUDA
  • PyTorch
  • TensorRT
  • ONNX
  • Containers
  • Orchestration
  • Distributed Systems
  • Video Understanding
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Software Engineer
Staff AI Software Engineer

Jobtailor • California (MO)

On-site
USD 170,000 - 210,000
Junior AI/ML Engineer
Junior AI/ML Engineer

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Jobtailor • North Carolina

On-site
USD 150,000 - 190,000
AI/ML Engineer
AI/ML Engineer

Jobtailor • Colorado

On-site
USD 130,000 - 180,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Jobtailor • New York (NY)

On-site
USD 120,000 - 160,000
Technical Program Manager, Model Performance
Technical Program Manager, Model Performance

Jobtailor • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

Jobtailor • North Carolina

Hybrid
USD 150,000 - 210,000
AI/ML Engineer, Precision Oncology
AI/ML Engineer, Precision Oncology

Jobtailor • Tampa (FL)

On-site
USD 120,000 - 180,000
Senior Machine Learning Engineer, Speech, LLM Training Data
Senior Machine Learning Engineer, Speech, LLM Training Data

Jobtailor • Overland Park (KS)

On-site
USD 140,000 - 190,000
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Jobtailor • Town of Texas (WI)

On-site
USD 120,000 - 180,000