Multimodal Inference Engineer — Scale GPU AI Models

OpenAI

San Francisco (CA)

On-site

USD 310,000 - 460,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If you thrive in fast-paced environments and enjoy tackling complex challenges, this opportunity offers a chance to make a significant impact in the AI landscape.

Qualifications

  • Experience with multimodal models and inference systems.
  • Understanding of GPU performance dynamics with complex data.

Responsibilities

  • Design and implement infrastructure for multimodal models.
  • Optimize systems for low-latency audio and image processing.
  • Collaborate with researchers and product teams.

Skills

Inference systems for LLMs
GPU-based ML workloads
Networking
Distributed compute
High-throughput data handling
Experimental research collaboration

Tools

vLLM
TensorRT-LLM
Custom model parallel systems

Job description

An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If you thrive in fast-paced environments and enjoy tackling complex challenges, this opportunity offers a chance to make a significant impact in the AI landscape.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Realtime Multimodal Inference Architect
Realtime Multimodal Inference Architect

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Senior AI Inference Engineer - Multimodal Large Models
Senior AI Inference Engineer - Multimodal Large Models

TikTok • San Jose (CA)

On-site
USD 212,000 - 450,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Sunnyvale (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Engineering Manager, AI Inference & Scale (Hybrid)
Engineering Manager, AI Inference & Scale (Hybrid)

Menlo Ventures • United States

Hybrid
USD 425,000 - 560,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
ML Engineer — AI Platform & Multimodal Inference
ML Engineer — AI Platform & Multimodal Inference

Corvic • Mountain View (CA)

On-site
USD 110,000 - 145,000
Inference Systems Engineer for Scalable Multimodal AI
Inference Systems Engineer for Scalable Multimodal AI

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Socket.dev • Boston (MA)

On-site
USD 168,000 - 227,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA)

On-site
USD 168,000 - 227,000