Multimodal Inference Engineer — Scale GPU AI Models
OpenAI
San Francisco (CA)
On-site
USD 310,000 - 460,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If you thrive in fast-paced environments and enjoy tackling complex challenges, this opportunity offers a chance to make a significant impact in the AI landscape.
Qualifications
Experience with multimodal models and inference systems.
Understanding of GPU performance dynamics with complex data.
Responsibilities
Design and implement infrastructure for multimodal models.
Optimize systems for low-latency audio and image processing.
Collaborate with researchers and product teams.
Skills
Inference systems for LLMs
GPU-based ML workloads
Networking
Distributed compute
High-throughput data handling
Experimental research collaboration
Tools
vLLM
TensorRT-LLM
Custom model parallel systems
Job description
An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If you thrive in fast-paced environments and enjoy tackling complex challenges, this opportunity offers a chance to make a significant impact in the AI landscape.