Vision-Language Models (VLMs)

TalentOla

Waukesha (WI)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading tech recruitment company is seeking a Senior Data Scientist with expertise in Vision-Language Models. This role involves designing and deploying AI solutions and developing Digital Twin frameworks using AWS services. The ideal candidate will have over 10 years of experience, a master's or Ph.D. in a related field, and proficiency in Python and machine learning frameworks. This opportunity is based in Waukesha or San Ramon with a focus on healthcare applications.

Qualifications

  • 10+ years of experience in machine learning or data science roles.
  • Proven expertise in deploying production-grade multimodal AI solutions.
  • Experience in self-driving cars and self-navigating robots.

Responsibilities

  • Design and deploy Vision-Language Models for multimodal applications.
  • Develop Digital Twin frameworks using AWS services.
  • Collaborate with cross-functional teams to define project requirements.

Skills

Proficiency in Python
Strong problem-solving skills
Excellent communication
Experience with Vision-Language Models

Education

Master's or Ph.D. in Computer Science or related field

Tools

PyTorch
TensorFlow
AWS SageMaker
CUDA
cuDNN
Docker
MLflow

Job description

2 weeks ago Be among the first 25 applicants

Get AI-powered advice on this job and more exclusive features.

Direct message the job poster from TalentOla

Overview

Role: Senior Data Scientist with expertise in Vision-Language Models (VLMs)

Experience: 10+ Years

Location: San Ramon, CA or Waukesha, WI (Onsite)

Responsibilities
  • Design, train, and deploy efficient Vision-Language Models (e.g., VILA, Isaac Sim) for multimodal applications including image captioning, visual search, and document understanding, pose understanding, pose comparison.
  • Develop and manage Digital Twin frameworks using AWS IoT TwinMaker, SiteWise, and Greengrass to simulate and optimize real-world systems.
  • Develop Digital Avatars using AWS services integrated with 3D rendering engines, animation pipelines, and real-time data feeds.
  • Explore cost-effective methods such as knowledge distillation, modal-adaptive pruning, and LoRA fine-tuning to optimize training and inference.
  • Implement scalable pipelines for training/testing VLMs on cloud platforms (AWS services such as SageMaker, Bedrock, Rekognition, Comprehend, and Textract).

Candidate should develop a blend of technical expertise, tool proficiency, and domain-specific knowledge on below NVIDIA Platforms:

  • NIM (NVIDIA Inference Microservices): Containerized VLM deployment.
  • NeMo Framework: Training and scaling VLMs across thousands of GPUs.
  • DeepStream SDK: Integrates pose models like TRTPose and OpenPose, Real-time video analytics and multi-stream processing.
  • Multimodal AI Solutions: Develop solutions that integrate vision and language capabilities for applications like image-text matching, visual question answering (VQA), and document data extraction.
  • Image Processing and Computer Vision: Develop solutions that integrate Vision based deep learning models for applications like live video streaming integration and processing, object detection, image segmentation, pose estimation, object tracking and image classification and defect detection on medical X-ray images.
  • Knowledge of real-time video analytics, multi-camera tracking, and object detection.
  • Training and testing the deep learning models on customized data.
  • Apply VLMs to healthcare-specific use cases such as medical imaging analysis, position detection, motion detection and measurements.
  • Ensure compliance with healthcare standards while handling sensitive data.
  • Evaluate trade-offs between model size, performance, and cost using techniques like elastic visual encoders or lightweight architectures.
  • Benchmark different VLMs (e.g., GPT-4V, Claude 3.5, Nova Lite) for accuracy, speed, and cost-effectiveness on specific tasks.
  • Benchmarking on GPU vs CPU.
  • Collaborate with cross-functional teams including engineers and domain experts to define project requirements.
  • Mentor junior team members and provide technical leadership on complex projects.
Qualifications
  • Education: Master’s or Ph.D. in Computer Science, Data Science, Machine Learning, or a related field.
  • Experience: Minimum of 10+ years of experience in machine learning or data science roles with a focus on vision-language models.
  • Proven expertise in deploying production-grade multimodal AI solutions.
  • Experience in self-driving cars and self-navigating robots.
  • Technical Skills: Proficiency in Python and ML frameworks (e.g., PyTorch, TensorFlow).
  • Hands-on experience with VLMs such as VILA, Isaac Sim, or VSS.
  • Familiarity with cloud platforms like AWS SageMaker or Azure ML Studio for scalable AI deployment.
  • CUDA, cuDNN
  • Domain Knowledge: Understanding of medical datasets (e.g., imaging data) and healthcare regulations.
  • Soft Skills: Strong problem-solving skills with the ability to optimize models for real-world constraints; Excellent communication skills to explain technical concepts to diverse stakeholders.
Preferred Technologies
  • Multimodal Techniques: Cross-attention layers, interleaved image-text datasets
  • MLOps Tools: Docker, MLflow
Job function
  • Information Technology
Industries
  • Information Services
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VLM Data Science Expert
VLM Data Science Expert

CitiusTech • San Ramon (CA)

On-site
USD 160,000 - 200,000
Comprehensive set of benefits
Flexible work culture
Professional development opportunities
Member of Technical Staff, Machine Learning - NomadicML
Member of Technical Staff, Machine Learning - NomadicML

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Member of Technical Staff, Machine Learning - NomadicML
Member of Technical Staff, Machine Learning - NomadicML

Pear VC • Austin (TX), California (MO)

On-site
USD 120,000 - 150,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
Senior/Principal Visual ML Engineer
Senior/Principal Visual ML Engineer

Socket.dev • Charleston (SC)

On-site
USD 150,000 - 260,000
Computer Vision Engineer
Computer Vision Engineer

OVA.Work • New York (NY)

On-site
USD 120,000 - 170,000
Senior Deep Learning Engineering - Autonomous Vehicles
Senior Deep Learning Engineering - Autonomous Vehicles

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
Machine Learning Engineer - Computer Vision
Machine Learning Engineer - Computer Vision

Dormont Manufacturing Co • Arlington (TX)

On-site
USD 90,000 - 120,000
Competitive Salary
Stock Option
Medical, Dental, and Vision Insurance
+4
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation

Jobtailor • Redwood City (CA)

On-site
USD 180,000 - 260,000
Machine Learning Engineer
Machine Learning Engineer

Escalon • Santa Monica (CA)

On-site
USD 100,000 - 120,000
Comprehensive health coverage
Flexible PTO
Collaborative, intellectually driven 팀