VLM Data Science Expert

CitiusTech

San Ramon (CA)

On-site

USD 160,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive set of benefits
Flexible work culture
Professional development opportunities

Job summary

A leading healthcare technology firm in San Ramon is seeking a Senior Data Scientist to develop and implement Vision-Language Models for advanced AI applications in healthcare. The ideal candidate will have over 10 years of experience in machine learning, proficiency in Python, and strong knowledge of AWS services. If you are passionate about transforming healthcare through technology, this role offers an opportunity to make a significant impact.

Qualifications

  • 10+ years of experience in Machine Learning or Data Science roles focused on Vision-Language Models.
  • Proven expertise in deploying production-grade multimodal AI solutions.
  • Experience with medical imaging and processing.

Responsibilities

  • Design, train, and deploy Vision-Language Models for multimodal applications.
  • Develop and manage Digital Twin frameworks using AWS IoT.
  • Implement scalable pipelines for training/testing VLMs on cloud platforms.

Skills

Vision-Language Models (VLMs)
Python
Deep Learning
Machine Learning Frameworks (PyTorch, TensorFlow)
AWS Services

Education

Master’s or Ph.D. in Computer Science, Data Science, Machine Learning, or related field

Tools

AWS SageMaker
CUDA
Docker

Job description

Overview

At CitiusTech, we constantly strive to solve the industry\'s greatest challenges with technology, creativity, and agility. With over 8,500 healthcare technology professionals worldwide, CitiusTech powers healthcare digital innovation, business transformation, and industry-wide convergence for over 140 organizations through next-generation technologies, solutions, and products. We aim to accelerate the transition to a human-first, sustainable, and digital healthcare ecosystem with the world\'s leading Healthcare and life sciences organizations and our partners.

Here is an opportunity for you to make a difference and collaborate with global leaders to shape the future of healthcare and positively impact human lives.

Our vision: To inspire new possibilities for the health ecosystem with technology and human ingenuity.

Base pay range

$160,000.00/yr - $200,000.00/yr

Direct message the job poster from CitiusTech

To learn more about CitiusTech, visitwww.citiustech.com

What is in it for you?

If you\'re a Senior Data Scientist with a strong background in Vision-Language Models (VLMs), this is a chance to lead the charge in building smart, scalable multimodal AI solutions. We’re looking for someone who’s worked hands-on with cutting-edge frameworks like VILA, Isaac, and VSS—and who knows how to take models from concept to production in real-world settings. If you\'ve got experience in healthcare, especially with medical devices, that\'s a big plus. You\'ll be diving into the latest VLM techniques and deploying them on cloud platforms like AWS, helping shape the future of AI in a meaningful, impactful way.

Key Responsibilities
  • Design, train, and deploy efficient Vision-Language Models (e.g., VILA, Isaac Sim) for multimodal applications including image captioning, visual search, and document understanding, pose understanding, pose comparison.
  • Develop and manage Digital Twin frameworks using AWS IoT TwinMaker, SiteWise, and Greengrass to simulate and optimize real-world systems.
  • Develop Digital Avatars using AWS services integrated with 3D rendering engines, animation pipelines, and real-time data feeds.
  • Explore cost-effective methods such as knowledge distillation, modal-adaptive pruning, and LoRA fine-tuning to optimize training and inference.
  • Implement scalable pipelines for training/testing VLMs on cloud platforms (AWS services such as SageMaker, Bedrock, Rekognition, Comprehend, and Textract.)
NVIDIA Platforms
  • Should develop a blend of technical expertise, tool proficiency, and domain- specific knowledge on below NVIDIA Platforms:
  • NeMo Framework: Training and scaling VLMs across thousands of GPUs.
  • DeepStream SDK: Integrates pose models like TRTPose and OpenPose, Real-time video analytics and multi-stream processing.
Multimodal AI Solutions
  • Develop solutions that integrate vision and language capabilities for applications like image-text matching, visual question answering (VQA), and document data extraction.
  • Leverage interleaved image-text datasets and advanced techniques (e.g., cross-attention layers) to enhance model performance.
Image Processing and Computer Vision
  • Develop solutions that integrate Vision based deep learning models for applications like live video streaming integration and processing, object detection, image segmentation, pose Estimation, Object Tracking and Image Classification and defect detection on medical Xray images
  • Knowledge of real-time video analytics, multi-camera tracking, and object detection.
  • Training and testing the deep learning models on customized data
Healthcare Domain Expertise (Nice to Have)
  • While it’s not a must, having experience in the healthcare space—especially with medical imaging, motion detection, or patient monitoring—can be a big advantage.
  • You’ll be applying Vision-Language Models to use cases like analyzing scans, detecting positioning and movement, and making precise measurements.
  • If you\'re familiar with healthcare standards and know how to handle sensitive data responsibly, that’s a definite plus.
  • Evaluate trade-offs between model size, performance, and cost using techniques like elastic visual encoders or lightweight architectures.
  • Benchmark different VLMs (e.g., GPT-4V, Claude 3.5, Nova Lite) for accuracy, speed, and cost-effectiveness on specific tasks.
  • Benchmarking on GPU vs CPU
  • Collaborate with cross-functional teams including engineers and domain experts to define project requirements.
  • Mentor junior team members and provide technical leadership on complex projects.
Experience
  • 10+ Years
Location
  • San Ramon, CA or Milwaukee, WI (Onsite)
Qualifications
  • Education: Master’s or Ph.D. in Computer Science, Data Science, Machine Learning, or a related field.
Experience
  • Minimum of 10+ years of experience in Machine Learning or Data Science roles with a focus on Vision-Language Models.
  • Proven expertise in deploying production-grade multimodal AI solutions.
  • Experience in self driving cars and self navigating robots.
Technical Skills
  • Proficiency in Python and ML frameworks (e.g., PyTorch, TensorFlow).
  • Hands-on experience with VLMs such as VILA, Isaac Sim, or VSS.
  • Familiarity with cloud platforms like AWS SageMaker or Azure ML Studio for scalable AI deployment.
  • CUDA, cuDNN
Domain Knowledge — A Valuable Bonus

It’s helpful if you’ve got a solid grasp of medical datasets, especially imaging data, and an understanding of healthcare regulations.

Knowing how to navigate the complexities of clinical data and compliance can really elevate your impact in this role

Soft Skills
  • Strong problem-solving skills with the ability to optimize models for real-world constraints.
  • Excellent communication skills to explain technical concepts to diverse stakeholders.
Preferred Technologies
  • Multimodal Techniques: Cross-attention layers, interleaved image-text datasets
  • MLOps Tools: Docker, MLflow
Life at CitiusTech

We focus on building highly motivated engineering teams and thought leaders with an entrepreneurial mindset, centered on our core values of Passion, Respect, Openness, Unity, and Depth (PROUD) of knowledge. Our success lies in creating a fun, transparent, non-hierarchical, diverse work culture that focuses on continuous learning and work-life balance.

Rated by our employees as the ‘Great Place to Work’ for’ according to the Great Place to Work survey. We offer you a comprehensive set of benefits to ensure that you have a long and rewarding career with us.

This posting complies with applicable pay transparency laws in states such as California, Colorado, New York, New Jersey, Illinois, and others.

Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status

Our EVP

Be You Be Awesome is our EVP and it reflects our continuing efforts to create CitiusTech as a great place to work where our employees can thrive, both personally and professionally. It encompasses the unique benefits and opportunities we offer to support your growth, well-being, and success throughout your journey with us and beyond. Together with our clients, we are solving some of the greatest healthcare challenges and positively impacting human lives. Welcome to the world of Faster Growth, Higher Learning, and Stronger Impact.

Join CitiusTech. Be You. Be Awesome.

To learn more about CitiusTech, visit www.citiustech.com

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Vision-Language Models (VLMs)
Vision-Language Models (VLMs)

TalentOla • Waukesha (WI)

On-site
USD 120,000 - 150,000
Senior/Principal Visual ML Engineer
Senior/Principal Visual ML Engineer

Socket.dev • Charleston (SC)

On-site
USD 150,000 - 260,000
AI Engineer
AI Engineer

bpd • United States

On-site
USD 110,000 - 170,000
Staff Data Engineer
Staff Data Engineer

LVT (LiveView Technologies) • Seattle (WA)

On-site
USD 171,000 - 221,000
Comprehensive health, dental, and vision coverage
401k match up to 4%
Flexible PTO
+1
Senior ML Engineer: Healthcare LLMs & Data Pipelines
Senior ML Engineer: Healthcare LLMs & Data Pipelines

C the Signs • Boston (MA)

Hybrid
USD 120,000 - 150,000
Competitive salary and benefits
Flexible working arrangements
Continuous learning opportunities
Computer VisionML Engineer
Computer VisionML Engineer

Norbert Health • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and equity
High autonomy and technical ownership
Mission-driven culture focused on learning
Applied AI Scientist
Applied AI Scientist

Vantor • United States

On-site
USD 140,000 - 216,000
401(k) with company match
Mental health resources
Student loan repayment assistance
+2
Sr AI/ML Engineer
Sr AI/ML Engineer

Vizient • Centennial (CO)

On-site
USD 107,000 - 179,000
Benefits package
Research, Vision Expertise
Research, Vision Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Videa Health, Inc. • Town of Boston (NY)

On-site
USD 110,000 - 150,000
Flexible PTO
Competitive pay and equity
Collaborative work culture