Video Understanding & Segmentation ML Engineer

Glint Tech Solutions

Santa Clara (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Glint Tech Solutions in Santa Clara, CA seeks a highly skilled Machine Learning Engineer to advance video understanding and segmentation. You will build scalable systems that convert raw egocentric and human-robot video into structured, searchable data across embedding, captioning, retrieval, and LLM reasoning stages.

The role emphasizes multi-modal learning, large-scale video pipelines, and agentic orchestration, with collaboration across teams to drive production-ready capabilities.

Qualifications

  • MS or PhD in CS/EE or related field, or equivalent practical experience.
  • 3+ years of hands-on experience in computer vision or multi-modal machine learning.
  • Strong proficiency in Python and PyTorch with solid software engineering fundamentals.
  • Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.
  • Experience building or fine-tuning LLM-based systems for video/image understanding (captioning, video QA, summarization).
  • Familiarity with agentic system design—tool use, multi-step reasoning, and orchestration frameworks.
  • Experience with large-scale video data pipelines and vector search/retrieval infrastructure (FAISS, Milvus).

Responsibilities

  • Build and optimize video/image embedding pipelines for large-scale video search and retrieval.
  • Develop LLM-based video understanding systems for indexing, summarization, and QA over long-form video.
  • Design segmentation algorithms to decompose long videos into structured clips.
  • Build automated captioning systems combining vision-language models and LLMs.
  • Architect agentic pipelines chaining embedding, captioning, retrieval, and LLM reasoning steps.
  • Develop and scale video search infrastructure (vector indexing, retrieval, ranking) for multi-modal queries.
  • Collaborate with annotation, data engineering, and robotics teams to integrate outputs into training pipelines.
  • Evaluate and benchmark embedding models, LLMs, and agentic frameworks; track frontier research.

Skills

Python
PyTorch
Computer vision
Multimodal ML
Video understanding
Vector search
LangChain
LLM integration

Education

MS or PhD in CS/EE or related

Tools

FAISS
Milvus

Job description

Glint Tech Solutions in Santa Clara, CA seeks a highly skilled Machine Learning Engineer to advance video understanding and segmentation. You will build scalable systems that convert raw egocentric and human-robot video into structured, searchable data across embedding, captioning, retrieval, and LLM reasoning stages.

The role emphasizes multi-modal learning, large-scale video pipelines, and agentic orchestration, with collaboration across teams to drive production-ready capabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer (Video Understanding & Segmentation)
Machine Learning Engineer (Video Understanding & Segmentation)

Glint Tech Solutions • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Senior Machine Learning Engineer – Video AI, Vision & Creative Systems
Senior Machine Learning Engineer – Video AI, Vision & Creative Systems

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Video AI Engineer - Creative Systems & Vision
Video AI Engineer - Creative Systems & Vision

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
ML Engineer – Vision-Language for Motion & Autonomy
ML Engineer – Vision-Language for Motion & Autonomy

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
In-Person ML Engineer: Vision & Embedding Systems
In-Person ML Engineer: Vision & Embedding Systems

Escalon Services, Inc. • Santa Monica (CA)

On-site
USD 100,000 - 120,000
Health coverage
Flexible PTO
AI/robotics projects
Machine Learning Engineer
Machine Learning Engineer

Escalon • Santa Monica (CA)

On-site
USD 100,000 - 120,000
Comprehensive health coverage
Flexible PTO
Collaborative, intellectually driven 팀
Vision & Embeddings ML Engineer (Real-Time)
Vision & Embeddings ML Engineer (Real-Time)

Escalon • Santa Monica (CA)

On-site
USD 100,000 - 120,000
Comprehensive health coverage
Flexible PTO
Collaborative, intellectually driven 팀
Senior Video AI Engineer: Production ML for Video
Senior Video AI Engineer: Production ML for Video

LinkedIn • California (MO)

Hybrid
USD 144,000 - 236,000
Health and wellness programs
Stock options
Annual bonus
ML Video Engineer — Edge AI for Image & Video
ML Video Engineer — Edge AI for Image & Video

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical coverage
Retirement benefits
Educational reimbursement
+1
Senior Video AI Engineer — Production ML & Personalization
Senior Video AI Engineer — Production ML & Personalization

Jobzhr • Mountain View (CA)

Hybrid
USD 144,000 - 236,000