Machine Learning Engineer (Video Understanding & Segmentation)

Glint Tech Solutions

Santa Clara (CA)

On-site

USD 140,000 - 210,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Glint Tech Solutions in Santa Clara, CA seeks a highly skilled Machine Learning Engineer to advance video understanding and segmentation. You will build scalable systems that convert raw egocentric and human-robot video into structured, searchable data across embedding, captioning, retrieval, and LLM reasoning stages.

The role emphasizes multi-modal learning, large-scale video pipelines, and agentic orchestration, with collaboration across teams to drive production-ready capabilities.

Qualifications

  • MS or PhD in CS/EE or related field, or equivalent practical experience.
  • 3+ years of hands-on experience in computer vision or multi-modal machine learning.
  • Strong proficiency in Python and PyTorch with solid software engineering fundamentals.
  • Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.
  • Experience building or fine-tuning LLM-based systems for video/image understanding (captioning, video QA, summarization).
  • Familiarity with agentic system design—tool use, multi-step reasoning, and orchestration frameworks.
  • Experience with large-scale video data pipelines and vector search/retrieval infrastructure (FAISS, Milvus).

Responsibilities

  • Build and optimize video/image embedding pipelines for large-scale video search and retrieval.
  • Develop LLM-based video understanding systems for indexing, summarization, and QA over long-form video.
  • Design segmentation algorithms to decompose long videos into structured clips.
  • Build automated captioning systems combining vision-language models and LLMs.
  • Architect agentic pipelines chaining embedding, captioning, retrieval, and LLM reasoning steps.
  • Develop and scale video search infrastructure (vector indexing, retrieval, ranking) for multi-modal queries.
  • Collaborate with annotation, data engineering, and robotics teams to integrate outputs into training pipelines.
  • Evaluate and benchmark embedding models, LLMs, and agentic frameworks; track frontier research.

Skills

Python
PyTorch
Computer vision
Multimodal ML
Video understanding
Vector search
LangChain
LLM integration

Education

MS or PhD in CS/EE or related

Tools

FAISS
Milvus

Job description

The Role

We are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale — turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.

Responsibilities
  • Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.
  • Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.
  • Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.
  • Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.
  • Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.
  • Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.
  • Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.
  • Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.
  • Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.
  • Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.
Minimum Qualifications
  • MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
  • 3+ years of hands‑on experience in computer vision or multi‑modal machine learning, with direct experience in video understanding tasks.
  • Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.
  • Hands‑on experience with CLIP or similar vision‑language/video embedding models for retrieval or representation learning.
  • Experience building or fine‑tuning LLM‑based systems for video/image understanding (e.g., captioning, video QA, summarization).
  • Familiarity with agentic system design — tool use, multi‑step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).
  • Experience working with large‑scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).
Preferred Qualifications
  • PhD with a research focus in video understanding, multi‑modal learning, or vision‑language models.
  • Experience with temporal action segmentation, action localization, or instruction‑level video chunking algorithms.
  • Experience working with egocentric video datasets or head‑mounted‑device (HMD) captured data.
  • Track record of deploying production‑scale video search or retrieval systems.
  • Experience integrating foundation or vision‑language models (e.g., CLIP, VideoCLIP, RT‑1/VLA variants) into perception or decision‑making pipelines.
  • Publications in top‑tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).
  • Experience with humanoid robotics or embodied AI data pipelines is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer – Video AI, Vision & Creative Systems
Senior Machine Learning Engineer – Video AI, Vision & Creative Systems

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Machine Learning Engineer
Machine Learning Engineer

Escalon • Santa Monica (CA)

On-site
USD 100,000 - 120,000
Comprehensive health coverage
Flexible PTO
Collaborative, intellectually driven 팀
Machine Learning Engineer
Machine Learning Engineer

Escalon Services, Inc. • Santa Monica (CA)

On-site
USD 100,000 - 120,000
Health coverage
Flexible PTO
AI/robotics projects
Machine Learning Engineer - Computer Vision
Machine Learning Engineer - Computer Vision

Dormont Manufacturing Co • Arlington (TX)

On-site
USD 90,000 - 120,000
Competitive Salary
Stock Option
Medical, Dental, and Vision Insurance
+4
Video Understanding & Segmentation ML Engineer
Video Understanding & Segmentation ML Engineer

Glint Tech Solutions • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Research Scientist, Video Understanding & World Models
Research Scientist, Video Understanding & World Models

Kindredventures • New York (NY)

On-site
USD 120,000 - 150,000
Senior Machine Learning Engineer, Content Engineering
Senior Machine Learning Engineer, Content Engineering

Paramount • New York (NY)

On-site
USD 130,000 - 160,000
Bonus eligibility
Diversity-focused workplace
Research Scientist, Video Understanding & World Models
Research Scientist, Video Understanding & World Models

Mecka AI • New York (NY)

On-site
USD 100,000 - 130,000
Member of Technical Staff, Vision/Language
Member of Technical Staff, Vision/Language

XDOF • San Mateo (CA)

On-site
USD 130,000 - 210,000
Member of Technical Staff, Vision / Language
Member of Technical Staff, Vision / Language

xdof.ai • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive compensation and equity
Comprehensive health and wellness benefits
Collaborative and fast-paced work environment