Multimodal Intelligence Systems Engineer

MaxIT Consulting - Max Corporate Group

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MaxIT Consulting - Max Corporate Group in the San Francisco Bay Area is seeking a Multimodal Intelligence Systems Engineer to own critical AI components, deploying vision-language systems for real-world industrial workflows and production environments. The role emphasizes practical delivery, end-to-end ownership, and on-site work in Burlingame transitioning to San Francisco.

The ideal candidate ships real systems, works with vision-language techniques, and thrives in hardware-constrained,

Qualifications

  • Approximately 1 to 5 years of relevant professional experience.
  • Hands‑on experience shipping computer vision, multimodal, or vision‑language systems to production.
  • Strong practical understanding of visual AI applied to images and/or video.
  • Experience owning technical systems end to end rather than working exclusively on research prototypes.
  • Comfort operating in real-time, edge, hardware-constrained, or connectivity-constrained environments.
  • Strong software engineering fundamentals and a high level of technical ownership.

Responsibilities

  • Build and deploy agentic vision-language systems for real-world industrial workflows.
  • Develop multimodal reasoning capabilities across images and video.
  • Work with detection, segmentation, visual reasoning, VLMs, and related applied computer vision techniques.
  • Own model orchestration and production deployment across hardware-constrained environments.
  • Design and maintain evaluation systems including ground-truth datasets, trajectory evaluation, and regression testing.
  • Iterate on production models based on real-world usage and measurable system performance.

Skills

Computer vision
Multimodal systems
Edge computing
Production deployment
Software engineering
Real-time systems

Job description

SF Bay Area, California | Fully On-site

We are seeking an Multimodal Intelligence Systems Engineer to join an early-stage technology company building AI systems for real-world industrial environments. This role focuses on multimodal visual intelligence running in hardware-constrained environments and supporting complex physical workflows.

The Opportunity

You will own critical parts of the AI layer, building and deploying vision-language systems that reason over images and video in real customer environments. The role combines applied AI, computer vision, evaluation systems, and production engineering.

Key Responsibilities
  • Build and deploy agentic vision-language systems for real-world industrial workflows.
  • Develop multimodal reasoning capabilities across images and video.
  • Work with detection, segmentation, visual reasoning, VLMs, and related applied computer vision techniques.
  • Own model orchestration and production deployment across hardware-constrained environments.
  • Design and maintain evaluation systems including ground-truth datasets, trajectory evaluation, and regression testing.
  • Iterate on production models based on real-world usage and measurable system performance.
Required Qualifications
  • Approximately 1 to 5 years of relevant professional experience.
  • Hands‑on experience shipping computer vision, multimodal, or vision‑language systems to production.
  • Strong practical understanding of visual AI applied to images and/or video.
  • Experience owning technical systems end to end rather than working exclusively on research prototypes.
  • Comfort operating in real-time, edge, hardware-constrained, or connectivity-constrained environments.
  • Strong software engineering fundamentals and a high level of technical ownership.
Candidate Profile

The strongest candidates are practical builders who have shipped real systems with real users. Experience in robotics, autonomous systems, drones, wearables, real-time video, edge computing, or adjacent physical-world AI domains is highly relevant.

Work Arrangement

This position is fully in person in the San Francisco Bay Area. The initial working location is in Burlingame, with the team expected to transition to a San Francisco office.

Work Authorization

Candidates must be able to work in the United States. Certain visa sponsorship, transfer, or change-of-status scenarios may be supported for candidates already located in the United States. New petitions requiring processing from abroad are not supported at this time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Multimodal Intelligence Systems Engineer
Multimodal Intelligence Systems Engineer

MaxIT Consulting - Max Corporate Group • California (MO)

On-site
USD 120,000 - 180,000
Applied AI Engineer – Computer Vision & VLMs
Applied AI Engineer – Computer Vision & VLMs

MaxIT Consulting - Max Corporate Group • California (MO)

On-site
USD 120,000 - 200,000
Applied AI Engineer – Computer Vision & VLMs
Applied AI Engineer – Computer Vision & VLMs

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 180,000
Edge Vision & Multimodal AI Engineer
Edge Vision & Multimodal AI Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 180,000
Industrial Edge Vision & Language Systems Engineer
Industrial Edge Vision & Language Systems Engineer

MaxIT Consulting - Max Corporate Group • California (MO)

On-site
USD 120,000 - 180,000
AI/ML Engineer (Computer Vision)
AI/ML Engineer (Computer Vision)

Blue-Signal-Search • San Francisco (CA)

On-site
USD 120,000 - 170,000
Research, Vision Expertise
Research, Vision Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Multimodal Perception & Authentication ML Engineer
Multimodal Perception & Authentication ML Engineer

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Relocation assistance
Junior AI/ML Engineer (Vision + Multimodal) On-site (SF)
Junior AI/ML Engineer (Vision + Multimodal) On-site (SF)

Blueprints AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Staff Engineer - Multimodal Vision & Video
Staff Engineer - Multimodal Vision & Video

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000