Multimodal Intelligence Systems Engineer

MaxIT Consulting - Max Corporate Group

California

On-site

USD 120,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

MaxIT Consulting - Max Corporate Group in the SF Bay Area is seeking a Multimodal Intelligence Systems Engineer to build and deploy vision-language systems for real-world industrial environments. This fully on-site role emphasizes edge hardware and production workflows.

You will own critical AI layers, develop multimodal reasoning over images and video, and deliver end-to-end systems with measurable performance in constrained settings.

Qualifications

  • Hands-on experience shipping computer vision or vision-language systems to production.
  • Strong practical understanding of visual AI applied to images/video.
  • Experience owning end-to-end technical systems, not just prototypes.

Responsibilities

  • Build and deploy agentic vision-language systems for real-world industrial workflows.
  • Develop multimodal reasoning across images and video.
  • Work with detection, segmentation, visual reasoning, VLMs and related CV techniques.
  • Own model orchestration and production deployment in hardware-constrained environments.
  • Design and maintain evaluation systems including ground-truth datasets and regression testing.
  • Iterate on production models based on real-world usage and system performance.

Skills

Computer vision
Multimodal systems
Edge deployment
Production engineering
Software ownership

Job description

Location

SF Bay Area, California | Fully On-site

We are seeking an Multimodal Intelligence Systems Engineer to join an early-stage technology company building AI systems for real-world industrial environments. This role focuses on multimodal visual intelligence running in hardware-constrained environments and supporting complex physical workflows.

The Opportunity

You will own critical parts of the AI layer, building and deploying vision-language systems that reason over images and video in real customer environments. The role combines applied AI, computer vision, evaluation systems, and production engineering.

Key Responsibilities
  • Build and deploy agentic vision-language systems for real-world industrial workflows.
  • Develop multimodal reasoning capabilities across images and video.
  • Work with detection, segmentation, visual reasoning, VLMs, and related applied computer vision techniques.
  • Own model orchestration and production deployment across hardware-constrained environments.
  • Design and maintain evaluation systems including ground-truth datasets, trajectory evaluation, and regression testing.
  • Iterate on production models based on real-world usage and measurable system performance.
Required Qualifications
  • Approximately 1 to 5 years of relevant professional experience.
  • Hands-on experience shipping computer vision, multimodal, or vision-language systems to production.
  • Strong practical understanding of visual AI applied to images and/or video.
  • Experience owning technical systems end to end rather than working exclusively on research prototypes.
  • Comfort operating in real-time, edge, hardware-constrained, or connectivity-constrained environments.
  • Strong software engineering fundamentals and a high level of technical ownership.
Candidate Profile

The strongest candidates are practical builders who have shipped real systems with real users. Experience in robotics, autonomous systems, drones, wearables, real-time video, edge computing, or adjacent physical-world AI domains is highly relevant.

Work Arrangement

This position is fully in person in the San Francisco Bay Area. The initial working location is in Burlingame, with the team expected to transition to a San Francisco office.

Work Authorization

Candidates must be able to work in the United States. Certain visa sponsorship, transfer, or change-of-status scenarios may be supported for candidates already located in the United States. New petitions requiring processing from abroad are not supported at this time.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Multimodal Intelligence Systems Engineer
Multimodal Intelligence Systems Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 180,000
Multimodal Intelligence Systems Engineer
Multimodal Intelligence Systems Engineer

MaxIT Consulting - Max Corporate Group • California (MO)

On-site
USD 120,000 - 180,000
Industrial Edge Vision & Language Systems Engineer
Industrial Edge Vision & Language Systems Engineer

MaxIT Consulting - Max Corporate Group • California (MO)

On-site
USD 120,000 - 180,000
Founding Computer Vision Engineer – Multimodal AI & VLMs
Founding Computer Vision Engineer – Multimodal AI & VLMs

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Edge Vision AI Engineer: Multimodal VLMs
Edge Vision AI Engineer: Multimodal VLMs

MaxIT Consulting - Max Corporate Group • California

On-site
USD 140,000 - 190,000
Applied AI Engineer – Computer Vision & VLMs
Applied AI Engineer – Computer Vision & VLMs

MaxIT Consulting - Max Corporate Group • California (MO)

On-site
USD 120,000 - 200,000
Applied AI Engineer – Computer Vision & VLMs
Applied AI Engineer – Computer Vision & VLMs

MaxIT Consulting - Max Corporate Group • California

On-site
USD 140,000 - 190,000
Applied AI Engineer – Computer Vision & VLMs
Applied AI Engineer – Computer Vision & VLMs

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 180,000
Applied AI Engineer, Vision & VLMs for Edge Systems
Applied AI Engineer, Vision & VLMs for Edge Systems

MaxIT Consulting - Max Corporate Group • California (MO)

On-site
USD 120,000 - 200,000
Edge Vision & Multimodal AI Engineer
Edge Vision & Multimodal AI Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 180,000