Multimodal AI Engineer (voice and vision)

Axiom Global Technologies

United States

Remote

USD 150,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Axiom Global Technologies seeks a Multimodal AI Engineer to own streaming speech and real-time computer vision for identity verification and proctoring. You will work on vision-language models in production and optimize models for constrained hardware while maintaining rigorous evaluation practices.

The role requires strong PyTorch, Python, Linux, SQL skills, and autonomous work with clean pull requests in a Kubernetes-deployed service. Excellent English communication is essential.

Qualifications

  • Must have streaming speech (recognition, synthesis, or both) shipped in production with latency numbers.
  • At least one vision model trained or fine-tuned on data they labeled or curated themselves, evaluated against incumbents, and shipped.
  • Strong PyTorch experience.
  • Audio fundamentals: voice activity detection, endpointing, resampling, feature extraction.
  • Vision fundamentals: face detection/embedding, gaze estimation, object detection (Ultralytics/YOLO), OpenCV.
  • Vision-language models in production or strong prompting/batching knowledge for structured output.
  • Model export and optimization for constrained hardware (ONNX Runtime plus TensorRT or INT8 or CPU inference).
  • Rigorous evaluation: held-out sets, confusion matrices, calibrated thresholds; honest reporting when models worsen.
  • Proficiency in Python, Linux, and SQL; comfortable with Kubernetes-deployed services.
  • Fluent written and spoken English; autonomous work style with clean PRs.

Responsibilities

  • Owns streaming speech in and out of the interview in a real-time proctoring setup.
  • Develops real-time computer vision for identity verification and proctoring.
  • Maintains vision-language models in production and optimizes for edge hardware.
  • Ensures reproducible experiments with proper evaluation protocols.

Skills

Streaming speech
Vision model training
PyTorch
Audio fundamentals
Vision fundamentals
Vision-language models
Model export/optimization
Evaluation discipline
Python/Linux/SQL
English communication

Tools

PyTorch
OpenCV
Ultralytics/YOLO
ONNX Runtime
TensorRT
Kubernetes
SQL
Linux

Job description

Role & responsibilities
Multimodal AI Engineer (voice and vision)

Owns streaming speech in and out of the interview and the real-time computer vision for identity verification and proctoring, including vision-language models.

Must have
  • Streaming speech (recognition, synthesis, or both) shipped in production under a latency budget, with real numbers for time to first audio and end-of-utterance latency.
  • At least one vision model trained or fine-tuned on data they labeled or curated themselves, evaluated against the incumbent, and shipped.
  • Strong PyTorch.
  • Audio fundamentals: voice activity detection, endpointing, resampling, and feature extraction.
  • Vision fundamentals: face detection and embedding (InsightFace or equivalent), gaze estimation, object detection (Ultralytics or YOLO), and OpenCV.
  • Vision-language models in production, or strong working knowledge of prompting and batching them for structured output.
  • Model export and optimization for constrained hardware: ONNX Runtime plus at least one of TensorRT, INT8 calibration, or CPU-targeted inference.
  • Evaluation discipline: held-out sets, confusion matrices, calibrated thresholds, and honest reporting when a new model is worse.
  • Python, Linux, and SQL at a professional level. Comfortable inside a Kubernetes-deployed service.
  • Fluent written and spoken English. Works autonomously with clean, reviewable pull requests.
Nice to have
  • Fine-tuning speech recognition on accented or domain audio, or training text-to-speech voices.
  • Speaker diarization, active speaker detection, or audio-visual sync.
  • Liveness, anti-spoofing, or deepfake and virtual-camera detection.
  • 3D reconstruction, depth estimation, or feature matching.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Real-Time Multimodal AI Engineer: Voice & Vision
Real-Time Multimodal AI Engineer: Voice & Vision

Axiom Global Technologies • United States

Remote
USD 150,000 - 190,000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 110,000 - 150,000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • Boston (MA)

On-site
USD 130,000 - 160,000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • Seattle (WA)

On-site
USD 120,000 - 150,000
AI Engineer, Voice and Realtime
AI Engineer, Voice and Realtime

Obble • United States

Remote
USD 120,000 - 180,000
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
Machine Learning Engineer, Multimodal Perception and Authentication
Machine Learning Engineer, Multimodal Perception and Authentication

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI/ML Engineer – LLM & Voice Pipeline
AI/ML Engineer – LLM & Voice Pipeline

Luxoft Germany • United States

Remote
USD 120,000 - 190,000
Multimodal ML Engineer - Vision, Audio & Text AI
Multimodal ML Engineer - Vision, Audio & Text AI

AI Breaking Wire • Mountain View (CA)

Hybrid
USD 200,000 - 320,000
Comprehensive health programs
Generous vacation & family leave
On-site gourmet cafeterias