Real-Time Multimodal AI Engineer: Voice & Vision

Axiom Global Technologies

United States

À distance

USD 150 000 - 190 000

Plein temps

14 jours+
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Résumé du poste

Axiom Global Technologies seeks a Multimodal AI Engineer to own streaming speech and real-time computer vision for identity verification and proctoring. You will work on vision-language models in production and optimize models for constrained hardware while maintaining rigorous evaluation practices.

The role requires strong PyTorch, Python, Linux, SQL skills, and autonomous work with clean pull requests in a Kubernetes-deployed service. Excellent English communication is essential.

Qualifications

  • Must have streaming speech (recognition, synthesis, or both) shipped in production with latency numbers.
  • At least one vision model trained or fine-tuned on data they labeled or curated themselves, evaluated against incumbents, and shipped.
  • Strong PyTorch experience.
  • Audio fundamentals: voice activity detection, endpointing, resampling, feature extraction.
  • Vision fundamentals: face detection/embedding, gaze estimation, object detection (Ultralytics/YOLO), OpenCV.
  • Vision-language models in production or strong prompting/batching knowledge for structured output.
  • Model export and optimization for constrained hardware (ONNX Runtime plus TensorRT or INT8 or CPU inference).
  • Rigorous evaluation: held-out sets, confusion matrices, calibrated thresholds; honest reporting when models worsen.
  • Proficiency in Python, Linux, and SQL; comfortable with Kubernetes-deployed services.
  • Fluent written and spoken English; autonomous work style with clean PRs.

Responsabilités

  • Owns streaming speech in and out of the interview in a real-time proctoring setup.
  • Develops real-time computer vision for identity verification and proctoring.
  • Maintains vision-language models in production and optimizes for edge hardware.
  • Ensures reproducible experiments with proper evaluation protocols.

Connaissances

Streaming speech
Vision model training
PyTorch
Audio fundamentals
Vision fundamentals
Vision-language models
Model export/optimization
Evaluation discipline
Python/Linux/SQL
English communication

Outils

PyTorch
OpenCV
Ultralytics/YOLO
ONNX Runtime
TensorRT
Kubernetes
SQL
Linux

Description du poste

Axiom Global Technologies seeks a Multimodal AI Engineer to own streaming speech and real-time computer vision for identity verification and proctoring. You will work on vision-language models in production and optimize models for constrained hardware while maintaining rigorous evaluation practices.

The role requires strong PyTorch, Python, Linux, SQL skills, and autonomous work with clean pull requests in a Kubernetes-deployed service. Excellent English communication is essential.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Multimodal AI Engineer (voice and vision)
Multimodal AI Engineer (voice and vision)

Axiom Global Technologies • États-Unis

À distance
USD 150 000 - 190 000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • San Francisco (CA)

Sur place
USD 120 000 - 160 000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • Oregon (WI)

Sur place
USD 110 000 - 150 000
Multimodal AI Engineer: Vision-Language in Production
Multimodal AI Engineer: Vision-Language in Production

Tether.io • Indiana (PA)

Sur place
USD 90 000 - 130 000
Real-Time Vision & Voice AI Architect
Real-Time Vision & Voice AI Architect

Hyphen Connect Limited • San Francisco (CA)

Sur place
USD 120 000 - 160 000
Engineering Manager - Multimodal & Real-Time AI
Engineering Manager - Multimodal & Real-Time AI

Perplexity AI • Iowa (LA)

Sur place
USD 140 000 - 210 000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • Boston (MA)

Sur place
USD 130 000 - 160 000
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • Seattle (WA)

Sur place
USD 120 000 - 150 000
Senior Multimodal AI Engineer
Senior Multimodal AI Engineer

Cohere • New York (NY)

Sur place
USD 150 000 - 210 000
Health benefits
Remote-friendly culture
Learning stipend
Real-Time Multimodal AI/ML Engineer for Vision
Real-Time Multimodal AI/ML Engineer for Vision

Apple Inc. • Sunnyvale (CA)

Hybride
USD 150 000 - 278 000
Comprehensive medical and dental cover
Retirement benefits
Discounted products and free services
+3