Multimodal AI Systems Engineer: Media & Speech

Meta

Columbia (SC)

Sur place

USD 154 000 - 217 000

Plein temps

Il y a 44 heures
Soyez parmi les premiers à postuler
Générateur de candidature

Une candidature conçue pour ce poste — un CV et une lettre de motivation personnalisés qui correspondent à l’offre.

Passez les filtres ATS

Résumé du poste

Meta's Applied AI (AAI) organization is hiring a Software Engineer for Multimedia & Multimodal AI. You will own data pipelines, evaluation, and model development across image, video, audio, and speech modalities.

This role emphasizes scaling data workflows, building robust ML systems, and mentoring engineers, with opportunities to advance state-of-the-art research while ensuring responsible AI practices.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or a related field.
  • 6+ years of programming experience in a relevant language or 3+ years with a PhD.
  • 3+ years building ML systems in production or research settings.
  • Strong Python and PyTorch.
  • Demonstrated experience with speech, audio, or music ML (ASR, TTS, codecs, MIR, etc.).
  • Experience with large-scale data pipelines and distributed training.
  • Track record of translating research ideas into working, measurable systems.

Responsabilités

  • Design and build agentic workflows and pipelines, including human-in-the-loop and expert-in-the-loop designs, to automate data production and scale output past what manual authoring supports.
  • Design and own data pipelines at scale: ingestion, filtering, pseudo-labeling and captioning with attribute classifiers, and provenance tracking for audio corpora.
  • Build evaluation infrastructure: objective metrics (speaker/style similarity, codec and generator quality), human listening-test pipelines, and the correlation analysis that ties the two together.
  • Improve training efficiency and reliability — distributed training, GPU utilization, codec and tokenizer retraining, experiment management.
  • Reproduce and extend state-of-the‑art research: implement new methods from papers into our codebases and run rigorous ablations.
  • Mentor engineers on the team, contribute to hiring and onboarding, and raise the bar on evaluation and quality practice.
  • Build and train generative and representation models for speech, sound, and music — including text-, audio-, and video-conditioned generation, infilling, editing, and style transfer.

Connaissances

Python
PyTorch
ML systems in production
Data pipelines
Research-to-production
Distributed training
Multimodal ML

Formation

Bachelor's degree in Computer Science / Computer Engineering or equivalent

Description du poste

Meta's Applied AI (AAI) organization is hiring a Software Engineer for Multimedia & Multimodal AI. You will own data pipelines, evaluation, and model development across image, video, audio, and speech modalities.

This role emphasizes scaling data workflows, building robust ML systems, and mentoring engineers, with opportunities to advance state-of-the-art research while ensuring responsible AI practices.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Multimodal AI Platform Engineer
Multimodal AI Platform Engineer

Meta • Tallahassee (FL)

Sur place
USD 154 000 - 217 000
Multimodal AI Engineer: Data Pipelines & Evaluation
Multimodal AI Engineer: Data Pipelines & Evaluation

Meta Careers • États-Unis

À distance
USD 154 000 - 217 000
Bonus
Equity
Benefits
Multimodal AI Engineer: Build Media Pipelines
Multimodal AI Engineer: Build Media Pipelines

Meta • Santa Fe (NM)

Sur place
USD 154 000 - 217 000
Multimodal AI Engineer: Pipelines, Evaluation & Synthesis
Multimodal AI Engineer: Pipelines, Evaluation & Synthesis

Meta • Indianapolis (IN)

Sur place
USD 154 000 - 217 000
Senior Multimodal AI Engineer - Multimedia Pipelines
Senior Multimodal AI Engineer - Multimedia Pipelines

Meta Careers • New York (NY)

Sur place
USD 184 000 - 257 000
Senior Multimodal AI Engineer
Senior Multimodal AI Engineer

Meta Careers • Menlo Park (CA)

Sur place
USD 184 000 - 257 000
Senior Multimodal AI Engineer — Multimedia Pipelines
Senior Multimodal AI Engineer — Multimedia Pipelines

Meta Careers • Bellevue (WA)

Sur place
USD 154 000 - 217 000
Equity
Benefits
Multimodal AI Engineer — Image/Video & Audio
Multimodal AI Engineer — Image/Video & Audio

xAI • Seattle (WA)

Sur place
USD 180 000 - 440 000
Equity
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
+2
Multimodal AI Engineer: Image/Video Generation & Systems
Multimodal AI Engineer: Image/Video Generation & Systems

xAI • Palo Alto (CA)

Sur place
USD 180 000 - 440 000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Short & long-term disability insurance
+3
Multimodal AI Systems Architect (AI Engineering)
Multimodal AI Systems Architect (AI Engineering)

Hyphen Connect Limited • San Francisco (CA)

Sur place
USD 120 000 - 160 000