Software Engineer, Multimedia & Multimodal AI

Meta

Columbia (SC)

On-site

USD 154,000 - 217,000

Full time

30 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Meta's Applied AI (AAI) organization is hiring a Software Engineer for Multimedia & Multimodal AI. You will own data pipelines, evaluation, and model development across image, video, audio, and speech modalities.

This role emphasizes scaling data workflows, building robust ML systems, and mentoring engineers, with opportunities to advance state-of-the-art research while ensuring responsible AI practices.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or a related field.
  • 6+ years of programming experience in a relevant language or 3+ years with a PhD.
  • 3+ years building ML systems in production or research settings.
  • Strong Python and PyTorch.
  • Demonstrated experience with speech, audio, or music ML (ASR, TTS, codecs, MIR, etc.).
  • Experience with large-scale data pipelines and distributed training.
  • Track record of translating research ideas into working, measurable systems.

Responsibilities

  • Design and build agentic workflows and pipelines, including human-in-the-loop and expert-in-the-loop designs, to automate data production and scale output past what manual authoring supports.
  • Design and own data pipelines at scale: ingestion, filtering, pseudo-labeling and captioning with attribute classifiers, and provenance tracking for audio corpora.
  • Build evaluation infrastructure: objective metrics (speaker/style similarity, codec and generator quality), human listening-test pipelines, and the correlation analysis that ties the two together.
  • Improve training efficiency and reliability — distributed training, GPU utilization, codec and tokenizer retraining, experiment management.
  • Reproduce and extend state-of-the‑art research: implement new methods from papers into our codebases and run rigorous ablations.
  • Mentor engineers on the team, contribute to hiring and onboarding, and raise the bar on evaluation and quality practice.
  • Build and train generative and representation models for speech, sound, and music — including text-, audio-, and video-conditioned generation, infilling, editing, and style transfer.

Skills

Python
PyTorch
ML systems in production
Data pipelines
Research-to-production
Distributed training
Multimodal ML

Education

Bachelor's degree in Computer Science / Computer Engineering or equivalent

Job description

Summary:

Applied AI (AAI) is Meta’s organization focused on making our AI models best-in-class, starting with coding. Within AAI, the Multimedia & MultiModality team covers the multimedia domain across every modality, on both the input and the output side of a model: image, video, audio, speech and music. We work directly with research, model-training and engineering partners across MSL, TBD and FAIR. Current problems include evaluating video experiences, diagnosing multimedia model behavior, producing domain-expert agent tasks, and building the data and measurement pipelines multimodal capabilities are trained and judged against.About the roleYou will take a modality or a capability area, decide what data is worth producing and how it should be measured, and carry it from an open question through to a pipeline that runs and a measurement the org relies on.This is a multimodal role, not a text-only role. You will work across image, video, audio and speech, as model inputs and as model outputs, and the data and evaluations you own will cover media, not text alone.You will choose where the pod invests, own outcomes beyond your individual contribution, set standards other engineers build against, and raise quality without becoming the review bottleneck.

Required Skills:

Software Engineer, Multimedia & Multimodal AI Responsibilities:

  1. Design and build agentic workflows and pipelines, including human-in-the-loop and expert-in-the-loop designs, to automate data production and scale output past what manual authoring supports.

  2. Design and own data pipelines at scale: ingestion, filtering, pseudo-labeling and captioning with attribute classifiers, and provenance tracking for audio corpora.

  3. Build evaluation infrastructure: objective metrics (speaker/style similarity, codec and generator quality), human listening-test pipelines, and the correlation analysis that ties the two together.

  4. Improve training efficiency and reliability — distributed training, GPU utilization, codec and tokenizer retraining, experiment management.

  5. Reproduce and extend state-of-the‑art research: implement new methods from papers into our codebases and run rigorous ablations.

  6. Mentor engineers on the team, contribute to hiring and onboarding, and raise the bar on evaluation and quality practice.

  7. Build and train generative and representation models for speech, sound, and music — including text-, audio-, and video-conditioned generation, infilling, editing, and style transfer.

Minimum Qualifications:
  1. Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience

  2. 6+ years of programming experience in a relevant language or 3+ years of experience + PhD

  3. 3+ years building ML systems in production or research settings

  4. strong Python and PyTorch

  5. Demonstrated experience with speech, audio, or music ML - ASR, TTS, audio codecs, music information retrieval, self-supervised audio representation learning, or audio generative modeling

  6. Experience with large-scale data pipelines and distributed training

  7. Track record of translating research ideas into working, measurable systems

Preferred Qualifications:
  1. Publications at top venues (ICASSP, Interspeech, ISMIR, NeurIPS, ICML, ICLR) in speech, audio, or music

  2. Generative modeling of continuous data (diffusion / flow matching, audio or vision), and demonstrated ability to switch domains and ramp quickly

  3. Audio DSP depth — pitch detection, FFT, real‑time signal processing

  4. Experience with disentangled or controllable generation (voice, emotion, style, instrumentation)

  5. Experience building evaluation harnesses and human‑eval pipelines for generative audio

  6. Music domain expertise: stem separation, mixing, lyrics/vocal conditioning

  7. Experience designing benchmarks or evaluations for model capability, with attention to grading reliability, reproducibility and label quality

  8. Experience building data pipelines for image, video, audio, speech or complex media formats, including versioning, lineage and provenance

  9. Experience designing AI agents, orchestration, or human‑in‑the‑loop systems

  10. Hands‑on experience evaluating or red‑team­ing multimodal models, or creating the data used to improve them

  11. Understanding of Responsible AI practices and building quality controls into AI output

  12. Experience with zero‑to‑one work: forming a charter and standing up process while priorities are still moving

  13. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)

  14. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)

  15. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies

Public Compensation:

$154,003/year to $217,000/year + bonus + equity + benefits

Industry:

Internet

Equal Opportunity:

Meta is proud to be an Equal Employment Opportunity and Aff

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Multimedia & Multimodal AI
Software Engineer, Multimedia & Multimodal AI

Meta • Tallahassee (FL)

On-site
USD 154,000 - 217,000
Software Engineer, Multimedia & Multimodal AI
Software Engineer, Multimedia & Multimodal AI

Meta Careers • Menlo Park (CA)

On-site
USD 184,000 - 257,000
Software Engineer, Multimedia & Multimodal AI
Software Engineer, Multimedia & Multimodal AI

Meta Careers • Bellevue (WA)

On-site
USD 154,000 - 217,000
Equity
Benefits
Software Engineer, Multimedia & Multimodal AI
Software Engineer, Multimedia & Multimodal AI

Meta Careers • New York (NY)

On-site
USD 184,000 - 257,000
Creative Coder
Creative Coder

Meta • Los Angeles (CA)

On-site
USD 154,000 - 216,000
Bonus
Equity
Health benefits
Software Engineer, Systems ML
Software Engineer, Systems ML

Meta • New York (NY)

On-site
USD 154,003 - 217,000
Bonus
Equity
Benefits
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Tallahassee (FL)

On-site
USD 184,000 - 257,000
Software Engineer, AI Native
Software Engineer, AI Native

SupportFinity™ • Seattle (WA)

On-site
USD 180,000 - 260,000
Bonus
Equity
Benefits
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Member of Technical Staff - Multimodal Understanding
Member of Technical Staff - Multimodal Understanding

Maven Ventures • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical coverage
401(k) retirement plan
+2