Software Engineer, Multimedia & Multimodal AI

Meta Careers

Menlo Park (CA)

On-site

USD 184,000 - 257,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Meta's Applied AI group is seeking a senior engineer to lead end-to-end multimodal data pipelines in Multimedia & Multimodal AI. You will own data production, evaluation, and deployment, spanning image, video, audio, and speech modalities, from research questions to scalable pipelines.

You will set direction, define success metrics, mentor engineers, and collaborate with research, model-training, and engineering teams.

Qualifications

  • Bachelor's degree in CS/Engineering or equivalent practical experience.
  • 8+ years programming or 4+ years with PhD.
  • 3+ years building ML systems in production or research.

Responsibilities

  • Set technical direction for a modality or capability area.
  • Design and build data pipelines and evaluation infrastructure.
  • Mentor engineers and contribute to hiring and onboarding.
  • Lead end-to-end data production and measurement pipelines.
  • Collaborate across research, model-training, and engineering partners.

Skills

Python
PyTorch
Distributed training
ML systems
Speech ML
Data pipelines
Research-to-production

Education

Bachelor's degree in CS/Engineering or equivalent

Job description

Applied AI (AAI) is Meta’s organization focused on making our AI models best-in-class, starting with coding. Within AAI, the Multimedia & MultiModality team covers the multimedia domain across every modality, on both the input and the output side of a model: image, video, audio, speech and music. We work directly with research, model-training and engineering partners across MSL, TBD and FAIR. Current problems include evaluating video experiences, diagnosing multimedia model behavior, producing domain-expert agent tasks, and building the data and measurement pipelines multimodal capabilities are trained and judged against.About the roleWe are hiring a senior engineer to lead this work end to end. You will take a modality or a capability area, decide what data is worth producing and how it should be measured, and carry it from an open question through to a pipeline that runs and a measurement the org relies on.This is a multimodal role, not a text-only role. You will work across image, video, audio and speech, as model inputs and as model outputs, and the data and evaluations you own will cover media, not text alone.You will choose where the pod invests, own outcomes beyond your individual contribution, set standards other engineers build against, and raise quality without becoming the review bottleneck.Software Engineer, Multimedia & Multimodal AI Responsibilities:Set technical direction for a modality or capability area.Determining where the team invests, what success means, and the tradeoffs behind both.Drive that direction across partner teams, not just inside your own.Design and build agentic workflows and pipelines, including human-in-the-loop and expert-in-the-loop designs, to automate data production and scale output past what manual authoring supports.Design and own data pipelines at scale: ingestion, filtering, pseudo-labeling and captioning with attribute classifiers, and provenance tracking for audio corpora.Build evaluation infrastructure: objective metrics (speaker/style similarity, codec and generator quality), human listening-test pipelines, and the correlation analysis that ties the two together.Improve training efficiency and reliability — distributed training, GPU utilization, codec and tokenizer retraining, experiment management.Reproduce and extend state-of-the-art research: implement new methods from papers into our codebases and run rigorous ablations.Mentor engineers on the team, contribute to hiring and onboarding, and raise the bar on evaluation and quality practice.Build and train generative and representation models for speech, sound, and music — including text-, audio-, and video-conditioned generation, infilling, editing, and style transfer.Minimum Qualifications:Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience8+ years of programming experience in a relevant language OR 4+ years experience with a PhD3+ years building ML systems in production or research settingsstrong Python and PyTorchDemonstrated experience with speech, audio, or music ML - ASR, TTS, audio codecs, music information retrieval, self-supervised audio representation learning, or audio generative modelingExperience with large-scale data pipelines and distributed trainingTrack record of translating research ideas into working, measurable systemsPreferred Qualifications:Publications at top venues (ICASSP, Interspeech, ISMIR, NeurIPS, ICML, ICLR) in speech, audio, or musicGenerative modeling of continuous data (diffusion / flow matching, audio or vision), and demonstrated ability to switch domains and ramp quicklyAudio DSP depth — pitch detection, FFT, real-time signal processingExperience with disentangled or controllable generation (voice, emotion, style, instrumentation)Experience building evaluation harnesses and human-eval pipelines for generative audioMusic domain expertise: stem separation, mixing, lyrics/vocal conditioningExperience designing benchmarks or evaluations for model capability, with attention to grading reliability, reproducibility and label qualityExperience building data pipelines for image, video, audio, speech or complex media formats, including versioning, lineage and provenanceExperience designing AI agents, orchestration, or human-in-the-loop systemsHands-on experience evaluating or red-teaming multimodal models, or creating the data used to improve themUnderstanding of Responsible AI practices and building quality controls into AI outputExperience with zero-to-one work: forming a charter and standing up process while priorities are still movingDemonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologiesAbout Meta:Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.$183,997/year to $257,000/year + bonus + equity + benefitsIndividual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Multimedia & Multimodal AI
Software Engineer, Multimedia & Multimodal AI

Meta Careers • Bellevue (WA)

On-site
USD 154,000 - 217,000
Equity
Benefits
Software Engineer, Multimedia & Multimodal AI
Software Engineer, Multimedia & Multimodal AI

Meta Careers • New York (NY)

On-site
USD 184,000 - 257,000
Software Engineer, Multimedia & Multimodal AI
Software Engineer, Multimedia & Multimodal AI

Meta • Bellevue (WA)

On-site
USD 184,000 - 257,000
Health insurance
Equity compensation
Software Engineer, Audio SWE
Software Engineer, Audio SWE

Meta • Burlingame (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Research Scientist, Contextual AI and Multimodal Agents
Research Scientist, Contextual AI and Multimodal Agents

Meta • Redmond (WA)

On-site
USD 219,000 - 301,000
Creative Coder
Creative Coder

Meta • Los Angeles (CA)

On-site
USD 154,000 - 216,000
Bonus
Equity
Health benefits
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Burlingame (CA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Research Engineer, Language - Wearables Polyglot AI
Research Engineer, Language - Wearables Polyglot AI

Meta • Redmond (WA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • San Francisco (CA)

On-site
USD 183,997 - 257,000
Software Engineer, Systems ML (Technical Leadership)
Software Engineer, Systems ML (Technical Leadership)

Meta • New York (NY)

On-site
USD 219,000 - 301,000
Bonus
Equity
Benefits