Member of Technical Staff - Multimodal Understanding

SpaceXAI

Palo Alto (CA)

On-site

USD 180,000 - 440,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Medical insurance
Vision insurance
Dental coverage
401(k) retirement plan
Disability insurance
Life insurance

Job summary

SpaceXAI is seeking a Member of Technical Staff - Multimodal Understanding to push toward superhuman multimodal intelligence. You will work across vision, audio, video, and text, building data pipelines, training infrastructure, and end‑to‑end product experiences.

You will collaborate with pre‑training, post‑training, and product teams to deliver multimodal reasoning, world modeling, tool use, and interactive human‑AI collaboration at web/petabyte scale.

Qualifications

  • Hands-on multimodal pre-training/post-training/fine-tuning experience.
  • Expert Python with PyTorch/JAX/XLA.
  • Experience building and optimizing large-scale distributed ML systems.
  • Deep experience designing data pipelines at scale.
  • Strong evaluation frameworks, benchmarks, or RL familiarity.
  • Proactive self-starter in high‑intensity environments.
  • Willingness to own end‑to‑end initiatives.

Responsibilities

  • Design, build, and optimize distributed systems for multimodal pre-training, post-training, inference and data processing.
  • Develop high-throughput pipelines for data acquisition, preprocessing, filtering, generation, decoding, loading, and management.
  • Advance multimodal capabilities including cross-modal alignment, world modeling, reasoning, and real-time interaction.
  • Drive data quality and studies; curate, filter, and scale pipelines for trillion-parameter models.
  • Create evaluation frameworks, benchmarks, and metrics capturing real‑world usage and failures.
  • Innovate on algorithms, scaling, and co-design for state‑of‑the‑art performance.
  • Build tooling, demos, and end-to-end product experiences with rapid iteration.
  • Collaborate across pre-training, post-training, and product teams for frontier capabilities.

Skills

Multimodal ML
Python expert
Distributed ML systems
Data pipelines at scale
Evaluation/ RL basics
Self-starter

Tools

PyTorch
JAX
XLA
Rust
C++
Spark
Kubernetes
Ray

Job description

Member of Technical Staff - Multimodal Understanding

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands‑on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

You will join the multimodal team to push toward superhuman multimodal intelligence. Advance understanding and generation across modalities—image, video, audio, and text—spanning the full stack: data curation/acquisition, tokenizer training, large-scale pre‑training, post‑training/alignment, infrastructure/scaling, evaluation, tooling/demos, and end‑to‑end product experiences.

Collaborate cross‑functionally with pre‑training, post‑training, reasoning, data, applied, and product teams to deliver frontier capabilities in multimodal reasoning, world modeling, tool use, agentic behaviors, and interactive human‑AI collaboration. Contribute to building models that can see, hear, reason about, and interact with the world in real time at unprecedented levels.

RESPONSIBILITIES:
  • Design, build, and optimize large‑scale distributed systems for multimodal pre‑training, post‑training, inference, data processing, and tokenization at web/petabyte scale.
  • Develop high‑throughput pipelines for data acquisition, preprocessing, filtering, generation, decoding, loading, crawling, visualization, and management (images, videos, audio + text).
  • Advance multimodal capabilities including spatial‑temporal compression, cross‑modal alignment, world modeling, reasoning, emergent abilities, audio/image/video understanding & generation, real‑time video processing, and noisy data handling.
  • Drive data quality and studies: curation (human/synthetic), filtering techniques, analysis, and scalable pipelines to support trillion‑parameter models.
  • Create evaluation frameworks, internal benchmarks, reward models, and metrics that capture real‑world usage, failure modes, interactive dynamics, and human‑AI synergy.
  • Innovate on algorithms, modeling approaches, hardware/software/algorithm co‑design, and scaling paradigms for state‑of‑the‑art performance.
  • Build research tooling, user‑friendly interfaces, prototypes/demos, full‑stack applications, and enable rapid iteration based on feedback.
  • Work across the stack (pre‑training → SFT/RL/post‑training) to enable reasoning, tool calling, agentic behaviors, orchestration, and seamless real‑time interactions.
BASIC QUALIFICATIONS:
  • Hands‑on experience with multimodal pre‑training, post‑training, or fine‑tuning (vision, audio, video, or cross‑modal).
  • Expert‑level proficiency in Python (core language), with strong experience in at least one of: JAX / PyTorch / XLA.
  • Proven track record building or optimizing large‑scale distributed ML systems (training/inference optimization, GPU utilization, multi‑GPU/TPU setups, hardware co‑design).
  • Deep experience designing and running data pipelines at scale: curation, filtering, generation, quality studies, especially for noisy/real‑world multimodal data.
  • Strong fundamentals in evaluation design, benchmarks, reward modeling, or RL techniques (particularly for interactive/agentic behaviors).
  • Proactive self‑starter who thrives in high‑intensity environments and is passionate about pushing multimodal AI frontiers.
  • Willingness to own end‑to‑end initiatives and do whatever it takes to deliver breakthrough user experiences.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience leading major improvements in model capabilities through better data, modeling, algorithms, or scaling.
  • Familiarity with state‑of‑the‑art in multimodal LLMs, scaling laws, tokenizers, compression techniques, reasoning, or agentic systems.
  • Proficiency in Rust and/or C++ for performance‑critical components.
  • Hands‑on work with large‑scale orchestration tools such as Spark, Ray, or Kubernetes.
  • Background building full‑stack tooling: performant interfaces, real‑time research demos/apps, or end‑to‑end product ownership.
  • Passion for end‑to‑end user experience in interactive, real‑time multimodal AI systems.
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long‑term disability insurance, life insurance, and various other discounts and perks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Imagine Model
Imagine Model

Maxxd • Palo Alto (CA), Northern (KY)

On-site
USD 180,000 - 440,000
Equity
Medical, Vision & Dental
401(k) plan
Senior Multimodal Engineer — Real-Time AI & World Modeling
Senior Multimodal Engineer — Real-Time AI & World Modeling

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical insurance
Vision insurance
+4
Operations Engineer - Human Engineer
Operations Engineer - Human Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 144,000 - 270,000
Equity
Medical Insurance
Vision & Dental
+4
Human Data - Engineer
Human Data - Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 144,000 - 270,000
Human Data Manager New Palo Alto, California
Human Data Manager New Palo Alto, California

x • Palo Alto (CA), Northern (KY)

On-site
USD 100,000 - 186,000
Equity
Medical coverage
Vision coverage
+5
Member of Technical Staff - Imagine Model
Member of Technical Staff - Imagine Model

xAI • Seattle (WA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
+2
Software Engineer - Voice Model
Software Engineer - Voice Model

SpaceXAI • Palo Alto (CA)

On-site
USD 150,000 - 450,000
Equity
Medical coverage
Vision coverage
+5
Expert Team Lead, Engineering
Expert Team Lead, Engineering

SpaceXAI • Palo Alto (CA)

Remote
USD 104,000 - 170,000
Software Engineer, Media
Software Engineer, Media

SpaceXAI • Palo Alto (CA), Seattle (WA)

On-site
USD 180,000 - 440,000
Equity
Medical, Vision and Dental coverage
401(k) retirement plan
+1
Software Engineer, Media
Software Engineer, Media

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical, Vision, and Dental coverage
401(k) retirement plan
+3