Software Engineer - Voice Model

Pantera Capital

Palo Alto (CA)

On-site

USD 150,000 - 450,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Medical, vision, dental coverage
401(k) retirement plan
Disability insurance
Life insurance
Discounts and perks

Job summary

SpaceXAI is building the world’s best voice AI. Join the Grok Voice Model team to design data pipelines, pre-train/post-train speech-language models, and push accuracy, latency, and multilingual fluency to the limit.

You will work across data curation, synthetic data generation, annotation workflows, and evaluation infrastructure to sustain high-performance, real-time voice experiences. The role requires hands-on engineers who thrive in a curious, fast-paced environment, with strong

Qualifications

  • Proficiency in Python for AI/ML systems via clean, efficient code.
  • Experience processing large-scale datasets using Spark and Ray.
  • Expertise in pre-training and post-training of speech-language models (JAX/PyTorch).
  • Ability to set up rigorous evaluation pipelines: metrics, A/B testing, factual checks.

Responsibilities

  • Design and execute large-scale speech data curation and processing pipelines for high-quality model training and evaluation.
  • Work on pre-training and post-training with supervised fine-tuning and reinforcement learning to ensure accurate, natural multilingual voice responses.
  • Build an evaluation framework covering objective metrics, human studies, content factuality, and real-time interaction quality.

Skills

Python
Pre-training & fine-tuning
Evaluation pipelines
Large-scale distributed systems
Communication skills

Tools

Spark
Ray
JAX
PyTorch
Kubernetes

Job description

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

You will join the Grok Voice Model team to help build the world’s best voice AI. We deliver smooth, natural, low-latency spoken interactions — expressive, multilingual, and reliable across devices and real-time scenarios. We own the full training pipeline: massive data curation, premium audio processing, frontier speech-language pre-training, and intensive post-training to push quality, speed, and stability to the limit.

Our goal: make talking to AI feel like conversing with the most charming, kind, and knowledgeable person imaginable. We’re seeking exceptionally smart, execution-oriented engineers to help us get there.

RESPONSIBILITIES:
  • Design and execute large-scale speech data curation and processing pipelines, including collection of diverse real-world audio, synthetic data generation, and automated annotation workflows to enable high-quality model training and evaluation.
  • Work on pre-training and post-training of speech-language models, with targeted enhancements through supervised fine-tuning, reinforcement learning, and other techniques to ensure Grok Voice responses are accurate, factually grounded, natural and idiomatic in spoken style, conversational in tone, and fluent across multiple languages.
  • Build and iterate a comprehensive evaluation framework covering objective metrics (accuracy, quality, latency, expressiveness), human preference studies, content factuality assessments, real-time interaction quality, and experimentation infrastructure to measure and improve performance.
  • Work closely with product teams to integrate voice models into applications and real-time environments, define spoken interaction specifications, and handle the full lifecycle from prototype to global-scale deployment for stable, low-latency, delightful voice experiences.
BASIC QUALIFICATIONS:
  • Python expert with deep proficiency in writing clean, efficient code for AI/ML systems.
  • Hands-on experience processing large-scale datasets using tools like Spark and Ray for cleaning, augmentation, and feature extraction.
  • Proficiency in pre-training and post-training speech-language models using JAX/PyTorch, including supervised fine-tuning, reinforcement learning, and optimizations for accuracy, factuality, natural spoken style, detail, and multilingual fluency.
  • Ability to set up and run rigorous evaluation pipelines: objective metrics, human preference studies, content factuality checks, and iterative A/B testing to drive model improvements.
  • Experience building or working with large-scale distributed training and inference systems on Kubernetes.
  • Proactive, self-driven attitude — ready to grind in a fast-paced, high-caliber team to deliver outstanding voice AI experiences.
COMPENSATION AND BENEFITS:

$150,000 - $450,000 USD

  • equity
  • comprehensive medical, vision, and dental coverage
  • access to a 401(k) retirement plan
  • short & long-term disability insurance
  • life insurance
  • various other discounts and perks

SpaceXAI is an equal opportunity employer.

For details on data processing, view our Recruitment Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Voice Model
Member of Technical Staff - Voice Model

Xai • Palo Alto (CA)

On-site
USD 150,000 - 450,000
Equity
Comprehensive medical coverage
401(k) retirement plan
+2
Voice AI Engineer – Equity, Multilingual, Low-Latency
Voice AI Engineer – Equity, Multilingual, Low-Latency

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 450,000
Equity
Medical, vision, dental coverage
401(k) retirement plan
+3
Member of Technical Staff - Voice Product
Member of Technical Staff - Voice Product

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical coverage
401(k) retirement plan
Member of Technical Staff - Voice Product
Member of Technical Staff - Voice Product

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
+3
Software Engineer - Training/Inference (C++)
Software Engineer - Training/Inference (C++)

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
AI Tutor - Spanish at xAI
AI Tutor - Spanish at xAI

aitrainer • Northern (KY)

Hybrid
USD 48,000 - 62,000
Health insurance
401(k) plan
Paid sick leave
Human Data - Engineer
Human Data - Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 144,000 - 270,000
Equity
Comprehensive medical, vision, anddent
401(k) retirement plan
+3
AI Tutor - French
AI Tutor - French

SpaceXAI • United States

On-site
USD 48,000 - 62,000
Health insurance
401(k) plan
Paid sick leave
AI Tutor - Spanish xAI · Remote United States $35/hr →
AI Tutor - Spanish xAI · Remote United States $35/hr →

Dorado • Northern (KY)

Hybrid
USD 48,000 - 62,000
Health insurance
401(k) plan
Paid sick leave
Operations Engineer - Human Engineer
Operations Engineer - Human Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 144,000 - 270,000
Equity
Medical Insurance
Vision & Dental
+4