Applied AI Modeling Engineer - Voice & Multimodal

Socket.dev

Los Altos (CA)

On-site

USD 140,000 - 200,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive Salary and Stock Options
Medical, dental, vision, retirement, &
Family Leave
Short Term & Long Term Disability
Paid time off and company holidays.
Learning and development support.

Job summary

Palona is seeking an applied AI Modeling Engineer to improve the intelligence, accuracy, safety, latency, and cost of its voice and multimodal agents in real restaurant environments. You will own problems across model selection and routing, prompting and context, fine-tuning or post-training when justified, speech and language quality, evaluation methodology, dataset development, and model behavior in production.

This is a product-facing modeling role.

Qualifications

  • 3+ years of industrial experience in relevant technical domain.
  • Strong ML foundations and hands-on experience developing or evaluating production AI systems.
  • Strong Python skills and experience with modern ML tooling such as PyTorch, JAX, Hugging Face, or equivalent systems.
  • Practical experience with LLMs, speech models, multimodal models, or agentic systems.
  • Ability to design reliable experiments, define useful metrics, analyze noisy results, and avoid optimizing against weak proxies.
  • Experience building datasets, evaluation harnesses, model services, or training and inference pipelines.
  • Strong software engineering judgment; your work is reproducible, tested, observable, and usable by other engineers.
  • Ability to connect modeling choices to product constraints including latency, cost, privacy, safety, and user experience.
  • Comfort operating in ambiguity and collaborating across research, engineering, product, and customer contexts.
  • AI-native working habits and genuine curiosity about new model capabilities and limitations.

Responsibilities

  • Develop modeling and experimentation strategies for high-impact agent problems in voice, language, reasoning, ordering, multilingual behavior, and multimodal understanding.
  • Build rigorous offline and online evaluations that measure task completion, accuracy, safety, latency, cost, conversational quality, and business outcomes.
  • Create and maintain representative datasets from simulations, human annotation, production feedback, and difficult edge cases while protecting sensitive data.
  • Evaluate frontier and open-source models and make clear build, buy, route, prompt, fine-tune, or distill decisions.
  • Improve prompting, context construction, memory, tool-use policies, structured outputs, model routing, and fallback behavior.
  • Design fine-tuning, preference optimization, distillation, or other post-training work when it offers a measurable advantage over simpler methods.
  • Partner with speech and real-time engineers to improve ASR, TTS, turn-taking, interruption handling, pronunciation, multilingual behavior, and end-to-end latency.
  • Develop analysis tools that explain model failures, slice performance by scenario, detect regressions, and accelerate iteration.
  • Ship model changes with production guardrails, staged rollouts, monitoring, rollback paths, and clear quality gates.
  • Translate new research and model releases into concrete product opportunities and communicate tradeoffs to technical and non-technical partners.
  • Raise scientific and engineering standards through reproducible experiments, thoughtful reviews, and clear documentation.

Skills

Machine learning foundations
Python
ML tooling
Experiment design
Production AI systems
LLMs
Speech models
Multimodal models
Model evaluation
Data annotation and datasets

Tools

PyTorch
JAX
Hugging Face

Job description

Palona is seeking an applied AI Modeling Engineer to improve the intelligence, accuracy, safety, latency, and cost of its voice and multimodal agents in real restaurant environments. You will own problems across model selection and routing, prompting and context, fine-tuning or post-training when justified, speech and language quality, evaluation methodology, dataset development, and model behavior in production.

This is a product-facing modeling role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Modeling Engineer: Voice & Multimodal (Stock Options)
AI Modeling Engineer: Voice & Multimodal (Stock Options)

Palona AI • Los Altos (CA)

On-site
USD 140,000 - 210,000
Stock options
Benefits: medical/dental/vision/ret/le
Family leave
+3
AI Modeling Engineer
AI Modeling Engineer

Socket.dev • Los Altos (CA)

On-site
USD 140,000 - 200,000
Competitive Salary and Stock Options
Medical, dental, vision, retirement, &
Family Leave
+3
AI Modeling Engineer
AI Modeling Engineer

Palona AI • Los Altos (CA)

On-site
USD 140,000 - 210,000
Stock options
Benefits: medical/dental/vision/ret/le
Family leave
+3
Senior PM, Model APIs & Developer Experience
Senior PM, Model APIs & Developer Experience

Doist • San Francisco (CA)

On-site
USD 200,000 - 280,000
Lead Real-Time Voice AI Architect
Lead Real-Time Voice AI Architect

Inflection AI, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 400,000 - 550,000
Robust medical, dental and vision with
401k matching
Flexible Time Off
+3
AI Product Engineer - End-to-End Restaurant Tech
AI Product Engineer - End-to-End Restaurant Tech

Palona AI • Los Altos (CA)

On-site
USD 140,000 - 190,000
Competitive salary and stock options
Health, dental, vision benefits
Family leave
+3
Staff AI Engineer — Speech & Voice Production
Staff AI Engineer — Speech & Voice Production

Salient • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical, dental, and vision coverage
Generous 401(k)
Catered lunches
Lead ML Engineer - Real-Time Voice & AI Agents
Lead ML Engineer - Real-Time Voice & AI Agents

AppFolio • Santa Barbara (CA)

Hybrid
USD 167,000 - 209,000
Hybrid work model
Founding Senior ML Engineer - Production Voice AI
Founding Senior ML Engineer - Production Voice AI

Visa Hunt • Redwood City (CA), Northern (KY)

Hybrid
USD 225,000 - 325,000
100% coverage for medical, dental, and
vision insurance
DoorDash credit
+4
Software Engineer, API Multimodal
Software Engineer, API Multimodal

OpenAI • San Francisco (CA)

On-site
USD 200,000 - 270,000