Senior Inference Engineer

DeepRec.ai

Palo Alto (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity
Health benefits
Monthly stipends
Company retreats
In-office culture

Job summary

DeepRec.ai in Palo Alto, CA seeks a Senior Inference Engineer to accelerate AI-driven video generation products. You will design and optimize inference pipelines, implement acceleration techniques, and push the boundaries of real-time AI deployment.

You will collaborate with researchers and engineers to deploy state-of-the-art video generation and language models, scale performance, and mentor fellow engineers in GPU programming.

Qualifications

  • 5+ years of engineering experience in inference acceleration and model deployment at scale.
  • Expertise in inference optimization, including quantization and attention acceleration.
  • Strong CUDA/NCCL and GPU programming knowledge; experience with SP, TP, PP parallelism.
  • Familiarity with video generation models and large language models (LLMs).
  • Excellent cross-functional collaboration and ownership in fast-paced environments.

Responsibilities

  • Accelerate inference pipelines and optimize for real-time video generation.
  • Improve GPU utilization with tensor, sequence, and pipeline parallelism.
  • Develop high-performance kernels using CUDA and NCCL.
  • Collaborate with researchers and engineers to deploy state-of-the-art models.
  • Mentor engineers, participate in code reviews, and promote best practices in GPU programming.

Skills

Inference acceleration
CUDA
NCCL
GPU programming
Quantization
Attention optimization
Model deployment
Distributed inference
Video generation
LLMs
Collaboration
Ownership mindset

Tools

CUDA
NCCL

Job description

Senior Inference Engineer AI Video Generation Company (Stealth) | Palo Alto, CA | Hybrid

About the Role We are seeking a Senior Inference Engineer to accelerate the performance of our AI-driven video generation products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.

You will design and optimize inference pipelines, implement state-of-the‑art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what's possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of our video and language models.

What You’ll Do
  • Accelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
  • Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
  • Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
  • Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the‑art video generation and large language models into production.
  • Improve Training Efficiency: Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle. (Bonus)
  • Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
What We’re Looking For
  • Experience: 5+ years of engineering experience, with a strong track record in inference acceleration and model deployment at scale.
  • Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.
  • GPU and Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
  • AI Domain Knowledge: Familiarity with video generation models and large language models (LLMs).
  • Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.
  • Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.
Nice to Have
  • Experience with high-throughput video or real-time streaming model deployment.
  • Familiarity with distributed training and optimization toolkits.
  • Contributions to open source projects in AI infrastructure or deep learning compilers.
  • Startup or rapid prototyping experience.
What We Offer
  • Competitive salary commensurate with AI industry benchmarks.
  • Equity in a fast-growing company shaping the future of generative AI.
  • Comprehensive health benefits, monthly stipends, and company retreats.
  • A collaborative, in-office culture focused on building and shipping together.
About the Company

A well-funded, early-stage AI video generation startup headquartered in Palo Alto, CA. The team is building technology to make video creation seamless, intuitive, and universally accessible through the transformative power of AI. Tight-knit and highly energetic, the company values efficiency, intellectual curiosity, and the ambition to make a meaningful impact on the world.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Engineer — Real-Time Video AI at Scale
Senior Inference Engineer — Real-Time Video AI at Scale

DeepRec.ai • Palo Alto (CA)

Hybrid
USD 180,000 - 240,000
Competitive salary
Equity
Health benefits
+3
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Software Engineer, AI Compute Infrastructure
Software Engineer, AI Compute Infrastructure

HeyGen • Palo Alto (CA), San Francisco (CA), Los Angeles (CA)

On-site
USD 120,000 - 160,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative culture
+1
Software Engineer – AI Inference Engine
Software Engineer – AI Inference Engine

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Inference Engineer
Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Office Bellevue
Competitive compensation
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package
Inference Specialist, Creative Technology - InterPositive
Inference Specialist, Creative Technology - InterPositive

Netflix, Inc. • Los Angeles (CA)

On-site
USD 165,000 - 265,000
Health Plans
401(k) Retirement Plan with employer match
Stock Option Program
+2
System Software Engineer - AI
System Software Engineer - AI

Delos Data Inc • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits