ML Research Engineer, Inference Optimization

IntellifAI Labs Inc

India

Hybrid

INR 1,500,000 - 2,500,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IntellifAI Labs Inc is seeking a Research Engineer to accelerate generative models for image, video, and audio at consumer scale. You will read papers, implement promising ideas, and build benchmarks to prove speedups without quality loss.

You will deploy successful optimizations to our GPU fleet and own them long-term, collaborating across teams and shipping real improvements.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Electronics, Electrical Engineering, or a related discipline.
  • Hands-on experience training or running inference for image, video, or audio generation models in PyTorch.
  • Strong PyTorch fundamentals: profile a model, read a trace, distinguish kernel-bound vs memory-bound bottlenecks.

Responsibilities

  • Read literature on generative-model acceleration and identify the 10% methods that may apply.
  • Implement ideas from papers and run them on our models with access to Claude Code and Codex.
  • Measure end-to-end latency, throughput, and cost per generation; ensure quality is not sacrificed.
  • Ship optimizations to our GPU fleet and own them afterward.

Skills

PyTorch inference
Model profiling
GPU performance tuning
Benchmarking

Education

Bachelor’s or Master’s degree in CS/EE

Tools

CUDA
TensorRT
torch.compile

Job description

ABOUT US

Strontium is one of the fastest-growing AI companies in Canada—fully bootstrapped, with 15+ years in the industry and 100,000+ customers across 60+ countries. We build and sell AI-powered creative tools directly to consumers.

Two of our products:

  • VideoExpress.ai—our video-generation product, rated 4.8★ on Capterra with 6,000+ reviews and ranked #1 in its category:
  • https://www.capterra.com/p/10019949/VideoExpress/reviews/
  • Artistly.ai—our image-generation product, rated 4.8★ on Trustpilot with 1,000+ reviews:
  • https://ca.trustpilot.com/review/artistly.ai

Strontium and IntellifAI Labs are collaborating to bring state-of-the-art AI technologies to everyday users. We ship fast, iterate constantly, and compete with companies 100× our size.

THE ROLE — RESEARCH ENGINEER, INFERENCE OPTIMIZATION

We run image, video, and audio generation models in production at consumer scale. Every second we remove from a generation improves the product and lowers our GPU costs.

Your job is to make our models dramatically faster without sacrificing output quality. This is a research role with a production target.

WHAT YOU’LL DO
  1. 1. READ

    Stay current with the literature on generative-model acceleration. Most published techniques will not work for our models; your job is to identify the 10% that might.

  2. 2. IMPLEMENT

    Take promising ideas from papers and get them running on our models quickly. You will have unlimited access to Claude Code and Codex, and we expect you to use them aggressively.

  3. 3. MEASURE

    Build and defend benchmarks covering end-to-end latency, throughput under real traffic, cost per generation, and quality regressions. A speedup that quietly degrades output is not a speedup.

  4. 4. SHIP

    Deploy successful optimizations to our own GPU fleet and own them afterward.

OPTIMIZATION AREAS
  • Model level: step distillation, few-step samplers, caching and feature reuse, quantization, pruning, and architecture surgery. (40%)
  • Kernel level: CUDA/Triton kernels, fused attention, torch.compile, and TensorRT. (40%)
  • Serving level: batching, scheduling, and autoscaling. (20%)
ROLE DETAILS
  • Flexible hours with reasonable overlap with Eastern Time
MINIMUM QUALIFICATIONS
  • Bachelor’s or Master’s degree in Computer Science, Electronics, Electrical Engineering, or a related discipline, with a strong coding background.
  • Hands‑on experience training or running inference for image, video, or audio generation models in PyTorch.
  • Strong PyTorch fundamentals: you can profile a model, read a trace, and distinguish a kernel‑bound bottleneck from a memory‑bound one.
NICE TO HAVE (GENUINELY OPTIONAL)
  • Publications or open-source contributions in efficient inference.
  • CUDA or Triton kernel‑authoring experience.
  • Experience operating serving stacks such as vLLM, ComfyUI, or custom schedulers under real traffic.

If you have spent the last year reading papers because you could not put them down, and keep thinking about these ideas all the time this is your job—whether that work happened during a Master’s program, an internship, or on nights and weekends.

Questions? Contact info@intellifailabs.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Engineer
Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 800,000 - 1,200,000
Inference Performance Engineer
Inference Performance Engineer

adaption • India

On-site
INR 12,440,000 - 18,182,000
Flexible work: In-person collaboration
Adaption Passport: Annual travel
Lunch Stipend
+1
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

EnCharge AI • India

On-site
INR 3,000,000 - 6,000,000
Research Engineer, AI Models
Research Engineer, AI Models

AlleyCorp • India

On-site
INR 1,600,000 - 2,400,000
Distributed Training & Inference Optimization Engineer
Distributed Training & Inference Optimization Engineer

Winzons • India

On-site
INR 3,000,000 - 5,000,000
AI Inference Engineer – LLM
AI Inference Engineer – LLM

GyanSys Inc. • Bengaluru

On-site
INR 1,000,000 - 1,600,000
Senior Inference Engineer
Senior Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 2,600,000 - 4,800,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work model
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)

EnCharge AI • Delhi

On-site
INR 4,500,000 - 7,500,000
Senior Scientist AI Research Engineer
Senior Scientist AI Research Engineer

Mulya Consulting • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000