Artificial Intelligence Engineer

Flam

Bengaluru

On-site

INR 1,100,000 - 1,700,000

Full time

41 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Flam is building the next generation of interactive media with AI-native tech and Flicks for the US market. We are a small team owning end-to-end training, evaluation, serving and latency budgets.

The role focuses on fine-tuning, optimizing inference, building evaluation harnesses, and shipping research to production with robust serving infra.

Qualifications

  • 2+ years building ML systems in production.

Responsibilities

  • Fine-tune and evaluate open-weight models (LoRA/QLoRA, full SFT, preference tuning) and ship with no quality regression.

Skills

Python
PyTorch
LoRA/QLoRA
Model Fine-Tuning
GPU Optimization
Research to Prod.
Inference Stack

Tools

vLLM
SGLang
TensorRTLLM
TGI

Job description

Flam is building the next generation of interactive media through its content format. We are an AI-native technology company transforming how brands and consumers interact through immersive, interactive content. Our technology enables rich, app-less experiences that can be launched instantly on smartphones, creating a fundamentally different way for brands to engage consumers. We are backed by leading investors and already work with some of the world's largest brands. We are now building Flicks, our interactive media format for the US market.

About the role

Flam builds multimodal AI systems that ship to real customers. Our stack spans four production surfaces:

Falcon — Our LLM system, built on a sparse-MoE backbone.

Finesse — Our TTS and cross-lingual voice cloning system.

We are a small team. Whoever joins will own systems end-to-end training, evaluation, serving, and the latency budget.

Responsibilities:

Fine-tune and evaluate open-weight models LoRA/QLoRA, full SFT, preference tuning) and get them into production without a quality regression.

Optimize inference: quantization FP8/NVFP4/INT4, speculative decoding, prefix and KV-cache strategies, batching and scheduling on vLLM or SGLang.

Build evaluation harnesses that tell us something true — benchmark suites, regression gates, and per-release comparisons against both our own prior checkpoints and external baselines.

Own latency. Profile the pipeline, find where the milliseconds go, and remove them.

Take research to production: read the paper, replicate it, decide honestly whether it's worth shipping, and then ship it.

Write and maintain the serving infrastructure around your models —containers, autoscaling, GPU scheduling, observability.

What we're looking for
Required

2+ years building ML systems that ran in production, not only in notebooks.

Strong Python and PyTorch. You can read a model implementation and modify it, not just call .fit.

Hands-on experience with at least one modern inference stack (vLLM, SGLang, TensorRTLLM, or TGI) and a real understanding of what makes it fast.

Demonstrated fine-tuning experience — you've trained something, evaluated it properly, and know why your eval numbers meant what you claimed.

Comfort with GPU-level reasoning: memory layout, precision trade-offs, where the bottleneck actually is.

Ability to work from a paper. We move on recent research and expect you to be able to read it.

Strongly preferred

Experience in one or more of: speech ASR/TTS, diffusion and image generation, or video/avatar generation.

Quantization experience beyond running an off-the-shelf script.

Real-time streaming systems WebRTC, low-latency audio/video pipelines).

Indic language NLP, multilingual tokenizer work, or code-mixed data.

Cloud GPU deployment on GCP, AWS, or RunPod — including the unglamorous parts.

Our stack
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Curation Engineer
Data Curation Engineer

Keka Technologies Private Limited • Bengaluru

On-site
INR 2,000,000 - 4,000,000
UI Visual Designer
UI Visual Designer

Flam • Bengaluru

On-site
INR 60,000 - 100,000
Flexible remote work environment
Competitive salary
ESOPs
+1
Software Engineering Intern
Software Engineering Intern

Keka Technologies Private Limited • Bengaluru

On-site
INR 167,000 - 279,000
Head of QA / Quality Engineering
Head of QA / Quality Engineering

Flam • Bengaluru Urban

On-site
INR 4,000,000 - 8,000,000
Sr. AI Engineer
Sr. AI Engineer

fulcrumdigital • Pune District

On-site
INR 4,500,000 - 7,000,000
Senior Product Designer
Senior Product Designer

Flam • Bengaluru Urban

On-site
INR 1,000,000 - 2,000,000
Sr Product Designer
Sr Product Designer

Flam • Bengaluru

On-site
INR 900,000 - 1,500,000
Senior Product Designer
Senior Product Designer

Keka Technologies Private Limited • Bengaluru

On-site
INR 1,200,000 - 2,400,000
AI Engineer Job ID: 379291
AI Engineer Job ID: 379291

Kuku FM, Inc. • India

On-site
INR 1,200,000 - 1,800,000
Visual Designer
Visual Designer

Keka Technologies Private Limited • Bengaluru

On-site
INR 600,000 - 1,200,000