Machine Learning Engineer, Vision

Neara

Bengaluru

Hybrid

INR 1,500,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI solutions company in Bengaluru is looking for a Machine Learning Engineer specialized in vision-language models. The role involves designing training pipelines, working with clients to develop tailored strategies, and optimizing models for real-world applications. Candidates should have strong Python and PyTorch skills, hands-on experience with large models, and a technical degree. Join a fast-paced team dedicated to solving impactful AI challenges in India.

Qualifications

  • Hands-on experience training or fine-tuning large models.
  • Solid grounding in transformer architectures and modern training techniques.
  • Comfort with ambiguity in project specifications.

Responsibilities

  • Design and run training and fine-tuning pipelines for large vision-language models.
  • Build multimodal data pipelines for data ingestion and quality assurance.
  • Optimize models for inference including quantisation and serving infrastructure.
  • Work directly with clients to translate their use cases into ML tasks.

Skills

Python
PyTorch
Data pipeline building
Secure coding practices
Transformer architectures

Education

Undergraduate degree in a technical discipline (CS, statistics, physics, or equivalent)

Job description

Machine Learning Engineer, Vision

Job type: Full Time · Department: Engineering · Work type: On-Site

Bengaluru, Karnataka, India

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

You will work across the full lifecycle of vision-language model (VLM) development — data, training, evaluation, and production. The team’s scope will evolve as the field does; we want engineers who are comfortable with that.

What You’ll Do
  • Design and run training and fine-tuning pipelines for large vision-language models on GPU clusters
  • Build multimodal data pipelines — ingestion, filtering, deduplication, synthetic generation, and quality assurance
  • Implement and experiment with new architectures and training techniques from research
  • Build evaluation harnesses, benchmarks, and automated regression tracking
  • Optimise models for inference — quantisation, batching, and serving infrastructure
  • Build robust pipelines and integrations that put vision model capabilities in the hands of end users
  • Translate real-world problems into well-scoped ML tasks with the right data and evaluation strategy
  • Work directly with clients to understand their use cases — document processing, visual search, form extraction — and own the solution end to end
  • Build production-grade systems on top of Sarvam Vision and open-source models: multimodal pipelines, retrieval-augmented workflows, and structured output extraction
  • Debug and improve deployed solutions — latency, accuracy, edge cases, and integration with client infrastructure
What We’re Looking For
  • Strong Python and PyTorch — comfortable reading and modifying model internals
  • Hands‑on experience training or fine‑tuning large models, including debugging broken runs
  • Experience building data pipelines at scale
  • Solid grounding in transformer architectures and modern training techniques
  • Comfort with ambiguity — the roadmap is not fully pre-specified
  • Strong focus on secure coding practices, code quality, and system reliability
  • Undergraduate degree in a technical discipline (CS, statistics, physics, or equivalent)
Bonus Points
  • Experience with vision-language models or multimodal systems
  • Distributed training (FSDP, DeepSpeed, Megatron‑LM)
  • Post‑training methods — RLHF, DPO, or alignment techniques
  • Inference optimisation — quantisation, distillation, serving
  • Prior exposure to vision‑based AI systems or document processing pipelines
  • Contributions to open-source projects or a solid GitHub portfolio
Why Sarvam?

Sarvam is a fast‑moving, high talent‑density team building full‑stack AI for India, working on problems that push the frontiers of AI with real population‑scale impact.

  • Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar
  • High ownership and high impact, from day one
  • Everything we do is AI‑first, from the way we build and ship to the way we think about problems
  • You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Researcher, Vision
Researcher, Vision

Neara • India

On-site
INR 4,000,000 - 7,000,000
Researcher, Vision
Researcher, Vision

Sarvam • Bengaluru

On-site
INR 1,200,000 - 2,000,000
High ownership and high impact from day one
Collaborative work with researchers and engineers
Backend Engineer, Vision
Backend Engineer, Vision

Neara • India

On-site
INR 1,000,000 - 1,500,000
Senior Backend Engineer,Vision
Senior Backend Engineer,Vision

Neara • India

On-site
INR 800,000 - 1,500,000
High-impact projects
Innovation-driven team
AI-first approach
ML Researcher, Foundational Models
ML Researcher, Foundational Models

Sarvam • Bengaluru

On-site
INR 1,000,000 - 2,000,000
ML Engineer (Training Infra), Foundational Models
ML Engineer (Training Infra), Foundational Models

Sarvam • Bengaluru

On-site
INR 1,200,000 - 2,000,000
High ownership and impact
AI-first approach
Forward Deployed Software Engineer
Forward Deployed Software Engineer

Neara • Bengaluru

On-site
INR 1,000,000 - 2,000,000
ML Engineer (Data), Foundational Models
ML Engineer (Data), Foundational Models

Sarvam • Bengaluru

On-site
INR 1,500,000 - 2,000,000
High ownership and impact
Work alongside top talent
AI-first environment
Embedded Data Scientist, Chanakya
Embedded Data Scientist, Chanakya

Neara • Delhi

On-site
INR 1,000,000 - 1,500,000
ML Ops Engineer, Chanakya
ML Ops Engineer, Chanakya

Neara • Delhi

On-site
INR 1,500,000 - 2,000,000
High ownership in projects
Collaborative team environment
Opportunity to work on impactful AI solutions