Senior Backend Engineer, Vision

Sarvam

Bengaluru

On-site

INR 3,000,000 - 6,000,000

Full time

8 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India.

You will own the architecture of the OCR and extraction serving harness for Sarvam's vision models, delivering frontier-grade extraction quality from 3B and 30B in-house models at national scale, while balancing latency and cost budgets that

Qualifications

  • 5–6+ years in backend engineering with production systems.
  • Strong proficiency in Go and/or Python and where each fits best.
  • Deep understanding of distributed systems design and trade-offs.
  • Production experience with Temporal or similar durable execution engines.
  • Kubernetes in production including autoscaling and GPU scheduling.
  • Experience serving ML/LLM inference in production with batching and caching.
  • Rigour in observability, reliability, and incident response.
  • Cost awareness to reduce per-page expenses without compromising quality.

Responsibilities

  • Own end-to-end architecture of the OCR and extraction serving harness: API layer, orchestration, inference layer, post-processing, delivery
  • Design the accuracy harness — multi-pass extraction, ensembling, cross verification, schema-constrained decoding, confidence calibration
  • Architect durable, resumable document workflows in Temporal: fan-out across pages, partial failure recovery, exactly-once side effects
  • Own the inference serving layer alongside infra: batching strategy, GPU pool management, autoscaling on real signals
  • Drive latency, throughput and unit economics down deliberately — profile, measure, and defend cost-per-page targets as volume scales
  • Build the observability substrate: distributed tracing, per-stage cost and quality metrics, SLOs, alerting
  • Design for multi-tenancy, tenant isolation, rate limiting and fair scheduling across enterprise customers
  • Support on-prem and constrained deployments where the harness runs inside customer environments
  • Set technical direction and mentor SDE 1–2 engineers

Skills

Go
Python
Distributed systems
Temporal
Kubernetes
ML inference
Observability
Cost optimization

Job description

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About The Team

Sarvam's research teams build our own vision-language models for OCR and structured extraction. This team builds everything around them — the serving harness that turns a 3B or 30B in-house model into a production document intelligence platform.

The bet is specific: with the right harness — routing, decomposition, retries, verification, ensembling, layout awareness, confidence calibration — a small sovereign model should match or beat what teams today get from frontier hosted models like Gemini Flash, at a fraction of the cost and fully within India. Closing that gap is an engineering problem, and it is this team's problem.

We run against the full messiness of Indian documents at population scale: PAN and Aadhaar, bank statements, GST filings, insurance and medical reports, 60-page

contracts, legal filings and RFPs — across languages, scan quality, and layouts that were never designed to be machine-read.

Stack: Go, Python, Temporal, REST, Kubernetes, PostgreSQL, Redis, object storage, OpenTelemetry-based observability.

About The Role

You will own the architecture of the serving harness for Sarvam's vision models — the system that has to deliver frontier-grade extraction quality out of 3B and 30B in-house models, at national scale, with cost and latency budgets that actually close.

This means owning the hard trade-off surface directly: accuracy versus latency versus rupees per page. Multi-pass inference, model routing and cascades, self-consistency and verification passes, confidence-driven escalation, batching and caching strategy, GPU utilisation. These are the levers that decide whether the product works, and you will be the person deciding how to pull them.

You will also set the reliability bar. These pipelines process documents that customers cannot afford to lose — KYC, loan underwriting, claims, contracts. Durability, idempotency, backpressure and graceful degradation are the baseline, not the roadmap.

The architecture you set will be inherited by everything the team builds after you.

What You'll Do

Own the end-to-end architecture of the OCR and extraction serving harness: API layer, orchestration, inference layer, post-processing, delivery

Design the accuracy harness — multi-pass extraction, ensembling, cross verification, schema-constrained decoding, confidence calibration, targeted re runs — and prove its gains against held-out evaluation sets

Architect durable, resumable document workflows in Temporal: fan-out across pages, partial failure recovery, exactly-once side effects, long-running jobs measured in minutes to hours

Own the inference serving layer alongside infra: batching strategy, GPU pool management, autoscaling on real signals, queue depth and admission control, multi-model routing

Drive latency, throughput and unit economics down deliberately — profile, measure, and defend cost-per-page targets as volume scales

Build the observability substrate: distributed tracing across the pipeline, per-stage cost and quality metrics, SLOs, alerting, and post-incident rigour

Design for multi-tenancy, tenant isolation, rate limiting and fair scheduling across enterprise customers with very different load shapes

Support on-prem and constrained deployments where the whole harness has to run inside a customer's environment

Set technical direction and raise the bar through design review and mentorship of SDE 1–2 engineers

What We're Looking For
  • 5–6+ years in backend engineering, with meaningful time spent operating high throughput production systems you were on-call for
  • Deep proficiency in Go and/or Python, and the judgement to know which belongs where
  • Strong distributed systems design: queues, workflow orchestration, idempotency, backpressure, retry and timeout semantics, consistency trade-offs, graceful degradation
  • Production experience with Temporal or an equivalent durable execution engine, on workflows that mattered
  • Kubernetes in production — autoscaling, resource management, rollouts, debugging under load; GPU workload scheduling is a strong plus
  • Demonstrable experience serving ML or LLM inference in production: batching, caching, model versioning, A/B rollout, latency budgeting
  • Rigour about observability and reliability — you have designed SLOs, run incidents, and shipped the fixes
  • Hard-won cost intuition: you have made a system materially cheaper without giving up quality
Bonus Points

Direct experience with OCR, IDP, or document AI systems — Textract, Document AI, Azure DI, or something you built yourself

GPU inference stacks: vLLM, TensorRT-LLM, Triton, SGLang, Ray Serve Evaluation infrastructure for ML systems — golden sets, regression gates, human in-the-loop review loops

Experience with BFSI, healthcare, or public-sector compliance and data-residency constraints

On-prem or air-gapped deployment experience

Note

We are looking for people who can own the outcomes described here, not people who match every line of this specification. If this problem excites you and you believe you can do this work, we want to hear from you.

Why Sarvam?

Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers ofAI with real population-scale impact.

Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar

High ownership and high impact, from day one

Everything we do is AI-first, from the way we build and ship to the way we think about problems

You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier ofAI in India, Sarvam is the place to be.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontend Engineer, Vision
Frontend Engineer, Vision

Sarvam • Bengaluru

On-site
INR 1,200,000 - 2,200,000
Staff Engineer, API Platform
Staff Engineer, API Platform

Sarvam • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Forward Deployed Software Engineer, Model API
Forward Deployed Software Engineer, Model API

Neara • Bengaluru

On-site
INR 1,000,000 - 1,500,000
High ownership and impact
Collaborative team environment
Opportunity to influence India's AI landscape
Forward Deployed Software Engineer
Forward Deployed Software Engineer

Neara • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Senior Backend Engineer,Vision
Senior Backend Engineer,Vision

Neara • India

On-site
INR 800,000 - 1,500,000
High-impact projects
Innovation-driven team
AI-first approach
Embedded Data Scientist, Chanakya
Embedded Data Scientist, Chanakya

Sarvam • Delhi

On-site
INR 1,000,000 - 2,000,000
Backend Engineer, Chanakya
Backend Engineer, Chanakya

Sarvam • Bengaluru

On-site
INR 1,500,000 - 2,500,000
High ownership and impact
AI-first work environment
Forward Deployed Engineer
Forward Deployed Engineer

Sarvam AI • Bengaluru

On-site
INR 1,800,000 - 3,000,000
ML Engineer (Data), Foundational Models
ML Engineer (Data), Foundational Models

Sarvam • Bengaluru

On-site
INR 1,500,000 - 2,000,000
High ownership and impact
Work alongside top talent
AI-first environment
Backend Engineer - Studio Media Platform
Backend Engineer - Studio Media Platform

Sarvam • Bengaluru

On-site
INR 1,200,000 - 1,800,000
High ownership and impact
Fast-paced, high talent-density team
Opportunity to work on AI-frontiers