Platform Engineer

Recrew AI

Bengaluru

On-site

INR 1,800,000 - 3,000,000

Full time

46 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Founding-team ownership over platform
Direct collaboration with founders
Deep-work engineering culture
Shape engineering culture and tooling

Job summary

Recrew AI, a research-first AI company incubated at IISc, is hiring a Senior Platform Engineer in Bengaluru to own the core backend, data systems, and MLOps infra that powers autonomous agents across critical networks.

You will drive end-to-end agent serving, data integration, learning feedback loops, and governance, reporting to the founders with strong ownership over platform direction at seed-stage.

Qualifications

  • 4+ years of backend engineering experience building and operating production-grade, scaled software systems end-to-end.
  • Hands-on experience operating production ML/AI platforms - model versioning, reproducible jobs, and safe rollout/rollback in live environments.
  • Strong data engineering skills: streaming pipelines, schedulers, SQL, object storage, and data provenance/quality practices.
  • Proficiency in async Python and FastAPI for building high-performance serving and API layers.
  • Experience with containerized, distributed deployments using Docker and Kubernetes on major cloud providers.
  • Demonstrated ability to build audit-grade logging and observability systems for systems that take real-world actions.
  • Distributed-systems fundamentals: retries, timeouts, circuit breakers, and graceful degradation under failure.

Responsibilities

  • Own the agent serving and runtime path end-to-end, including model versioning, deployment pipelines, and rollback mechanisms using LangChain/LangGraph.
  • Build and maintain the tool and data integration layer connecting agents to external data sources, APIs, and domain-specific infrastructure systems.
  • Design and operate the learning and feedback pipeline - capturing agent outcomes, labeling signals, and closing the loop for continuous improvement.
  • Enforce platform governance: budget controls, rate limits, and approval gates for agent actions on critical infrastructure.
  • Build audit-grade tracing, observability, and logging systems (using Langfuse or equivalent) that meet real-world accountability requirements for autonomous action systems.
  • Own platform reliability, latency, and uptime against defined SLOs - including designing for retries, timeouts, and graceful degradation in distributed environments.
  • Collaborate closely with agent, data, and domain teams to ensure platform abstractions support evolving research and product requirements.

Skills

Backend engineering
ML/AI platforms
Data engineering
Python (async)
FastAPI
Docker
Kubernetes
Observability
Distributed systems
Audit logging

Education

Master's degree in CS/Engineering

Tools

LangChain
LangGraph
Langfuse
Local model serving

Job description

Type: Full-time


Industry: Artificial Intelligence, Critical Infrastructure


About Company


A research-first AI company incubated at the Indian Institute of Science (IISc). The company is building advanced AI for the planning and operations of critical networks.


Its core technology is a World Model for critical networks - frontier AI, not another LLM wrapper. The founding team previously built and scaled a deep-tech company in the private 5G and cellular connectivity space.


Position Overview


The company's autonomous agents diagnose complex problems and plan actions across high-stakes infrastructure networks - and this role owns the platform that makes them run. As Senior Platform Engineer, you will build and own the core backend, data systems, and MLOps infrastructure that enable agents to operate reliably at scale, learn from every case, and remain fully observable and auditable. This is a deep-work IC role with direct reporting to the founders and outsized ownership over platform architecture from the ground up.


Role & Responsibilities



  • Own the agent serving and runtime path end-to-end, including model versioning, deployment pipelines, and rollback mechanisms using LangChain/LangGraph.

  • Build and maintain the tool and data integration layer connecting agents to external data sources, APIs, and domain-specific infrastructure systems.

  • Design and operate the learning and feedback pipeline - capturing agent outcomes, labeling signals, and closing the loop for continuous improvement.

  • Enforce platform governance: budget controls, rate limits, and approval gates for agent actions on critical infrastructure.

  • Build audit-grade tracing, observability, and logging systems (using Langfuse or equivalent) that meet real-world accountability requirements for autonomous action systems.

  • Own platform reliability, latency, and uptime against defined SLOs - including designing for retries, timeouts, and graceful degradation in distributed environments.

  • Collaborate closely with agent, data, and domain teams to ensure platform abstractions support evolving research and product requirements.


Must Have Criteria



  • 4+ years of backend engineering experience building and operating production-grade, scaled software systems end-to-end.

  • Hands-on experience operating production ML/AI platforms - model versioning, reproducible jobs, and safe rollout/rollback in live environments.

  • Strong data engineering skills: streaming pipelines, schedulers, SQL, object storage, and data provenance/quality practices.

  • Proficiency in async Python and FastAPI for building high-performance serving and API layers.

  • Experience with containerized, distributed deployments using Docker and Kubernetes on major cloud providers.

  • Demonstrated ability to build audit-grade logging and observability systems for systems that take real-world actions.

  • Distributed-systems fundamentals: retries, timeouts, circuit breakers, and graceful degradation under failure.


Nice to Have



  • Prior experience with LangChain, LangGraph, or similar agent orchestration frameworks.

  • Familiarity with Langfuse or other LLM tracing/observability tooling.

  • Experience serving open-source LLMs on-premises (local model serving infrastructure).

  • Master's degree in Computer Science, Engineering, or a related field.

  • Background in network infrastructure, telecom, or critical systems domains.


What We Offer



  • Founding-team-level ownership over platform architecture at a seed-stage, IISc-incubated AI company.

  • Direct collaboration with founders and exposure to frontier AI research at the intersection of World Models, LLMs, and critical infrastructure.

  • Deep-work engineering culture - small team, high autonomy, minimal process overhead.

  • Opportunity to shape engineering culture, tooling choices, and platform direction from zero.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead – Agentic AI Platform
Tech Lead – Agentic AI Platform

Multiscale AI • Hyderabad

On-site
INR 3,600,000 - 6,000,000
Senior AI/ML Engineer: LLM & Agent Stack
Senior AI/ML Engineer: LLM & Agent Stack

TrueFoundry • Bengaluru

On-site
INR 4,500,000 - 6,500,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Story Terrace Inc. • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Platform Engineer
Platform Engineer

SourcingXPress • Bengaluru

On-site
INR 7,604,000 - 11,407,000
Lunch and dinner provided
$200/month learning budget
$1,000/month tool experimentation budget
Founding Engineer
Founding Engineer

Blaugarnet Inc. • Pune District

Hybrid
INR 5,000,000 - 6,000,000
Founding equity
Senior Software Engineer - AI Platform Engineer
Senior Software Engineer - AI Platform Engineer

CloudBees • Chennai District

On-site
INR 3,500,000 - 5,500,000
Senior AI Engineer - Agentic AI & Knowledge Systems
Senior AI Engineer - Agentic AI & Knowledge Systems

Ontio AI • Pune District

On-site
INR 3,000,000 - 5,000,000
Applied AI Engineer
Applied AI Engineer

Simbian™, Inc. • Bengaluru

Hybrid
INR 1,200,000 - 2,400,000
Founding Engineer, AI Hyderabad · onsite · Full Time Apply →
Founding Engineer, AI Hyderabad · onsite · Full Time Apply →

CENNA Systems Inc. • Hyderabad

On-site
INR 1,200,000 - 2,400,000
SDE (Platform Engineer)
SDE (Platform Engineer)

Loadshare Networks • Chennai District, Bengaluru

Hybrid
INR 2,500,000 - 4,500,000