Platform Engineer

Recrew AI

Bengaluru

On-site

INR 4,200,000 - 7,000,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Founding-team ownership over platform
Seed-stage exposure
Autonomy and deep-work culture

Job summary

Recrew AI in Bengaluru is seeking a Senior Platform Engineer to own the backend, data systems and MLOps infrastructure that powers autonomous agents for critical networks. This deep-work IC role reports to the founders and carries outsized ownership over platform architecture from ground up.

You will end-to-end manage agent serving paths, data integration, learning pipelines, and observability. The role requires hands-on experience with Docker/Kubernetes, async Python, and scalable distributed

Qualifications

  • 4+ years backend engineering experience building production-grade systems.
  • Hands-on experience operating production ML/AI platforms with safe rollout/rollback.
  • Strong data engineering skills: streaming pipelines, SQL, object storage, data quality.
  • Proficiency in async Python and FastAPI for high-performance serving and APIs.
  • Experience with Docker and Kubernetes in cloud deployments.
  • Auditable logging and observability for real-world autonomous actions.
  • Distributed-systems fundamentals: retries, timeouts, circuit breakers.

Responsibilities

  • Own the agent serving and runtime path end-to-end, including model versioning and deployment pipelines.
  • Build and maintain the data integration layer to external sources and domain systems.
  • Design and operate the learning/feedback pipeline and continuous improvement loop.
  • Enforce platform governance: budgets, rate limits, approval gates for actions.
  • Build audit-grade tracing and logging systems meeting accountability requirements.
  • Ensure platform reliability, latency, and uptime against defined SLOs.
  • Collaborate with agent, data, and domain teams to support evolving research and product needs.

Skills

Backend engineering
Async Python
FastAPI
Distributed systems
Data engineering
ML/AI platforms
Observability/logging
Cloud deployments

Education

Bachelor's degree in CS/Engineering

Tools

Docker
Kubernetes
LangChain/LangGraph
Langfuse
Open-source LLMs on-prem

Job description

Type: Full-time

Industry: Artificial Intelligence, Critical Infrastructure

About Company

A research-first AI company incubated at the Indian Institute of Science (IISc). The company is building advanced AI for the planning and operations of critical networks.

Its core technology is a World Model for critical networks — frontier AI, not another LLM wrapper. The founding team previously built and scaled a deep-tech company in the private 5G and cellular connectivity space.

Position Overview

The company's autonomous agents diagnose complex problems and plan actions across high-stakes infrastructure networks — and this role owns the platform that makes them run. As Senior Platform Engineer, you will build and own the core backend, data systems, and MLOps infrastructure that enable agents to operate reliably at scale, learn from every case, and remain fully observable and auditable. This is a deep-work IC role with direct reporting to the founders and outsized ownership over platform architecture from the ground up.

Role & Responsibilities
  • Own the agent serving and runtime path end-to-end, including model versioning, deployment pipelines, and rollback mechanisms using LangChain/LangGraph.
  • Build and maintain the tool and data integration layer connecting agents to external data sources, APIs, and domain-specific infrastructure systems.
  • Design and operate the learning and feedback pipeline — capturing agent outcomes, labeling signals, and closing the loop for continuous improvement.
  • Enforce platform governance: budget controls, rate limits, and approval gates for agent actions on critical infrastructure.
  • Build audit-grade tracing, observability, and logging systems (using Langfuse or equivalent) that meet real-world accountability requirements for autonomous action systems.
  • Own platform reliability, latency, and uptime against defined SLOs — including designing for retries, timeouts, and graceful degradation in distributed environments.
  • Collaborate closely with agent, data, and domain teams to ensure platform abstractions support evolving research and product requirements.
Must Have Criteria
  • 4+ years of backend engineering experience building and operating production-grade, scaled software systems end-to-end.
  • Hands-on experience operating production ML/AI platforms — model versioning, reproducible jobs, and safe rollout/rollback in live environments.
  • Strong data engineering skills: streaming pipelines, schedulers, SQL, object storage, and data provenance/quality practices.
  • Proficiency in async Python and FastAPI for building high-performance serving and API layers.
  • Experience with containerized, distributed deployments using Docker and Kubernetes on major cloud providers.
  • Demonstrated ability to build audit-grade logging and observability systems for systems that take real-world actions.
  • Distributed-systems fundamentals: retries, timeouts, circuit breakers, and graceful degradation under failure.
Nice to Have
  • Prior experience with LangChain, LangGraph, or similar agent orchestration frameworks.
  • Familiarity with Langfuse or other LLM tracing/observability tooling.
  • Experience serving open-source LLMs on-premises (local model serving infrastructure).
  • Master's degree in Computer Science, Engineering, or a related field.
  • Background in network infrastructure, telecom, or critical systems domains.
What We Offer
  • Founding-team-level ownership over platform architecture at a seed-stage, IISc-incubated AI company.
  • Direct collaboration with founders and exposure to frontier AI research at the intersection of World Models, LLMs, and critical infrastructure.
  • Deep-work engineering culture — small team, high autonomy, minimal process overhead.
  • Opportunity to shape engineering culture, tooling choices, and platform direction from zero.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Engineer: LLM & Agent Stack
Senior AI/ML Engineer: LLM & Agent Stack

TrueFoundry • Bengaluru

On-site
INR 4,500,000 - 6,500,000
AI/ML Engineer
AI/ML Engineer

Recrew AI • Bengaluru Urban

On-site
INR 1,200,000 - 1,800,000
Ownership of agentic AI systems
Access to latest LLM models and AI工具
Global delivery network collaboration
+1
Senior AI Platform Engineer
Senior AI Platform Engineer

Story Terrace Inc. • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Founding Engineer, AI Hyderabad · onsite · Full Time Apply →
Founding Engineer, AI Hyderabad · onsite · Full Time Apply →

CENNA Systems Inc. • Hyderabad

On-site
INR 1,200,000 - 2,400,000
AI Engineer
AI Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI/ML Engineer – Agentic AI & LLM Systems
AI/ML Engineer – Agentic AI & LLM Systems

Recrew AI • Bengaluru Urban

On-site
INR 1,500,000 - 2,400,000
Senior Technical Lead
Senior Technical Lead

Impetus • Dadri

On-site
INR 4,200,000 - 6,600,000
Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence Inc. • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Platform Engineer
Platform Engineer

SourcingXPress • Bengaluru

On-site
INR 7,604,000 - 11,407,000
Lunch and dinner provided
$200/month learning budget
$1,000/month tool experimentation budget
Senior AI Platform Engineer
Senior AI Platform Engineer

Lexsi Labs • Bengaluru

On-site
INR 450,000 - 800,000