Senior Software Engineer - Infrastructure

Story Terrace Inc.

Mumbai

On-site

INR 1,500,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company seeks a Senior Infrastructure Software Engineer to architect and manage backend systems powering an innovative AI platform. The ideal candidate should possess strong Python skills and expertise with Docker and Kubernetes, along with a solid understanding of multi-cloud infrastructures. You will collaborate across teams ensuring scalable and reliable deployments. This role offers a dynamic environment focused on building production-grade systems in a cutting-edge field.

Qualifications

  • 3+ years of experience in backend engineering, platform engineering, DevOps, or infrastructure roles.
  • Strong Python expertise with experience in production services and frameworks.
  • Deep experience with Docker and Kubernetes in production environments.

Responsibilities

  • Design and build Python-based backend services for the AI platform.
  • Architect and operate AI infrastructure at production scale.
  • Implement Infrastructure as Code using tools like Terraform.

Skills

Python expertise
Infrastructure as Code
DevOps practices
Containerization
Multi-cloud infrastructure
Distributed systems

Education

3+ years in backend engineering or related roles

Tools

Docker
Kubernetes
Terraform
FastAPI

Job description

LexsiLabs is one of the leading frontier labs focused on building aligned, interpretable, and safe Superintelligence. Our work spans efficient alignment methods, interpretability‑led system design, and scalable AI platforms that operate reliably across enterprise and regulated environments. Our mission is to build AI systems that are powerful, transparent, and production‑grade by design.

Our team operates with deep technical ownership, minimal hierarchy, and a strong bias toward building systems that work at scale. At Lexsi.ai, infrastructure is not support work. It is a core product capability.

As a Senior Infrastructure Software Engineer, you will architect, build, and operate the core backend and deployment systems that power the Lexsi AI platform. You will own multi‑cloud, serverless, and stateless AI deployments, ensuring our systems scale seamlessly across environments of any size while maintaining correctness, reliability, and cost‑efficiency.

This role is ideal for someone who thinks like a SDE+ Platform Engineer+ DevOps, and takes pride in owning end‑to‑end systems.

Responsibilities
  • Design and build Python‑based backend services that power core platform functionality and AI workflows.
  • Architect and operate AI/LLM inference and serving infrastructure at production scale.
  • Build stateless, serverless, and horizontally scalable systems that can run across environments of varying sizes.
  • Design multi‑cloud infrastructure across AWS, Azure, and GCP with portability and reliability as first‑class goals.
  • Deploy and manage containerized workloads using Docker, Kubernetes, ECS, or equivalents.
  • Build and operate distributed compute systems for AI workloads, including inference‑heavy and RL‑style execution patterns.
  • Implement Infrastructure as Code using Terraform, CloudFormation, Pulumi, or similar tools.
  • Own CI/CD pipelines for backend services, infrastructure, and AI workloads.
  • Optimize GPU and compute usage for performance, cost, batching, and autoscaling.
  • Define and enforce reliability standards targeting 99%+ uptime across critical services.
  • Build observability systems for latency, throughput, failures, and resource utilization.
  • Implement security best practices across IAM, networking, secrets, and encryption.
  • Support compliance requirements (SOC2, ISO, HIPAA) through system design and evidence‑ready infrastructure.
  • Lead incident response, root‑cause analysis, and long‑term reliability improvements.
  • Collaborate closely with ML engineers, product, and leadership to translate AI requirements into infrastructure design.
Required Qualifications
  • 3+ years of hands‑on experience in backend engineering, platform engineering, DevOps, or infrastructure‑focused SDE roles.
  • Strong Python expertise with experience building and running production backend services.
  • Experience with Python frameworks such as FastAPI, Django, Flask, or equivalent.
  • Deep hands‑on experience with Docker and Kubernetes in production environments.
  • Practical experience designing and operating multi‑cloud infrastructure.
  • Strong understanding of Infrastructure as Code and declarative infrastructure workflows.
  • Experience building and maintaining CI/CD pipelines for complex systems.
  • Solid understanding of distributed systems, async processing, and cloud networking.
  • Strong ownership mindset with the ability to build, run, debug, and improve systems end‑to‑end.
Nice to Have
  • Experience deploying AI or LLM workloads in production environments.
  • Familiarity with model serving frameworks such as KServe, Kubeflow, Ray, etc.
  • Experience running GPU workloads, inference batching, and rollout strategies.
  • Exposure to serverless or hybrid serverless architectures for AI systems.
  • Experience implementing SLO/SLA‑driven reliability and monitoring strategies.
  • Prior involvement in security audits or compliance‑driven infrastructure work.
  • Contributions to open‑source infrastructure or platform projects.
  • Strong system design documentation and architectural reasoning skills.
What Success Looks Like
  • Lexsi AI platforms scale cleanly across cloud providers and deployment sizes.
  • Inference systems are reliable, observable, and cost‑efficient under real load.
  • Engineering teams ship faster because infrastructure is predictable and well‑designed.
  • Incidents are rare, understood deeply when they occur, and lead to durable fixes.
  • Infrastructure decisions support long‑term platform scalability, not short‑term hacks.
Next Steps & Interview Process
  • Take‑home assignment focused on real infrastructure and scaling problems.
  • Deep technical interview covering system design, failure modes, and trade‑offs.
  • Final discussion focused on ownership, reliability mindset, and execution style.

We avoid process theatre. If you can design systems that don’t fall apart under pressure, we’ll move fast.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer
Senior AI Platform Engineer

Lexsi Labs • Bengaluru

On-site
INR 450,000 - 800,000
Sr Solution Sales - AI Platform
Sr Solution Sales - AI Platform

Lexsi Labs • Mumbai

On-site
INR 4,500,000 - 7,500,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Story Terrace Inc. • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Technical Content Lead
Technical Content Lead

Lexsi Labs • India

On-site
INR 1,200,000 - 1,800,000
Developer Advocate
Developer Advocate

Lexsi Labs • India

On-site
INR 800,000 - 1,200,000
Product Manager (AI Platform)
Product Manager (AI Platform)

Lexsi Labs • Mumbai

On-site
INR 1,500,000 - 2,500,000
Senior Technical Lead
Senior Technical Lead

Impetus • Dadri

On-site
INR 4,200,000 - 6,600,000
Staff DevOps Engineer
Staff DevOps Engineer

Sia Partners' • Mumbai

On-site
INR 1,500,000 - 2,500,000
Opportunity to lead cutting-edge AI projects
Dynamic and collaborative team environment
Staff DevOps Engineer
Staff DevOps Engineer

Sia • Mumbai

On-site
INR 1,800,000 - 3,000,000
Opportunity to lead AI projects
Collaborative team environment
Equal opportunity employer
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,200,000
Significant equity in a venture-backed company
Opportunity to work with modern tech stack