Senior AI Platform Engineer

Lexsi Labs

Bengaluru

On-site

INR 450,000 - 800,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lexsi Labs is hiring a Senior AI Platform Engineer to build the platform layer that operationalizes Lexsi’s core AI systems. You will work across model pipelines, training systems, inference infrastructure, and backend platform architecture to deliver scalable, reliable platform capabilities.

This role emphasizes R&D-to-platform rollout, with emphasis on fine-tuning, RL, alignment, interpretability, agent execution, and inference optimization across LLMs and tabular models.

Qualifications

  • Strong experience building and shipping complex AI/ML systems in production.
  • Deep backend and platform engineering experience, especially in Python, distributed services, workflow orchestration, data systems, and cloud infrastructure.
  • Hands-on experience with fine-tuning systems, RL pipelines, inference infrastructure, distributed training, model serving, evaluation systems.
  • Strong understanding of systems implications of modern model workflows across LLMs, agents, and structured / tabular model systems.
  • Experience scaling workloads across clusters and production environments with reliability, observability, and performance.
  • Ability to work across research code, systems code, and product infrastructure without losing rigor.
  • Strong technical judgment around tradeoffs between model quality, infra complexity, scalability, interpretability, and operational cost.

Responsibilities

  • Own the platformization of Lexsi’s internal AI libraries, turning research-heavy systems into robust platform capabilities with stable APIs, execution layers, observability, and deployment paths.
  • Build and scale training and post-training infrastructure for workflows including SFT, RL, evaluation, model adaptation, and agent optimization.
  • Design the integration layer between research systems and product infrastructure, including job orchestration, artifact management, dataset versioning, experiment lineage, and runtime control surfaces.
  • Build inference systems that can support complex model behaviors under production constraints, including latency, throughput, cost efficiency, debuggability, and safety.
  • Design for multi-cluster and distributed execution, including scheduling, fault tolerance, checkpointing, retries, workload isolation, and heterogeneous compute environments.
  • Operationalize systems such as AlignTune for fine-tuning and RL pipelines, TabTune for tabular foundation model workflows, and DLBacktrace for interpretability, tracing, and behavioral inspection.
  • Build common platform primitives for model lifecycle management across training, evaluation, serving, rollback, and monitoring.
  • Partner closely with research teams to translate model-science complexity into production architecture without flattening away the core technical value.
  • Improve platform reliability for long-running and failure-prone AI workloads, especially where model behavior, system behavior, and infrastructure behavior interact in non-trivial ways.
  • Ensure that alignment, interpretability, and auditability are embedded into system design, especially for enterprise and regulated deployments where model outputs and decisions must be explainable.

Skills

Python
Distributed systems
Backend engineering
Model training
Platform infra

Job description

Lexsi Labs is a frontier AI lab building aligned, interpretable, and safe superintelligent systems. Our work spans alignment methods, interpretability-led system design, and foundational model research across LLMs, agents, and tabular / structured-data models.

A core part of our work is turning advanced AI research into production systems. That means taking internally developed libraries and model workflows, such as AlignTune, TabTune, and DLBacktrace, and integrating them into scalable platform infrastructure that supports training, inference, evaluation, observability, and enterprise deployment.

Lexsi Labs is a frontier AI lab building aligned, interpretable, and safe superintelligent systems. Our work spans alignment methods, interpretability-led system design, and foundational model research across LLMs, agents, and tabular / structured-data models.

A core part of our work is turning advanced AI research into production systems. That means taking internally developed libraries and model workflows, such as AlignTune, TabTune, and DLBacktrace, and integrating them into scalable platform infrastructure that supports training, inference, evaluation, observability, and enterprise deployment.

The Role

We are hiring a Senior AI Platform Engineer to build the platform layer that operationalizes Lexsi’s core AI systems.

This role is centered on R&D-to-platform rollout : taking technically sophisticated model systems and making them usable, reliable, and scalable inside the product stack. You will work across model pipelines, training systems, inference infrastructure, distributed execution, and backend platform architecture.

This is not a thin integration role. It requires strong engineering depth and enough model understanding to work effectively with systems involving fine-tuning, RL, alignment, interpretability, agent execution, and inference optimization across LLMs, agents, and tabular foundation models.

Responsibilities
  • Own the platformization of Lexsi’s internal AI libraries, turning research-heavy systems into robust platform capabilities with stable APIs, execution layers, observability, and deployment paths.
  • Build and scale training and post-training infrastructure for workflows including SFT, RL, evaluation, model adaptation, and agent optimization .
  • Design the integration layer between research systems and product infrastructure, including job orchestration, artifact management, dataset versioning, experiment lineage, and runtime control surfaces.
  • Build inference systems that can support complex model behaviors under production constraints, including latency, throughput, cost efficiency, debuggability, and safety.
  • Design for multi-cluster and distributed execution , including scheduling, fault tolerance, checkpointing, retries, workload isolation, and heterogeneous compute environments.
  • Operationalize systems such as AlignTune for fine-tuning and RL pipelines, TabTune for tabular foundation model workflows, and DLBacktrace for interpretability, tracing, and behavioral inspection.
  • Build common platform primitives for model lifecycle management across training, evaluation, serving, rollback, and monitoring.
  • Partner closely with research teams to translate model-science complexity into production architecture without flattening away the core technical value.
  • Improve platform reliability for long-running and failure-prone AI workloads, especially where model behavior, system behavior, and infrastructure behavior interact in non-trivial ways.
  • Ensure that alignment, interpretability, and auditability are embedded into system design, especially for enterprise and regulated deployments where model outputs and decisions must be explainable.
Example Problems You Might Work On
  • Turn AlignTune into a production-grade internal service for supervised fine-tuning and reinforcement learning across multiple model families, datasets, and evaluation loops.
  • Build rollout infrastructure for new model-science capabilities so research systems can be exposed safely and incrementally inside the platform.
  • Integrate DLBacktrace into training and inference pipelines so model behavior can be traced, debugged, and surfaced through internal and external product surfaces.
  • Build inference architecture for large models and agent systems that must balance cost, performance, explainability, and runtime control.
  • Design distributed execution flows across clusters for long-running training, evaluation, and analysis workloads with strong guarantees around recovery and reproducibility.
  • Unify workflows across LLMs, agents, and tabular models without collapsing their distinct operational and scientific requirements into a one-size-fits-none abstraction.
  • Build the platform interfaces that let downstream teams launch, inspect, evaluate, and deploy complex model workflows without needing to reimplement research infrastructure.
Requirements
  • Strong experience building and shipping complex AI / ML systems in production
  • Deep backend and platform engineering experience, especially in Python, distributed services, workflow orchestration, data systems, and cloud infrastructure
  • Hands‑on experience with one or more of: fine-tuning systems, RL pipelines, inference infrastructure, distributed training, model serving, evaluation systems
  • Strong understanding of the systems implications of modern model workflows across LLMs, agents, and structured / tabular model systems
  • Experience scaling workloads across clusters and production environments, with strong instincts around reliability, observability, and performance
  • Ability to work across research code, systems code, and product infrastructure without losing rigor at either layer
  • Strong technical judgment around the tradeoffs between model quality, infra complexity, scalability, interpretability, and operational cost
Strong Bonus Signals
  • Experience with alignment, interpretability, or AI safety systems
  • Experience with multi-cluster scheduling, inference optimization, or serving infrastructure for large models
  • Experience converting internal research frameworks into reusable platform capabilities
  • Experience debugging production failures caused by interactions between model behavior, orchestration systems, and infrastructure
  • Experience with agent runtimes, tool orchestration, long-horizon execution, or stateful model systems
What This Role Is Not
  • Not a prompt-engineering role
  • Not a glue-code integration role
  • Not a research-only role disconnected from deployment
  • Not a junior role

This role is for engineers who want to work on the hardest layer in applied AI: the boundary where model science, platform architecture, and production systems collide.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer
Senior AI Platform Engineer

Story Terrace Inc. • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Senior Software Engineer - Infrastructure
Senior Software Engineer - Infrastructure

Story Terrace Inc. • Mumbai

On-site
INR 1,500,000 - 2,500,000
Senior Technical Lead
Senior Technical Lead

Impetus • Dadri

On-site
INR 4,200,000 - 6,600,000
Sr Solution Sales - AI Platform
Sr Solution Sales - AI Platform

Lexsi Labs • Mumbai

On-site
INR 4,500,000 - 7,500,000
Product Manager (AI Platform)
Product Manager (AI Platform)

Lexsi Labs • Mumbai

On-site
INR 1,500,000 - 2,500,000
AI Platform Architect
AI Platform Architect

CLOUDSUFI • Dadri

On-site
INR 4,200,000 - 7,000,000
Technical Lead- Platform
Technical Lead- Platform

Juniper Square • India

On-site
INR 1,500,000 - 2,300,000
AI-ML Engineer
AI-ML Engineer

KanthamAi • Mumbai

On-site
INR 2,000,000 - 3,000,000
Platform Engineer
Platform Engineer

Recrew AI • Bengaluru

On-site
INR 4,200,000 - 7,000,000
Founding-team ownership over platform
Seed-stage exposure
Autonomy and deep-work culture
Principal Machine Learning Engineer
Principal Machine Learning Engineer

SourcingXPress • Hyderabad

On-site
INR 1,906,000 - 2,860,000