Sr. Software Engineer - AI Platform

sghcorp.com

Bengaluru

On-site

INR 4,000,000 - 7,500,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Penguin Solutions in Bengaluru is seeking a Senior AI Platform Engineer to build and operate the production AI platform that powers ClusterWareAI. You will develop infrastructure, services, and tooling to deploy, retrieve knowledge, execute workflows, and operate AI workloads at scale in enterprise environments.

You will work with AI engineers, platform engineers, and software developers to create secure, scalable, observable AI services for production deployments.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field.
  • 7+ years of software engineering, platform engineering, or infrastructure engineering experience.
  • Strong programming skills in Python.
  • Experience developing and operating production AI or LLM-powered applications.
  • Experience with Kubernetes, Docker, and cloud-native architectures.
  • Experience with model serving technologies such as vLLM, NVIDIA Triton Inference Server, or similar platforms.
  • Experience building REST APIs, microservices, and distributed systems.
  • Familiarity with vector databases, retrieval systems, and RAG architectures.

Responsibilities

  • Build and operate the production AI platform that supports model serving, inference, and AI services.
  • Develop scalable retrieval, knowledge management, and data ingestion pipelines to support AI-driven operations.
  • Build APIs, platform services, and automation that integrate AI capabilities into ClusterWareAI.
  • Implement AI observability, monitoring, evaluation, and operational tooling for production AI services.
  • Optimize inference performance, scalability, reliability, and operational efficiency.
  • Build and maintain CI/CD pipelines and deployment automation for AI applications.
  • Partner with AI, platform, and software engineering teams to deploy AI capabilities into production.
  • Contribute to platform security, reliability, and operational excellence.

Skills

Python
Kubernetes
Docker
LLM/AI platforms
REST APIs
Distributed systems
Vector databases

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

NVIDIA Triton Inference Server
vLLM
LangSmith
MLflow
OpenTelemetry
ClusterWareAI

Job description

Select how often (in days) to receive an alert:

At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.

Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.

Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.

At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.

Job Overview

We are seeking a Senior AI Platform Engineer to build and operate the production AI platform that powers ClusterWareAI. In this role, you will develop the infrastructure, services, and tooling that enable our AI Operational Agent to reliably deploy, retrieve knowledge, execute workflows, and operate at scale in enterprise environments.

You will work closely with AI engineers, platform engineers, and software developers to build secure, scalable, observable, and highly available AI services for production deployments.

Responsibilities

  • Build and operate the production AI platform that supports model serving, inference, and AI services.
  • Develop scalable retrieval, knowledge management, and data ingestion pipelines to support AI-driven operations.
  • Build APIs, platform services, and automation that integrate AI capabilities into ClusterWareAI.
  • Implement AI observability, monitoring, evaluation, and operational tooling for production AI services.
  • Optimize inference performance, scalability, reliability, and operational efficiency.
  • Build and maintain CI/CD pipelines and deployment automation for AI applications.
  • Partner with AI, platform, and software engineering teams to deploy AI capabilities into production.
  • Contribute to platform security, reliability, and operational excellence.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field.
  • 7+ years of software engineering, platform engineering, or infrastructure engineering experience.
  • Strong programming skills in Python.
  • Experience developing and operating production AI or LLM-powered applications.
  • Experience with Kubernetes, Docker, and cloud-native architectures.
  • Experience with model serving technologies such as vLLM, NVIDIA Triton Inference Server, or similar platforms.
  • Experience building REST APIs, microservices, and distributed systems.
  • Familiarity with vector databases, retrieval systems, and RAG architectures.

Preferred Qualifications

  • Experience implementing monitoring, observability, and operational tooling for AI services.
  • Experience with managed AI platforms such as Azure AI Foundry, AWS Bedrock, or Google Vertex AI.
  • Experience with AI engineering tools such as LangSmith, MLflow, OpenTelemetry, or similar technologies.
  • Experience deploying GPU-based AI workloads in Kubernetes environments.
  • Experience integrating AI services with enterprise applications, APIs, or infrastructure platforms.
  • Background in enterprise infrastructure, cloud platforms, or systems management software.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Software Engineer - AI Platform
Sr. Software Engineer - AI Platform

Penguin Solutions • Bengaluru

Hybrid
INR 4,500,000 - 7,500,000
Hybrid work model
Sr. Software Engineer - Applied AI
Sr. Software Engineer - Applied AI

sghcorp.com • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Sr. System Integration Engineer
Sr. System Integration Engineer

sghcorp.com • Bengaluru

On-site
INR 2,800,000 - 4,000,000
Sr. Software Engineer - Applied AI
Sr. Software Engineer - Applied AI

Penguin Solutions • Bengaluru

Hybrid
INR 600,000 - 900,000
Sr. Software Engineer - Applied AI
Sr. Software Engineer - Applied AI

Spore N Sprouts • Bengaluru

Hybrid
INR 2,500,000 - 4,500,000
Software Engineer II - Kubernetes
Software Engineer II - Kubernetes

sghcorp.com • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Sr. Software Engineer - Applied AI (Bangalore, IN, 560048)
Sr. Software Engineer - Applied AI (Bangalore, IN, 560048)

Penguin Computing • India

Hybrid
INR 4,000,000 - 7,000,000
Principal System Integration Engineer
Principal System Integration Engineer

sghcorp.com • Bengaluru

On-site
INR 5,500,000 - 8,000,000
Software Engineer II - Kubernetes
Software Engineer II - Kubernetes

Penguin Solutions • Bengaluru

Hybrid
INR 1,200,000 - 2,400,000
Sr. System Integration Engineer
Sr. System Integration Engineer

Penguin Solutions • Bengaluru

Hybrid
INR 300,000 - 520,000