Associate Architect - Platform

Quantiphi

Bengaluru

On-site

INR 1,500,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions provider in Bengaluru is looking for an Associate Architect specializing in MLOps/LLMOps. This mid-senior level role requires expertise in design and implementation of GPU computing clusters and Kubernetes. The ideal candidate will have 6-8 years of experience in ML infrastructure, a collaborative spirit, and proficiency in Python and Bash scripting. The position offers full-time employment with opportunities to innovate in the AI domain.

Qualifications

  • Over 6-8 years of experience working with developing ML Infrastructure.
  • Hands-on experience in building services on Kubernetes.
  • Understanding of GPU computing and AI infrastructure.

Responsibilities

  • Design and implement GPU compute clusters.
  • Develop automated testing strategies for commerce flows.
  • Troubleshoot system failures and optimize cluster operations.

Skills

Kubernetes
GPU computing
Python programming
Bash scripting
Docker
Collaboration
Open-source contribution

Tools

Terraform
MySQL
Ansible

Job description

Talent Acquisition Lead for Nvidia AI Business @ Quantiphi Analytics – Strategic Hiring - Product Engineering, R&D, GTM and Marketing

Role : Associate Architect - MLOps / LLMOps

Experience : 6 to 8 Years

Job Summary:

Join our dynamic team as a Platform Architect and leverage your expertise in production-scale platforms within the GenAI or ML domain. In this role, you'll be instrumental in designing, developing and maintaining cutting-edge build and test environments for critical GenAI workloads running on foundational cloud infrastructure.

You'll partner with architects to design and implement highly robust and scalable systems, while also providing crucial development support to SRE/Operations teams as they tackle complex distributed systems challenges at scale. We're seeking an engineer who champions Quantiphi's dedication to Cloud-Native development, with a particular emphasis on Kubernetes.

Job Responsibilities:
  • Designing and implementing state-of-the-art GPU compute clusters to support critical workloads.
  • Designing comprehensive automated testing strategies and frameworks across unit, integration, API, and end-to-end levels for critical commerce flows.
  • Developing robust performance testing frameworks to validate platform scalability, resilience, and identify optimization opportunities.
  • Planning of comprehensive monitoring solutions with alerting systems to track platform health and ensure SLA compliance.
  • Designing specialized test frameworks for security controls and ensuring compliance validation across payment and personal data.
  • Architecting a scalable automation infrastructure that supports growing platform capabilities with consistent test environments.
  • Troubleshooting, diagnosing, and performing root cause analysis of system failures, isolating components and failure scenarios in collaboration with internal and external partners.
  • Optimizing cluster operations for maximum reliability, efficiency, and performance.
Job Requirements:
  • Over 6-8 years of experience working with developing ML Infrastructure.
  • Over 3 years of hands‑on experience in large-scale direct experience building and deploying production‑ready services on Kubernetes.
  • A proven history of engaging with and contributing to open‑source projects.
  • A collaborative spirit, demonstrated by prior work developing scalable software solutions for cloud services.
  • The ability to effectively communicate complex technical designs and quality approaches across various mediums.
  • A deep understanding of GPU computing and AI infrastructure.
  • A strong passion for solving complex technical challenges and optimizing system performance.
  • Working knowledge of cluster configuration management tools such as BCM or Ansible, and infrastructure‑level applications including Kubernetes, Terraform, and MySQL.
  • In‑depth understanding of container technologies like Docker and Containers.
  • Proficiency in programming with Python and Bash scripting.
Ways To Stand Out From The Crowd:
  • Significant experience with sophisticated infrastructure tooling, including Kubernetes Cluster API, Terraform, Helm, and Operator Framework.
  • Practical, production‑level experience across major cloud platforms: Azure, Google Cloud Platform (GCP), or Amazon Web Services (AWS).
  • Ability to adapt to new technologies and Frameworks in ML/GenAI landscape.
  • A strong track record of successfully refactoring and optimising software for deployment within Kubernetes environments.
  • Comfort discussing and working with core Kubernetes concepts like CSI, CNI, and CRI.
  • Comprehensive understanding of the CNCF landscape and its associated tooling.
  • The ability to decompose complex problems into simpler sub‑problems and leverage existing solutions for efficient implementation, along with designing simple, self‑sustaining systems.
  • Experience leveraging AI/ML to proactively detect and resolve incidents, automate alert triaging, perform log analysis, and streamline repetitive workflows.
Seniority level

Mid‑Senior level

Employment type

Full‑time

Job function

Information Technology

Industries

IT Services and IT Consulting and Software Development

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Architect - DevOps and ML Op's
Senior Architect - DevOps and ML Op's

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 3,500,000 - 6,000,000
AI Platform Architect
AI Platform Architect

HCLTech • Dadri

On-site
INR 4,200,000 - 7,000,000
Senior Engineer - AI Platform
Senior Engineer - AI Platform

NetConnectGlobal • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Senior DevOps Engineer (Kubernetes & AI Infra)
Senior DevOps Engineer (Kubernetes & AI Infra)

Navikenz • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Architect / Sr. Architect – GenAI & AI Platforms (Data and AI)
Architect / Sr. Architect – GenAI & AI Platforms (Data and AI)

Tiger Analytics • Chennai

On-site
INR 4,000,000 - 8,000,000
MLOps Architect
MLOps Architect

Anblicks • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Cutting-edge technologies
Mentorship opportunities
Impactful initiatives in AI and Data
Lead Platform Engineer/Lead DevOps Engineer
Lead Platform Engineer/Lead DevOps Engineer

NexTurn Inc. • Hyderabad

On-site
INR 4,500,000 - 6,000,000
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure

NVIDIA Gruppe • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Generous benefits package