AI DevOps Engineer

Zoho

Hyderabad

On-site

INR 3,500,000 - 5,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Anthrobyte is seeking a founding Platform Engineer to own the full AI infrastructure, MLOps pipelines, model serving, observability, and cloud architecture. You will shape a production-grade platform and work with AI engineers, product leadership, and enterprise clients to deploy scalable AI transformations.

You will lead platform strategy, mentor a DevOps team, and define how Anthrobyte’s AI infrastructure scales across client portfolios. This is a founding role with ownership across the stack.

Qualifications

  • 4–6 years in DevOps or platform engineering, with at least 1–2 years specifically in AI or ML infrastructure
  • Demonstrated hyperscale experience — infrastructure supporting millions of daily requests, petabyte-scale data, or multi-region distributed systems
  • Deep Kubernetes expertise: GPU node pools, resource quotas, network policies, and production incident ownership
  • Hands-on production experience with at least one major LLM serving stack (vLLM, Triton, TGI, Ray Serve, or BentoML)
  • Strong Python and scripting capability — you write automation, not just configuration
  • Proficiency with IaC (Terraform preferred) and GitOps workflows as standard practice
  • Cloud practitioner depth in AWS, GCP, or Azure — compute, networking, and storage for AI workloads
  • Clear communicator — able to translate infrastructure complexity into language that resonates with engineers, product leads, and enterprise clients alike

Responsibilities

  • Design and operate AI infrastructure across multi-cloud environments supporting LLM inference, fine-tuning, and RAG pipelines at production scale
  • Architect GPU cluster management and optimise inference throughput using relevant serving frameworks
  • Own infrastructure-as-code — reproducible, version-controlled, disaster-recovery-ready environments across deployments
  • Drive multi-region, high-availability architecture decisions reflecting enterprise reliability standards
  • Build and maintain MLOps platforms — model versioning, experiment tracking, automated retraining, and deployment pipelines
  • Implement CI/CD pipelines that support rapid model iteration with production stability
  • Define promotion workflows from development to staging to production with rollback and canary strategies
  • Lead container orchestration at scale — Kubernetes, Helm, service mesh, autoscaling for AI workloads
  • Configure GPU node pools, resource quotas, taints and tolerations, network policies for secure AI workloads
  • Own production incident response — from detection to resolution and post-mortem

Skills

AI infrastructure
Cloud architecture
Platform engineering
Observability
Python scripting

Tools

Kubernetes
Terraform
Pulumi
GitOps
Kubeflow
MLflow
ArgoCD
AWS
GCP
Azure

Job description

A founding infrastructureleadership role.


Anthrobyte builds enterprise AIsystems that move organisations from pilot to production-grade adoption. As wescale, we need a platform engineer who can own the full infrastructure vision:model serving, MLOps pipelines, GPU cluster management, observability, andcloud architecture — all as one coherent, production-grade system.


This is less a traditionalDevOps role and more a founding platform seat. You will work directly with AIengineers, product leadership, and enterprise clients to define how AItransformation is deployed, scaled, and made reliable inside complex organisations.You will be the person who makes AI products real.


Own the full platform layer— AI infrastructure, MLOps pipelines, model serving, observability, and cloudarchitecture — as one coherent, production-grade system.


Lead/Architect/Head of Platform Engineering


Grow into engineeringleadership — shaping platform strategy, building and mentoring a DevOps team,and defining how AI infrastructure scales with Anthrobyte's client portfolio.


Requirements

Own the platform. Power thetransformation.



  • Design and operate AI infrastructure acrossmulti-cloud environments (AWS, GCP, Azure) supporting LLM inference,fine-tuning, and RAG pipelines at production scale

  • Architect GPU cluster management and optimiseinference throughput using vLLM, Triton Inference Server, TensorRT, orequivalent serving frameworks

  • Own infrastructure-as-code (Terraform, Pulumi) —reproducible, version-controlled, disaster-recovery-ready environments acrossall deployments

  • Drive multi-region, high-availabilityarchitecture decisions that reflect the reliability standards enterpriseclients require


MLOps& Model Lifecycle



  • Build and maintain MLOps platforms — modelversioning, experiment tracking, automated retraining, and deployment pipelinesusing MLflow, Kubeflow, or equivalent

  • Implement CI/CD pipelines (GitHub Actions,ArgoCD, Tekton) that support rapid model iteration without sacrificingproduction stability

  • Define promotion workflows from development tostaging to production — with rollback, canary, and blue-green strategies asstandard practice

  • Lead container orchestration at scale —Kubernetes (EKS/GKE/AKS), Helm charts, service mesh configuration, andauto-scaling strategies for variable AI workloads

  • Configure GPU node pools, resource quotas,taints and tolerations, and network policies for secure, efficient AI workloadscheduling

  • Own production incident response — fromdetection through resolution to post-mortem and systemic fix


Observability& FinOps



  • Own observability end-to-end: latency, GPUutilisation, cost-per-inference, model drift detection, and SLO/SLA dashboards(Prometheus, Grafana, or equivalent)

  • Lead FinOps strategy for GPU compute — spotinstance management, reserved capacity planning, cost attribution across teamsand client engagements

  • Surface infrastructure cost and reliability datato leadership and clients in clear, actionable terms


Security,Governance & Compliance



  • Enforce security and data governance standardsacross AI deployments — access controls, audit logging, secret management, andPII handling in inference pipelines

  • Support enterprise client compliancerequirements — including data residency, model access controls, and audit traildocumentation


Cross-FunctionalPartnership



  • Translate AI engineer and product requirementsinto platform specifications — and push back with alternatives whenrequirements are unrealistic or unsafe

  • Partner with client engineering teams duringenterprise AI deployments, acting as the technical infrastructure authority


The Growth Pathway


Demonstrate consistentexcellence as a platform engineering leader and the scope expands. You willgrow into Lead/Architect/Head of Platform Engineering — building andmentoring a DevOps and MLOps team, shaping infrastructure strategy acrossAnthrobyte's full client portfolio, and defining what production-gradeenterprise AI deployment looks like at scale. This is not a title — it is alevel of ownership that must be earned and continually re-earned.


The profile we are searching for.


You think across the full stack— from YAML to architecture, from GPU cost to enterprise reliability SLAs. Youare the kind of engineer who has felt the weight of a production incident at2am and built the systems that prevent the next one. You are as comfortablepresenting infrastructure trade-offs to a CTO as you are deep in a Terraformmodule.


Youbring:



  • 4–6 years in DevOps or platform engineering,with at least 1–2 years specifically in AI or ML infrastructure

  • Demonstrated hyperscale experience —infrastructure supporting millions of daily requests, petabyte-scale data, ormulti-region distributed systems

  • Deep Kubernetes expertise: GPU node pools,resource quotas, network policies, and production incident ownership

  • Hands-on production experience with at least onemajor LLM serving stack (vLLM, Triton, TGI, Ray Serve, or BentoML)

  • Strong Python and scripting capability — youwrite automation, not just configuration

  • Proficiency with IaC (Terraform preferred) andGitOps workflows as standard practice

  • Cloud practitioner depth in AWS, GCP, or Azure —particularly compute, networking, and storage for AI workloads

  • Clear communicator — able to translateinfrastructure complexity into language that resonates with engineers, productleads, and enterprise clients alike


Bonussignals:



  • Edge /On-Premise LLM Serving

  • AI Consultancy Experience

  • Regulated Industry Deployments



  • Multi-CloudArchitecture

  • MLflow or Kubeflow Ownership

  • Startup 0→1 Environment


A rare kind of opportunity.



  • Founding platform engineering ownership —greenfield infrastructure built your way, with your architectural decisions

  • Direct access to engineering and productleadership from day one

  • The mandate to build the platform function theway it should be built — AI-native, observable, and enterprise-reliable

  • Active involvement in enterprise AI deploymentengagements — real infrastructure challenges, real clients, real consequences

  • Access to GPU compute resources, premium cloudcredits, and AI tooling subscriptions

  • Competitive compensation benchmarked to seniorengineering market rates in Hyderabad

  • A culture where great infrastructure work isvisible and celebrated — not invisible

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI FDE Lead/Architect - Enterprise Solutions
AI FDE Lead/Architect - Enterprise Solutions

Zoho • Hyderabad

On-site
INR 900,000 - 1,300,000
Tech Lead – Agentic AI Platform
Tech Lead – Agentic AI Platform

Multiscale AI • Hyderabad

On-site
INR 3,600,000 - 6,000,000
Senior AI Engineer
Senior AI Engineer

Blaugarnet Inc. • Pune District

Hybrid
INR 3,000,000 - 5,000,000
Platform Engineer
Platform Engineer

SourcingXPress • Bengaluru

On-site
INR 7,604,562 - 11,406,843
Lunch and dinner provided
$200/month learning budget
$1,000/month tool experimentation budget
Founding Engineer
Founding Engineer

Blaugarnet Inc. • Pune District

On-site
INR 5,000,000 - 6,000,000
Founding equity
Sr. AI Engineer
Sr. AI Engineer

Viamagus Technologies • Hyderabad

On-site
INR 4,000,000 - 6,500,000
Lead Engineer
Lead Engineer

Arcadia • Chennai District

On-site
INR 4,000,000 - 7,000,000
Competitive compensation
Hybrid model with remote-first policy
Flexible Leave Policy
+7
Lead Site Reliability Engineer or Platform Engineer
Lead Site Reliability Engineer or Platform Engineer

Weekday 1 • Bengaluru

On-site
INR 5,000,000 - 10,000,000
AI Infrastructure / DevOps Engineer (AI-Native, Agentic) 5–10 Years
AI Infrastructure / DevOps Engineer (AI-Native, Agentic) 5–10 Years

Sprouts.ai • Chandigarh

On-site
INR 1,200,000 - 1,800,000
Ownership of infra decisions
Fast track to leadership roles
Exposure to frontier AI systems
AI Technical Lead
AI Technical Lead

Benchmarkit • Pune District

On-site
INR 4,000,000 - 7,000,000