Staff Software Engineer - AI

Tekion

Bengaluru

On-site

INR 3,200,000 - 5,200,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Tekion seeks an experienced platform/ML engineer to build and operate the LLM control plane and gateway in a multi-tenant SaaS setup. You will design smart routing, quotas, and cost tracking while delivering a unified REST/gRPC API and SDKs for product teams.

Emphasis on safety, governance, and cross-vendor LLM integration. You will own the agent runtime, orchestration patterns, and data pipelines, enabling real-time context and retrieval with hybrid search.

Qualifications

  • 8+ years building large-scale data/ML or platform systems.
  • Production experience with Python plus one of Java/Scala/Go; microservices and API design.
  • MLOps at scale: pipelines (Airflow/Kubeflow), tracking/registry (MLflow), CI/CD for models, A/B testing, shadow/canary.
  • Cloud and containers: AWS (preferred), plus Docker/Kubernetes; multitenant SaaS considerations.
  • Practical ML knowledge (feature engineering, training, evaluation, drift detection).
  • Built or operated an LLM gateway/control plane: provider adapters, routing/policies, caching, quota/ratelimit, cost.

Responsibilities

  • Build and run the LLM control plane/gateway: smart routing, rate limits/quotas, failover, and token/cost tracking.
  • Ship a unified API and SDKs (REST/gRPC) with normalized schemas, structured outputs, caching, and observability.
  • Enforce safety and privacy by default: content filtering, prompt validation, and PII redaction.
  • Enable multimodal, multivendor LLMs with automated canarying and versioning.
  • Own the agent runtime: tool registry, permissions, function calling, grounding, and retrieval.
  • Design orchestration patterns and manage agent state and long-running workflows.
  • Enable platform components for training and scoring pipelines for ML models; track experiments.
  • Create components to monitor model/data drift and retrain/tune for accuracy.
  • Add human-in-the-loop review and safe actioning before agents touch dealer systems.
  • Evolve domain graph and entity resolution; build data ingestion pipelines.
  • Provide real-time context to agents with access controls.
  • Power retrieval with hybrid search and smart cache to balance accuracy and cost.
  • Run continuous offline/online evaluations for quality and safety.
  • Define SLOs for latency, uptime, and cost view; enable autoscaling.
  • Maintain model/agent registry, versioning, audits, and reproducibility; ensure compliance.
  • Provide templates/CLIs, sandboxes, and docs; mentor engineers on MLOps and AI safety.

Skills

Python
Java/Scala/Go
MLOps
Airflow/Kubeflow
ML pipelines
Docker/Kubernetes
Cloud AWS
Knowledge graph
GraphQL
Vector search
Hybrid retrieval
LLM gateway
Agent systems

Tools

MLflow
Airflow
Kubeflow
Spark
Flink
Kafka
Neo4j
Qdrant
Milvus
pgvector

Job description

Responsibilities:
  • Build and run the LLM control plane/gateway: smart routing, rate limits/quotas, failover, and token/cost tracking.
  • Ship a unified API and SDKs (REST/gRPC) with normalized schemas, structured outputs, caching, and full observability (traces/logs/metrics).
  • Enforce safety and privacy by default: content filtering, prompt/response validation, and PII redaction.
  • Enable multimodal, multivendor use of LLMs with automated canarying and versioning.
  • Own the agent runtime: tool registry, permissions, function calling, grounding, and retrieval.
  • Design orchestration patterns (sequential, planner-executor, streaming) and manage agent state and long-running workflows.
  • Enable platform components for training and scoring pipelines for classical ML (e. g., XGBoost/LightGBM/linear/trees) and deep models; standardize experiment tracking and packaging.
  • Create components to monitor model and data drift, retraining and tuning models as needed to maintain accuracy and relevance.
  • Add human-in-the-loop review and safe actioning before agents touch dealer systems.
  • Evolve the domain graph and entity resolution; build reliable data ingestion pipelines.
  • Serve real-time context to agents (profiles, inventory, pricing, appointments, service history) with access controls and lineage.
  • Power retrieval with hybrid search (graph + vector + keyword) and smart cache/TTL to balance accuracy, latency, and cost.
  • Run continuous offline/online evaluations for quality, factuality, bias, and safety for platform sanity.
  • Define SLOs for latency (p50/p95), uptime, and cost view capabilities; enable autoscaling and spend controls.
  • Maintain a model/agent registry, versioning, approvals, audit trails, and reproducibility; support compliance where needed.
  • Provide templates/CLIs, sandboxes, and docs so product teams can build and ship fast; mentor engineers and champion MLOps and AI safety best practices.
Requirements:
  • 8+ years building large-scale data/ML or platform systems; strong software engineering fundamentals (abstracted API design, concurrency, distributed systems).
  • Production experience with Python plus one of Java/Scala/Go; microservices and API design.
  • MLOps at scale: pipelines (Airflow/Kubeflow), tracking/registry (MLflow), CI/CD for models, A/B testing, shadow/canary, and online feature computation (Spark/Flink/Kafka).
  • Cloud and containers: AWS (preferred), plus Docker/Kubernetes; performance, reliability, and cost engineering in multitenant SaaS.
  • Practical ML knowledge (feature engineering, training, evaluation, and drift detection); experience deploying models that power user-facing workflows.
  • Built or operated an LLM gateway/control plane: provider adapters, routing/policies, caching, quota/ratelimit, cost, and token accounting.
  • Agentic systems: tool use/function calling, orchestration frameworks, human-in-the-loop, safety/guardrails, and online evaluation/telemetry.
  • Graph and retrieval: knowledge graphs (e. g., Neo4j/Neptune/TigerGraph), GraphQL, vector search (e. g., pgvector/Qdrant/Milvus), and hybrid retrieval patterns.
Preferred Mindset:
  • Platform-as-product: obsess over developer experience, paved roads, and clear SLAs.
  • Thinks in systems; observability, fallback, and access control are core, not afterthoughts.
  • Passionate about AI; enjoys enabling real-world LLM and agentic use cases.
  • Cost-aware builder: you treat latency and dollars as first-class metrics and design for graceful degradation.
  • Vendor-agnostic thinker: choose the right model/provider per use case; build for portability and resilience.
  • Documentation and teaching: you make complex systems understandable; you uplevel teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer
ML Engineer

Tekion • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
AI Platform Engineer
AI Platform Engineer

e-Hireo • Bengaluru

On-site
INR 400,000 - 650,000
Lead AI Engineer
Lead AI Engineer

Keka Technologies Private Limited • Nagar

On-site
INR 1,500,000 - 2,100,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Aptus Data Labs • Bengaluru

Hybrid
INR 1,500,000 - 2,100,000
Sr. AI Engineer (SLM and LLM)
Sr. AI Engineer (SLM and LLM)

NexTurn • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Sr AI/ML Engineer
Sr AI/ML Engineer

Staples India • Chennai District

On-site
INR 3,200,000 - 5,200,000
AI ML Platform Engineer
AI ML Platform Engineer

The Standard India • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Senior Agentic AI Developer
Senior Agentic AI Developer

Coretek-Services • Telangana

On-site
INR 3,500,000 - 7,000,000
Senior Agentic AI Developer
Senior Agentic AI Developer

Coretek Services India • Hyderabad

Hybrid
INR 300,000 - 540,000