AI Orchestration / Platform Engineer

NTT DATA, Inc.

Gurgaon

On-site

INR 1,800,000 - 3,200,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA is seeking an AI Orchestration / Platform Engineer to join our Hyderabad-based team, focusing on agent workflow orchestration, model routing, and runtime capabilities. You will integrate with cloud, security, policy, CI/CD, and observability, delivering reliable multi-step agent workflows in production.

The role demands strong experience in containers, Kubernetes, IaC, and AI tooling, with emphasis on state management, retries, and secure deployment.

Qualifications

  • 8+ years of platform, DevOps, SRE, cloud, distributed-systems, or software-engineering experience.
  • Production experience supporting AI/ML, LLM, agentic, workflow, or high-scale distributed application platforms.
  • Strong experience with containers, Kubernetes or equivalent runtimes, CI/CD, infrastructure as code, configuration, secrets, and automated deployment.
  • Understanding of agent orchestration concepts including state, checkpoints, retries, timeouts, queues, long-running tasks, human approvals, and failure recovery.
  • Experience with logging, metrics, distributed tracing, OpenTelemetry or equivalent observability, alerting, dashboards, and incident response.
  • Strong scripting or development skills in Python, Go, Java, TypeScript, or comparable languages.
  • Ability to integrate platform capabilities with security, identity, policy, data, network, and enterprise approval requirements.
  • Experience balancing reliability, delivery speed, latency, throughput, portability, and operating cost.

Responsibilities

  • Design and implement orchestration patterns for multi-step agents, multi-agent collaboration, deterministic workflows, long-running tasks, approvals, and event-driven execution.
  • Implement durable state, checkpoints, queues, retries, backoff, idempotency, timeouts, compensation, dead-letter handling, and recovery for agent workflows.
  • Integrate model gateways and routing logic that select models based on capability, sensitivity, latency, cost, availability, or policy requirements.
  • Build deployment and release patterns for agent services, orchestration components, prompts, tool definitions, configurations, and evaluation assets.
  • Integrate with existing CI/CD, infrastructure-as-code, policy-as-code, secrets, identity, container, and cloud-runtime capabilities.
  • Create observability for agent traces, model calls, tool calls, workflow state, token usage, cost, latency, errors, quality signals, and dependency health.
  • Implement automated quality and security gates, including unit and integration tests, evaluation suites, policy checks, vulnerability scans, and rollback criteria.
  • Optimize runtime performance, concurrency, throughput, caching, context use, model selection, and infrastructure cost.
  • Support production incidents, root-cause analysis, capacity planning, resiliency testing, disaster recovery, and operational runbooks.
  • Develop reusable platform templates, SDKs, reference pipelines, dashboards, and onboarding guidance for agent-development teams.

Skills

Python
Go
Java
TypeScript
SRE mindset
CI/CD automation

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Docker
Terraform
Pulumi
Helm
GitOps

Job description

Select how often (in days) to receive an alert: Create Alert

NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization.

We are currently seeking a AI Orchestration / Platform Engineer to join our team in Hyderabad, Telangana (IN-TG), India (IN).

Agent runtime | Model routing | LLMOps | CI/CD | Observability and reliability
Number of positions

1

Level
Primary locations

Hyderabad or Noida preferred; exceptional onshore candidates may be considered

Target / alternate titles
Core keywords

agent orchestration, workflow engine, multi-agent systems, model gateway, model routing, LLMOps, platform engineering, Kubernetes, containers, CI/CD, infrastructure as code, policy as code, observability, OpenTelemetry, queues, state, retries, resilience, cost telemetry, SRE

Recruiter red flags

Traditional DevOps profile with no AI runtime understanding; framework-only orchestration without production platform depth; manual deployments; no observability or incident ownership; cannot explain model routing, state, retries, evaluation, or AI-specific controls.

Build and operate the orchestration and runtime capabilities that allow agentic applications to move reliably from development into production. The role will integrate agent workflows with existing cloud, security, policy, CI/CD, and observability capabilities, supporting multi-step execution, model routing, state, resiliency, evaluation, and production operations without rebuilding the client's established platform foundations.

Client and delivery context
  • The client already has delivery pipelines, policy-based architecture, security controls, and production review processes. The engineer must integrate with and enhance those capabilities.
  • The platform should support multiple models, clouds, frameworks, and developer-productivity ecosystems without forcing avoidable lock-in.
  • The engineer will support shared patterns used across different business-process agents and may work across two parallel delivery tracks.
  • Platform work must be pragmatic and outcome-driven, with emphasis on enabling engineers to release and operate reliable agents quickly.
Primary ownership
  • Agent workflow and runtime orchestration, including state, routing, queues, retries, timeouts, scheduling, persistence, and exception management.
  • Model gateway and routing patterns, provider abstraction, policy-based selection, fallback, rate limits, quotas, and cost controls.
  • CI/CD, environment promotion, configuration, secrets, infrastructure integration, release validation, rollback, and operational readiness.
  • Application and platform observability, reliability engineering, incident response, capacity, performance, and production support.
Key responsibilities
  • Design and implement orchestration patterns for multi-step agents, multi-agent collaboration, deterministic workflows, long-running tasks, approvals, and event-driven execution.
  • Implement durable state, checkpoints, queues, retries, backoff, idempotency, timeouts, compensation, dead-letter handling, and recovery for agent workflows.
  • Integrate model gateways and routing logic that select models based on capability, sensitivity, latency, cost, availability, or policy requirements.
  • Build deployment and release patterns for agent services, orchestration components, prompts, tool definitions, configurations, and evaluation assets.
  • Integrate with existing CI/CD, infrastructure-as-code, policy-as-code, secrets, identity, container, and cloud-runtime capabilities.
  • Create observability for agent traces, model calls, tool calls, workflow state, token usage, cost, latency, errors, quality signals, and dependency health.
  • Implement automated quality and security gates, including unit and integration tests, evaluation suites, policy checks, vulnerability scans, and rollback criteria.
  • Optimize runtime performance, concurrency, throughput, caching, context use, model selection, and infrastructure cost.
  • Support production incidents, root-cause analysis, capacity planning, resiliency testing, disaster recovery, and operational runbooks.
  • Develop reusable platform templates, SDKs, reference pipelines, dashboards, and onboarding guidance for agent-development teams.
Must-have candidate profile
  • 8+ years of platform, DevOps, SRE, cloud, distributed-systems, or software-engineering experience.
  • Production experience supporting AI/ML, LLM, agentic, workflow, or high-scale distributed application platforms.
  • Strong experience with containers, Kubernetes or equivalent runtimes, CI/CD, infrastructure as code, configuration, secrets, and automated deployment.
  • Understanding of agent orchestration concepts including state, checkpoints, retries, timeouts, queues, long-running tasks, human approvals, and failure recovery.
  • Experience with logging, metrics, distributed tracing, OpenTelemetry or equivalent observability, alerting, dashboards, and incident response.
  • Strong scripting or development skills in Python, Go, Java, TypeScript, or comparable languages.
  • Ability to integrate platform capabilities with security, identity, policy, data, network, and enterprise approval requirements.
  • Experience balancing reliability, delivery speed, latency, throughput, portability, and operating cost.
Preferred experience
  • Experience with model gateways, multi-model routing, provider abstraction, fallback, quotas, or AI cost controls.
  • Experience with LangGraph, Temporal, Airflow, Argo Workflows, Durable Functions, Step Functions, Kubernetes operators, or equivalent orchestration technologies.
  • Experience with AI tracing and evaluation platforms, prompt/model registries, feature flags, canary releases, and regression gates.
  • Experience with Terraform, Pulumi, Helm, GitOps, policy engines, service mesh, event streaming, and API gateways.
  • Experience operating platforms across AWS, Azure, GCP, hybrid, or private environments.
  • Experience in regulated enterprises or systems with sensitive proprietary data and formal production controls.

Kubernetes, Docker, serverless or managed container platforms; Terraform/Pulumi/Helm/GitOps; GitHub Actions, GitLab, Jenkins, Azure DevOps, or equivalent; LangGraph, Temporal, Airflow, Argo, Step Functions, Durable Functions, or equivalent orchestration; model gateways and provider APIs; queues and events; OpenTelemetry, Prometheus, Grafana, cloud monitoring, AI tracing and evaluation tools. Exact products are flexible.

Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees.

NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us . This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here . If you'd like more information on your EEO rights under the law, please click here . For Pay Transparency information, please click here .

Job Segment

Cloud, Testing, Developer, Java, Technology

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Orchestration / Platform Engineer
AI Orchestration / Platform Engineer

NTT DATA North America • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Principal Agentic AI Engineer / Hands-on Technical Lead
Principal Agentic AI Engineer / Hands-on Technical Lead

NTT DATA, Inc. • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Principal Agentic AI Engineer / Hands-on Technical Lead
Principal Agentic AI Engineer / Hands-on Technical Lead

NTT DATA, Inc. • Gurgaon

On-site
INR 3,500,000 - 5,500,000
Senior Agentic Full-Stack Engineers
Senior Agentic Full-Stack Engineers

NTT DATA, Inc. • Gurgaon

Hybrid
INR 1,800,000 - 2,400,000
Principal Agentic AI Engineer / Hands-on Technical Lead
Principal Agentic AI Engineer / Hands-on Technical Lead

NTT DATA North America • Gurugram District

Hybrid
INR 4,000,000 - 9,000,000
Senior Agentic Full-Stack Engineers
Senior Agentic Full-Stack Engineers

NTT DATA North America • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior Agentic Full-Stack Engineers
Senior Agentic Full-Stack Engineers

NTT DATA North America • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Required. AI & Data Analytics Engineer AI & Data Analytics Engineer
Required. AI & Data Analytics Engineer AI & Data Analytics Engineer

NTT DATA, Inc. • Bengaluru

Hybrid
INR 1,800,000 - 3,000,000
AI Engineer
AI Engineer

NTT DATA, Inc. • Gurugram District

On-site
INR 300,000 - 600,000
GenAI/ML Engineers
GenAI/ML Engineers

NTT DATA, Inc. • Hyderabad

On-site
INR 3,500,000 - 7,000,000