Senior Generative AI Operations (GenAI Ops) Engineer

EPAM Systems

Turkey

On-site

TRY 2,917,000 - 4,375,000

Full time

23 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Private health insurance
English courses
Continuous upskilling

Job summary

EPAM Systems is seeking a GenAI Ops Engineer to build, deploy and sustain the operational infra for large-scale generative AI models across major clouds. You will work with data scientists, ML engineers and software developers to ensure scalable, reliable GenAI services and agent ecosystems.

The role requires strong experience in CI/CD, IaC, container orchestration, and monitoring, with a focus on multi-agent workflows and tool integrations. Fluent English is required.

Qualifications

  • 3+ years in a DevOps/SRE/MLOps role focused on cloud infra.
  • Experience deploying and operating LLM inference services.
  • Familiarity with IaC, CI/CD, and container orchestration.
  • Hands-on with tracing/metrics and eval pipelines.
  • Experience with multi-agent workflows and tool integration.
  • Knowledge of guardrails, data security and compliance.
  • Fluent English (B2+).

Responsibilities

  • Build and manage CI/CD pipelines for LLMs and AI agents.
  • Orchestrate multi-agent workflows and A2A communications.
  • Integrate tools via MCP and secure external APIs.
  • Apply IaC across cloud platforms to support GenAI workloads.
  • Monitor, log and optimize model serving and inference.
  • Ensure security, governance and compliance in GenAI infra.

Skills

Cloud platforms (AWS,GCP,Azure)
CI/CD pipelines
Scripting (Python,Bash)
IaC (Terraform,AWS CDK,CloudFormation)
Docker
Kubernetes
LLM inference & tooling (vLLM, Triton,
Observability & tracing (OpenTelemetry
Retrieval pipelines & vector stores (P
Multi-agent workflows (LangGraph, Crew
Security & guardrails

Education

Master's degree or PhD in CS/AI

Tools

Terraform
Docker
Kubernetes
AWS
GCP
Azure
Model Context Protocol (MCP)

Job description

We are seeking a highly motivated and experienced Generative AI Operations (GenAI Ops) Engineer to join our innovative team.

In this role, you will be at the forefront of the AI revolution, responsible for building, deploying, and maintaining the operational infrastructure for our cutting-edge generative AI models and services. You will work closely with data scientists, machine learning engineers and software developers to ensure our GenAI applications—especially complex multi-agent systems—are scalable, reliable and efficient across major cloud platforms. If you are passionate about operationalizing large-scale AI systems and want to make a significant impact, this is the role for you.

Responsibilities
  • Build and Manage CI/CD Pipelines: Design, implement and maintain robust, automated CI/CD pipelines for training, evaluating and deploying large language models (LLMs) and AI agents
  • Orchestrate Agentic AI Workflows: Design, deploy and manage sophisticated multi-agent systems Ensure seamless Agent-to-Agent (A2A) communication and collaboration between specialized agents to automate complex business processes
  • Manage Tool Integration: Implement and manage secure, scalable integrations between AI agents and external tools/APIs, leveraging open standards like the Model Context Protocol (MCP) to ensure interoperability
  • Leverage AI-Powered Development: Utilize AI-powered development tools to accelerate the entire software development lifecycle from writing infrastructure code and tests to troubleshooting operational issues in cloud environments
  • Infrastructure as Code (IaC): Utilize cloud-native IaC services or cloud-agnostic tools like Terraform to define and manage the infrastructure required for GenAI workloads
  • Model Monitoring and Observability: Implement comprehensive monitoring and logging solutions to track model and agent performance, resource utilization and system health For agentic systems, this includes tracing the agent's actions and logging the multi-step conversational flow
  • Scalability and Performance Optimization: Design and implement scalable architectures for model serving and inference Continuously optimize the performance and cost-effectiveness of our GenAI services
  • Security and Compliance: Implement and enforce security best practices for our GenAI infrastructure and data Ensure compliance with industry standards and regulations
Requirements
  • 3+ years in a DevOps, SRE or MLOps role with a focus on cloud infrastructure and a background in cloud services (AWS, GCP, Azure)
  • Skilled in building and managing CI/CD pipelines (Jenkins, GitLab CI or cloud-native services) and proficiency in at least one scripting language (e.g. Python, Bash)
  • Familiarity with IaC tools (e.g. AWS CDK, CloudFormation, Terraform) and in containerization and orchestration (Docker, Kubernetes)
  • Track record deploying and operating LLM inference (e.g. vLLM, Triton, TGI, Ray Serve, KServe/Seldon)
  • Hands‑on with LLM/app tracing and metrics (e.g. OpenTelemetry + Langfuse, Arize Phoenix, WhyLabs) and building eval pipelines (offline/online regression suites)
  • Skilled in operating retrieval pipelines: embedding generation, indexing/refresh strategies, vector DBs (Pinecone, Weaviate, Milvus, FAISS) and relevance monitoring
  • Practice running multi‑agent workflows (LangGraph, CrewAI, AutoGen‑like), including state management, retries, rate limits, tool‑failure handling and step‑level auditing
  • Experience in implementing guardrails: secrets isolation, tool/API permissions, prompt‑injection defenses, data leakage prevention, PII redaction and policy enforcement
  • Fluent English (B2+ level)
Nice to have
  • Master's degree or PhD in Computer Science, AI, Machine Learning or a related field
  • Background integrating agents with external tools using MCP (or similar tool‑calling standards) and operating tool registries
  • Experience with cloud‑native GenAI services like AWS Bedrock, Azure AI Foundry or Google Vertex AI
  • Familiarity with the architecture and operational challenges of Large Language Models (LLMs)
  • Experience designing or managing multi‑agent systems or complex orchestrated workflows
  • Knowledge of monitoring and observability tools like Prometheus, Grafana or Datadog
  • Relevant cloud or DevOps certifications
  • Strong problem‑solving skills and the ability to work effectively in a fast‑paced collaborative environment
We offer
  • CONTINUOUS UPSKILLING, LEARNING & DEVELOPMENT
    • Diversity of tasks and projects
    • Assessment center for objective review of competency level
    • Personal development plan
    • Mentoring programs and leadership development
    • Certification and professional development support
    • Access to learning platforms including more than 2,500 internal courses
    • English courses taught by certified teachers
  • CORPORATE BENEFITS
    • Extra leave days
    • Referral bonuses
  • COMPENSATION PACKAGE
    • Competitive compensation paid in USD
    • Regular salary and performance reviews
  • MEDICAL & HEALTHCARE
    • Private health insurance
    • Well‑being events
  • WORKING ENVIRONMENT
    • Recreation areas and kitchens
    • Tea, coffee and snacks
    • Sports equipment and game consoles
    • IT Equipment
    • Microsoft’s Software Assurance Home Use Program (HUP)

Please note that our Talent Attraction Team reviews applications and CVs submitted in English.

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI‑Native enterprises, driving measurable value from innovation and digital investments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Generative AI Operations Engineer (GenAI Ops)
Generative AI Operations Engineer (GenAI Ops)

EPAM Systems • Turkey

On-site
TRY 300,000 - 460,000
Private health insurance
English courses
Professional development support
+2
Forward Deployed Engineer/Chief Role
Forward Deployed Engineer/Chief Role

EPAM Systems • Turkey

On-site
TRY 4,375,000 - 5,834,000
Private health insurance
English courses and learning stipend
Mentoring programs
Senior AI Engineer with ReactJS
Senior AI Engineer with ReactJS

EPAM Systems • Turkey

On-site
TRY 4,375,000 - 6,806,000
Continuous upskilling
Mentoring programs and leadership dev
Private health insurance
+4
Senior Data Engineer, AI
Senior Data Engineer, AI

EPAM Systems • Turkey

On-site
TRY 4,375,000 - 6,806,000
Continuous upskilling
Diversity of tasks and projects
Mentoring programs
+5
Lead Azure Cloud Engineer
Lead Azure Cloud Engineer

EPAM Systems • Turkey

On-site
TRY 400,000 - 800,000
Upskilling program
Private health insurance
Mentoring programs
+1
Lead DevOps Engineer (Azure AI / GenAI Solutions)
Lead DevOps Engineer (Azure AI / GenAI Solutions)

EPAM Systems • Turkey

On-site
TRY 4,375,000 - 6,320,000
Continuous upskilling
Private health insurance
USD compensation
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • Turkey

On-site
TRY 4,859,000 - 6,803,000
Private health insurance
Professional development
English courses
+2
Senior Python GenAI Engineer
Senior Python GenAI Engineer

EPAM Systems • Turkey

On-site
TRY 4,375,000 - 5,834,000
Private health insurance
Continuous upskilling
English courses
+2
Forward Deployed Engineer/Chief Role
Forward Deployed Engineer/Chief Role

EPAM Systems, Inc. • Turkey

On-site
TRY 5,772,000 - 8,658,000
Extra leave days
Referral bonuses
Private health insurance
+1
Senior AI Engineer with AWS Bedrock
Senior AI Engineer with AWS Bedrock

EPAM Systems • Turkey

Hybrid
TRY 3,885,000 - 6,799,000
Continuous upskilling
Diversity of tasks
Mentoring programs
+5