Position Summary:
Systems Ltd is seeking an experienced, highly technical, and hands-on Agentic AI Engineer to join our cutting-edge technology team onsite in Dubai, UAE. In this specialized artificial intelligence and systems reliability role, you will spearhead the deployment, optimization, and scaling of advanced LLM inference serving, high-throughput agentic workflows, and distributed GPU infrastructure. You will work closely with cross-functional AI engineering, DevOps, and site reliability teams to guarantee ultra-low latency, optimal token throughput, and absolute production stability across mission-critical generative AI deployments. Ideal candidates bring a rigorous engineering background, deep operational mastery of vLLM, NVIDIA GPU tuning, Kubernetes orchestration, and complex distributed tracing frameworks.
Detailed Job Description:
As an Agentic AI Engineer at Systems Ltd in Dubai, you will take full ownership of the end-to-end performance, scalability, and resilience of our enterprise LLM serving engines and agentic pipelines. Your day-to-day responsibilities encompass managing vLLM inference instances, debugging GPT-OSS and Harmony parsers, optimizing KV-cache and prompt caching, and tuning Tensor Parallelism (TP2/TP4) across NVIDIA GPU clusters. You will dive deep into concurrency control, dynamic batching, and throughput optimization while managing complex infrastructure components including Docker, Kubernetes, API gateways, reverse proxies, and AWS ALB/NLB load balancers. You will also lead comprehensive load, stress, and soak testing, conduct granular latency analyses (TTFT, tokens/sec, P50/P95/P99), manage high-stakes production incidents, write Root Cause Analyses (RCAs), and ensure robust observability via OpenTelemetry, Langfuse, and Kibana.
Key Responsibilities:
- Deploy, configure, and manage vLLM and high-performance LLM inference serving engines in production environments.
- Troubleshoot and optimize GPT-OSS models and Harmony parsers for complex agentic tool and function-calling workflows.
- Implement prompt caching, KV-cache optimization, and advanced Tensor Parallelism (TP2/TP4) tuning strategies.
- Monitor NVIDIA GPU performance, memory allocation, and utilization metrics to maximize compute efficiency.
- Tune concurrency, dynamic request batching, and token generation throughput across distributed LLM clusters.
- Operate containerized AI applications and microservices using Docker and Kubernetes in enterprise production settings.
- Troubleshoot complex networking issues across API gateways, reverse proxies, HTTP protocols, TLS/certificates, and connection resets.
- Administer and debug load balancers (AWS ALB/NLB) and implement robust retry, timeout, backoff, and circuit-breaker patterns.
- Establish comprehensive observability using OpenTelemetry, distributed tracing, Langfuse LLM monitoring, and Kibana centralized logging.
- Conduct rigorous load, stress, and soak testing alongside granular latency analysis (P50/P95/P99, TTFT, tokens/sec) and capacity planning.
Required Qualifications & Skills:
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Software Engineering, or a related technical discipline.
- Extensive hands-on professional experience in AI infrastructure engineering, LLM inference serving, or Site Reliability Engineering for AI systems.
- Deep technical expertise in vLLM, LLM inference tuning, KV-cache optimization, and prompt caching mechanisms.
- Proven proficiency in NVIDIA GPU monitoring, performance tuning, and Tensor Parallelism configuration (TP2/TP4).
- Strong practical background in Python-based AI application debugging, concurrency management, and throughput optimization.
- Extensive production operations experience with Docker, Kubernetes, API gateways, and reverse proxy troubleshooting.
- Solid command of networking fundamentals including HTTP, TLS/SSL certificates, connection resets, and ALB/NLB load balancer administration.
- Practical experience implementing distributed tracing, OpenTelemetry, Langfuse observability, and Kibana centralized logging.
- Demonstrated ability in conducting load testing, capacity planning, saturation analysis, and high-stakes production incident management (RCA).
- Professional availability to work onsite in Dubai, UAE, with excellent communication and cross-functional leadership skills.
Nice-to-Have Skills:
- Contributions to open-source LLM serving projects or advanced AI agent orchestration frameworks (LangChain, AutoGen, CrewAI).
- Certifications such as Certified Kubernetes Administrator (CKA), AWS Certified Solutions Architect, or NVIDIA Deep Learning Institute credentials.
- Experience with custom CUDA kernel optimization or quantized model formats (AWQ, GPTQ, GGUF).
- Prior working experience in regional Middle East enterprise or telecommunications tech hubs.
- Experience automating CI/CD pipelines, configuration versioning, and zero-downtime rolling rollbacks for AI services.
Application Information:
- Recruiter: Systems Ltd Recruitment Team
- Contact Name: Heena Shaikh (CHRAM)
- Email: heena.shaikh@systemsltd.com
- Phone: Unspecified
- Application URL: Unspecified
- Salary/Rate: Competitive and commensurate with experience
- Deadline: Open until filled
- Notice Period: Immediate to short notice preferred
- Contract Duration: Permanent / Full-Time (Onsite)
Recruitment Pro Tip:
When applying for senior agentic AI and LLM infrastructure roles, ensure your resume highlights specific metrics you have improved—such as lowering Time-to-First-Token (TTFT), increasing tokens-per-second throughput, or successfully optimizing GPU memory utilization under heavy concurrent loads.