Ai/ml/llm Systems Engineer - Enterprise Ai Platform Engineer

Saudi Aramco (ASC)

Saudi Arabia

On-site

SAR 299,962 - 449,943

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Saudi Aramco (ASC) is seeking an AI ML LLM Systems Engineer to contribute to the development of enterprise-scale AI platforms. You will deploy and manage LLMs and vision models on NVIDIA SuperPods Cloud and build scalable inference pipelines using Kubernetes and Docker.

The ideal candidate has a Master's degree, 8 years in AI/ML systems, and strong proficiency in Python and SQL. The role focuses on ensuring high application performance and compliance with SLAs.

Qualifications

  • Hold a Master's degree in Computer Science, Software Engineering, or a related field.
  • Have 8 years of experience in AI/ML systems or cloud-native infrastructure, including at least 4 years in LLM deployment and optimization.
  • Experience deploying and optimizing LLMs and vision models on NVIDIA GPU clusters and HPC environments.

Responsibilities

  • Deploy and manage LLMs and vision models on NVIDIA SuperPods Cloud.
  • Build and maintain scalable inference pipelines using Kubernetes, Docker, and OpenShift.
  • Optimize performance through multiple techniques.

Skills

Proficiency in Python
Proficiency in SQL
Kubernetes (K8s)
Docker
OpenShift
Elasticsearch
PostgreSQL
Monitoring and dashboarding techniques

Education

Master's degree in Computer Science, Software Engineering, or related field

Tools

NVIDIA SuperPods Cloud
Git
Bitbucket
Jenkins
ArgoCD

Job description

We are seeking an AI ML LLM Systems Engineer to join our Digital & AI Center of Excellence and contribute to the development of enterprise-scale AI platforms that support advanced machine learning and language model inference across Saudi Aramco's operations.

Responsibilities
  • Deploy and manage LLMs and vision models on NVIDIA SuperPods Cloud ensuring high performance and efficient use of GPU resources.
  • Build and maintain scalable inference pipelines using Kubernetes (K8s), Docker and OpenShift for enterprise AI platforms.
  • Optimize inference performance through multiple techniques.
  • Benchmark and evaluate LLMs for performance, accuracy, latency and resource utilization across different hardware and software configurations.
  • Implement and support LLMOps frameworks with full observability including logging, tracing and model performance tracking.
  • Integrate and manage vector databases Elasticsearch and relational databases PostgreSQL for efficient data retrieval and user interaction history tracking.
  • Implement and maintain CI/CD (Continuous Integration and Continuous Delivery) pipelines for model and platform updates using Git, Bitbucket, Jenkins and ArgoCD.
  • Ensure high availability and reliability of AI application workflows using frameworks like Haystack.
  • Collaborate with infrastructure teams on GPU provisioning and resource allocation for AI workloads.
  • Develop and maintain monitoring, alerting and dashboarding systems for AI/ML workloads to ensure SLA/SLO compliance.
Qualifications
  • Hold a Master's degree in Computer Science, Software Engineering, or a related field.
  • Have 8 years of experience in AI/ML systems or cloud-native infrastructure, including at least 4 years in LLM deployment and optimization.
  • Proficiency in Python and SQL is required, with experience in building and optimizing AI/ML applications.
  • Ability to work with Kubernetes (K8s), Docker, and OpenShift in production environments.
  • Experience deploying and optimizing LLMs and vision models on NVIDIA GPU clusters and high-performance computing (HPC) environments and cloud environment.
  • Ability to demonstrate proficiency in inference scaling, distributed computing, and SLA/SLO planning for AI workloads.
  • Strong knowledge in Elasticsearch, PostgreSQL, and workflow frameworks like Haystack for AI application development.
  • Ability to implement CI/CD pipelines using tools like Git, Bitbucket, Jenkins, and ArgoCD.
  • Experience in benchmarking and evaluating LLMs for performance, accuracy, and efficiency is required.
  • Monitoring and dashboarding for AI/ML systems is also necessary.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Enterprise AI Platform Engineer | LLM & GPU Inference
Enterprise AI Platform Engineer | LLM & GPU Inference

Saudi Aramco (ASC) • Saudi Arabia

On-site
SAR 299,000 - 450,000
AI Engineer
AI Engineer

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
AI Engineer
AI Engineer

Latitude • Riyadh

On-site
SAR 420,000 - 660,000
Senior AI Platform Engineer — MLOps, LLMs & RAG Systems
Senior AI Platform Engineer — MLOps, LLMs & RAG Systems

Latitude • Riyadh

On-site
SAR 420,000 - 660,000
Senior AI Systems Engineer & Platform Lead
Senior AI Systems Engineer & Platform Lead

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
AI/ML Automation Analyst
AI/ML Automation Analyst

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 180,000 - 240,000
Lead Engineer, AI & Machine Learning II
Lead Engineer, AI & Machine Learning II

International Women in Mining • Saudi Arabia

On-site
SAR 300,000 - 450,000
AI-ML Support Analyst
AI-ML Support Analyst

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 180,000 - 300,000
Senior AI Engineer
Senior AI Engineer

Mozn • Saudi Arabia

On-site
SAR 260,000 - 420,000
Competitive compensation
Top-tier health insurance
Enabling culture to focus on strengths
+1
Senior AI Technical Lead for Generative & Enterprise ML
Senior AI Technical Lead for Generative & Enterprise ML

Systems Arabia • Riyadh

On-site
SAR 350,000 - 600,000