AI Platform Engineer

Sharon AI, Inc

Sydney

On-site

AUD 120,000 - 180,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Sharon AI is building the GPU‑as‑a‑Service platform in a cutting‑edge AI infrastructure environment. The AI Platform Engineer L2 will help build, operate and support the platform layer powering Sharon AI's GPUaaS offering across Kubernetes, Slurm, and MLOps tooling.

Reporting to the Head of Operations, you will collaborate with Network Engineering and Infrastructure teams to automate provisioning, troubleshoot platform services, and enable customers to train and run AI/ML workloads securely and

Qualifications

  • 2–4 years of experience in platform engineering, DevOps, MLOps or SRE, ideally supporting GPU or AI/ML workloads.
  • Bachelor's degree in Computer Science or related field.
  • Hands‑on production experience with Kubernetes.
  • Experience with CI/CD and Infrastructure‑as‑Code.
  • Solid understanding of Kubernetes and GPU scheduling frameworks, including Slurm, Kubernetes device plugins and NVIDIA GPU Operator.
  • Experience with MLOps tooling and ML pipeline orchestration.
  • Proficiency in scripting and automation using Python and Bash.
  • Working knowledge of GPU infrastructure and distributed training concepts, including NCCL and data/model parallelism.
  • Experience with observability tools such as Prometheus and Grafana.
  • Understanding of Linux systems administration and networking fundamentals.
  • Strong troubleshooting and problem‑solving skills.
  • Strong communication and collaboration skills, with the ability to work across Infrastructure, Network Engineering and customer‑facing teams.
  • Exposure to GPU‑based infrastructure or high‑performance computing environments.

Responsibilities

  • Build and operate the AI platform layer, including Kubernetes, Slurm and container orchestration, supporting Sharon AI's GPU infrastructure
  • Develop and maintain CI/CD pipelines for model training, fine‑tuning and inference workloads
  • Implement and support MLOps tooling for experiment tracking, model registry and deployment
  • Configure and manage multi‑tenant GPU resource scheduling and quota management across customer workloads
  • Support model serving infrastructure for training and inference, ensuring reliability and performance
  • Build monitoring, logging and alerting to track platform health, GPU utilisation and workload performance
  • Automate platform provisioning and configuration using Infrastructure‑as‑Code tools such as Terraform and Ansible
  • Collaborate with Network Engineering and Infrastructure teams to ensure the platform layer aligns with the underlying InfiniBand/RDMA fabric
  • Troubleshoot platform‑level issues affecting customer AI/ML workloads
  • Participate in on‑call rotations and incident response for platform‑related issues
  • Contribute to platform documentation, runbooks and the internal knowledge base
  • Support customer onboarding onto the GPUaaS platform, including workload configuration and troubleshooting

Skills

Kubernetes
CI/CD
Python
Bash
Linux
Troubleshooting
Communication
Multi-tenant GPU scheduling
MLOps
NVIDIA GPU Operator
NCCL
PyTorch
TensorFlow
Slurm
Terraform
Ansible

Education

Bachelor's degree in Computer Science

Tools

Kubernetes
Slurm
NVIDIA GPU Operator
CI/CD pipelines
Terraform
Ansible
MLflow
Kubeflow
Ray

Job description

About Sharon AI

Sharon AI is building the infrastructure powering the next generation of artificial intelligence.

Operating across AI infrastructure, high-performance compute, cloud platforms and large-scale technology environments, Sharon AI delivers scalable, secure and reliable infrastructure for demanding AI, ML and HPC workloads.

The Role

As an AI Platform Engineer L2, you'll help build, operate and support the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering. You'll work across Kubernetes, Slurm, container orchestration, MLOps tooling and model serving infrastructure to enable customers to train and run AI/ML workloads reliably and efficiently on Sharon AI's neocloud platform.

Reporting to the Head of Operations, you'll work closely with Network Engineering, Infrastructure and customer-facing teams to implement, automate and troubleshoot the platform services sitting above Sharon AI's underlying GPU and network fabric. This is a hands‑on opportunity for a platform, DevOps, MLOps or SRE engineer looking to deepen their expertise in GPU infrastructure and AI‑native platform operations.

Key Responsibilities

  • Build and operate the AI platform layer, including Kubernetes, Slurm and container orchestration, supporting Sharon AI's GPU infrastructure
  • Develop and maintain CI/CD pipelines for model training, fine‑tuning and inference workloads
  • Implement and support MLOps tooling for experiment tracking, model registry and deployment
  • Configure and manage multi‑tenant GPU resource scheduling and quota management across customer workloads
  • Support model serving infrastructure for training and inference, ensuring reliability and performance
  • Build monitoring, logging and alerting to track platform health, GPU utilisation and workload performance
  • Automate platform provisioning and configuration using Infrastructure‑as‑Code tools such as Terraform and Ansible
  • Collaborate with Network Engineering and Infrastructure teams to ensure the platform layer aligns with the underlying InfiniBand/RDMA fabric
  • Troubleshoot platform‑level issues affecting customer AI/ML workloads
  • Participate in on‑call rotations and incident response for platform‑related issues
  • Contribute to platform documentation, runbooks and the internal knowledge base
  • Support customer onboarding onto the GPUaaS platform, including workload configuration and troubleshooting

Skills & Experience

  • 2–4 years' experience in platform engineering, DevOps, MLOps or SRE, ideally supporting GPU or AI/ML workloads
  • Bachelor's degree in Computer Science or a related field
  • Hands‑on production experience with Kubernetes
  • Experience with CI/CD and Infrastructure‑as‑Code
  • Solid understanding of Kubernetes and GPU scheduling frameworks, including Slurm, Kubernetes device plugins and NVIDIA GPU Operator
  • Experience with MLOps tooling and ML pipeline orchestration
  • Proficiency in scripting and automation using Python and Bash
  • Working knowledge of GPU infrastructure and distributed training concepts, including NCCL and data/model parallelism
  • Experience with observability tools such as Prometheus and Grafana
  • Understanding of Linux systems administration and networking fundamentals
  • Strong troubleshooting and problem‑solving skills
  • Strong communication and collaboration skills, with the ability to work across Infrastructure, Network Engineering and customer‑facing teams
  • Exposure to GPU‑based infrastructure or high‑performance computing environments

Experience with Slurm, NVIDIA GPU Operator, NCCL and distributed training frameworks such as PyTorch or TensorFlow is advantageous, as is experience with MLOps platforms including MLflow, Kubeflow or Ray. Kubernetes certifications such as CKA/CKAD, cloud certifications across AWS/GCP/Azure, and exposure to InfiniBand/RDMA networking concepts are also advantageous.

Why Join Sharon AI

  • Help build and operate the platform powering a growing GPU‑as‑a‑Service and AI neocloud business
  • Work hands‑on with GPU infrastructure, Kubernetes, Slurm and AI‑native platform technologies
  • Develop deeper expertise across AI/HPC infrastructure and distributed workloads
  • Work closely with Network Engineering and Infrastructure teams across the underlying GPU and network fabric
  • Help enable customers to reliably train, fine‑tune and run AI/ML workloads at scale
  • Contribute to automation, observability and platform capabilities in a fast‑moving environment
  • Join a highly technical and ambitious team operating at the forefront of AI infrastructure

Our Values

Integrity | Innovation | Collaboration | Wellbeing | Inclusion

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Engineer — GPU Infrastructure & MLOps
AI Platform Engineer — GPU Infrastructure & MLOps

Sharon AI, Inc • Sydney

On-site
AUD 120,000 - 180,000
Systems Administrator
Systems Administrator

Sharon AI, Inc • Council of the City of Sydney

On-site
AUD 90,000 - 130,000
Platforms Development Lead — Remote/Hybrid SaaS AI Infra
Platforms Development Lead — Remote/Hybrid SaaS AI Infra

Sharon AI, Inc • Sydney

Hybrid
AUD 140,000 - 210,000
Platforms Development Team Lead
Platforms Development Team Lead

Sharon AI, Inc • Sydney

Hybrid
AUD 140,000 - 210,000
Principal Network Architect
Principal Network Architect

Sharon AI, Inc • Sydney

Hybrid
AUD 180,000 - 260,000
Platforms Development Team Lead
Platforms Development Team Lead

Sharon Ai, Inc • Sydney

Hybrid
AUD 180,000 - 240,000
Product Marketing Manager
Product Marketing Manager

Sharon AI, Inc • Sydney

On-site
AUD 120,000 - 160,000
Technical Project Manager
Technical Project Manager

Sharon AI, Inc • Sydney

On-site
AUD 110,000 - 150,000
AI Infrastructure Systems Engineer
AI Infrastructure Systems Engineer

Sharon AI, Inc • Council of the City of Sydney

On-site
AUD 90,000 - 130,000
Platform DevOps Engineer
Platform DevOps Engineer

Future Secure AI • Sydney

Hybrid
AUD 120,000 - 150,000
Competitive salary
Flexible work environment
Diversity and creativity