AI Architecture & Infrastructure Engineer

CMB Wing Lung Bank Limited

Hong Kong

On-site

HKD 900,000 - 1,300,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

CMB Wing Lung Bank Limited is seeking an AI Architecture & Infrastructure Engineer to design and build bank-grade AI infrastructure, including data processing, training, deployment, and lifecycle management. You will focus on security hardening, reliability, and scalable AI platform engineering within a banking context, reporting to the Head of AI Platform.

The role requires strong Kubernetes experience, cloud-native skills, and hands-on model serving, with responsibilities spanning data, model

Qualifications

  • Bachelor’s degree in a related field or above.
  • 5+ years in AI platforms, cloud, or distributed infra.
  • Strong programming in Python/Go/Java/C++ and system design.

Responsibilities

  • Design and deliver AI infra architecture and platform.
  • Build scalable model-training and data processing infra.
  • Develop large-model training and fine-tuning pipelines.
  • Operate Kubernetes-based AI workloads and GPU management.
  • Create unified model serving and inference optimization.
  • Implement MLOps/LLMOps pipelines and governance.
  • Lead security hardening, reliability, and resilience.
  • Build observability and production operations for AI workloads.
  • Provide technical leadership and cross-functional delivery.

Skills

Python
Go
Java
C++
Data structures
Distributed systems
Kubernetes
Linux
Cloud-native

Education

Bachelor’s degree

Tools

Kubernetes
Docker
Terraform
Helm
Argo CD
Prometheus
Grafana
OpenTelemetry
Kafka
Redis
PostgreSQL

Job description

Position Overview

We are hiring an AI Architecture & Infrastructure Engineer to design and build bank-grade AI infrastructure, model platforms, and intelligent-application foundations. The role is responsible for turning AI capabilities into secure, reliable, scalable, auditable, and operationally sustainable enterprise platforms, covering data processing, model training, fine-tuning, evaluation, deployment, online inference, monitoring, and lifecycle management.

The position focuses on AI architecture, distributed computing, GPU resource management, Kubernetes, MLOps/LLMOps, model serving, cloud-native engineering, security hardening, reliability, and business continuity. In this job description, “hardness engineering” is interpreted as security hardening, reliability engineering, and resilience engineering for AI infrastructure.

Item

Details

Job title

AI Architecture & Infrastructure Engineer

Function

AI Platform, Technology Architecture, Cloud Infrastructure, and Machine Learning Engineering

Reporting to

Head of AI Platform / Head of Technology Architecture

Location

[City / Working arrangement]

Employment type

Full-time

Number of openings

[Number]

Key Responsibilities

1. AI Architecture and Platform Engineering

Design and deliver the architecture of the bank’s AI infrastructure and platform, covering data, training, fine-tuning, evaluation, model registry, model serving, inference gateway, identity, auditability, observability, and operations. Develop architectures for offline training, batch prediction, real-time inference, high-concurrency serving, and multi-model collaboration. Establish architecture standards, interface specifications, technical baselines, and platform evolution roadmaps.

2. Model Training and Data Processing Infrastructure

Build repeatable, scalable, and auditable model-training infrastructure for supervised learning, deep learning, large language models, and other machine-learning workloads. Own training orchestration, dataset management, data preparation, feature processing, experiment tracking, model evaluation, resource scheduling, training logs, metric collection, and model-artifact management. Optimize data, compute, network, storage, and communication efficiency for training workloads.

3. Model Fine-Tuning and Large-Model Engineering

Design and implement infrastructure for fine-tuning large language models and domain-specific models, supporting instruction tuning, supervised fine-tuning, parameter-efficient fine-tuning, and relevant preference-optimization workflows. Understand and apply techniques such as LoRA, QLoRA, adapters, quantization, distillation, mixed precision, gradient accumulation, checkpoint management, and distributed training. Integrate these capabilities into standardized training, evaluation, release, and rollback processes. Work with data, model-risk, business, and security teams to ensure that training data and model artifacts meet internal governance requirements.

4. Kubernetes and Cloud-Native Platform Engineering

Build and operate AI workloads on Kubernetes (K8s), including containerization, Pod scheduling, GPU allocation, node-pool management, autoscaling, job queues, network policies, service discovery, storage orchestration, namespace isolation, quota management, and multi-tenant governance. Work with Helm, Operators, Ingress, service mesh, CI/CD, GitOps, Terraform, or equivalent technologies. Troubleshoot reliability, performance, scheduling, and resource-isolation issues for training and inference workloads running on Kubernetes.

5. Model Serving and Inference Optimization

Build a unified model-serving platform and model gateway for conventional machine-learning models, deep-learning models, and large language models. Own model containerization, version management, canary releases, A/B testing, rate limiting, circuit breaking, caching, batch inference, asynchronous invocation, and multi-model routing. Continuously optimize inference latency, throughput, GPU/CPU utilization, memory consumption, concurrency, and cost per request. Experience with model quantization, continuous batching, KV cache, inference caching, or other acceleration techniques is preferred.

6. MLOps/LLMOps and Delivery Automation

Build automated workflows covering data preparation, training, fine-tuning, evaluation, model registration, approval, release, and production monitoring. Maintain traceability among model versions, dataset versions, code versions, configuration parameters, and experiment results to ensure reproducibility, rollback, and auditability. Provide standardized SDKs, APIs, templates, pipelines, and self-service tools for data-science, machine-learning, and application teams.

7. Security Hardening, Reliability, and Resilience Engineering

Establish security baselines and production-reliability mechanisms for the AI platform, including authentication, fine-grained authorization, secrets and key management, network segmentation, data masking, sensitive-data protection, image and dependency scanning, software supply-chain security, runtime protection, vulnerability remediation, audit trails, and anomaly detection. Define service-level objectives, capacity-management practices, incident-response procedures, backup and recovery plans, business-continuity controls, disaster-recovery failover, and regular resilience testing. Improve recoverability under traffic surges, dependency failures, hardware failures, model anomalies, and security incidents.

8. Observability and Production Operations

Build unified observability across infrastructure, training jobs, model services, data pipelines, and business indicators. Use logs, metrics, distributed tracing, model-quality indicators, data-drift detection, performance metrics, and cost metrics to diagnose platform and model issues. Participate in production on-call rotations, major-incident reviews, capacity planning, performance testing, and platform upgrades, driving root-cause remediation rather than temporary fixes.

9. Technical Leadership and Cross-Functional Delivery

Own technical design reviews, proof-of-concept validation, solution decomposition, production acceptance, and documentation for key initiatives. Mentor software, machine-learning, platform, SRE, data, security, and risk engineers. Participate in technical hiring, interviews, code reviews, and architecture reviews, helping the organization build standardized, reusable, and continuously evolving AI engineering capabilities.

Minimum Qualifications

Capability area

Requirements

Education

Bachelor’s degree or above in Computer Science, Software Engineering, Artificial Intelligence, Network Engineering, Information Security, or a related field.

Experience

5+ years of experience in AI platforms, cloud platforms, distributed systems, MLOps/LLMOps, platform engineering, or related infrastructure development. Experience in banking, financial services, or other highly regulated industries is preferred.

Programming

Strong proficiency in at least one of Python, Go, Java, or C++, with solid knowledge of data structures, algorithms, concurrency, networking, and system design.

Linux and cloud-native engineering

Hands-on experience with Linux, Docker, Kubernetes/K8s, container networking, service discovery, CI/CD, GitOps, and infrastructure as code.

Kubernetes expertise

Understanding of K8s scheduling, Deployment, StatefulSet, Job, CronJob, Service, Ingress, ConfigMap, Secret, RBAC, NetworkPolicy, resource quotas, and autoscaling. Experience with GPU workloads and multi-tenant clusters is preferred.

Model training

Understanding of data preparation, training orchestration, distributed training, mixed precision, checkpointing, experiment tracking, model evaluation, and model registration.

Model fine-tuning

Experience with supervised fine-tuning, instruction tuning, LoRA/QLoRA, adapters, quantization, distillation, parameter-efficient fine-tuning, or related large-model training techniques.

AI frameworks

Experience with at least one of PyTorch, TensorFlow, JAX, Hugging Face Transformers, DeepSpeed, FSDP, or equivalent frameworks and tools.

Model serving

Understanding of model deployment, online inference, model gateways, version control, canary releases, continuous/dynamic batching, caching, rate limiting, circuit breaking, and inference-performance optimization.

Data and storage

Experience with one or more of object storage, relational databases, NoSQL, message queues, data lakes, or feature stores, together with an understanding of data access, lineage, and version governance.

Security and reliability

Experience with IAM, secrets management, network segmentation, vulnerability management, supply-chain security, auditability, disaster recovery, incident response, and security hardening.

Engineering discipline

Strong focus on automation, testing, observability, documentation, code quality, change management, and production operations; able to own delivery from design through production.

Communication

Able to communicate effectively with business, technology, data, risk, compliance, audit, and information-security stakeholders, translating complex technical issues into clear solutions and decisions.

Preferred Qualifications

Experience with core banking systems, financial data platforms, risk management, anti-money laundering, customer service, intelligent operations, or other financial AI use cases. Experience with GPU clusters, NVIDIA technologies, distributed training, inference acceleration, model gateways, retrieval-augmented generation, AI agents, model security, or AI developer platforms. Familiarity with Terraform, Helm, Argo CD, Prometheus, Grafana, OpenTelemetry, Kafka, Redis, PostgreSQL, object storage, or major cloud services. Demonstrated experience building an AI platform from the ground up, scaling platform adoption, or handling major production incidents. Relevant certifications in cloud, Kubernetes, information security, data, or related technologies are a plus.

Expected Outcomes and Success Measures

Area

Expected outcomes

Platform delivery

Establish a unified platform for AI training, fine-tuning, evaluation, deployment, and inference, reducing duplicated implementation.

Training efficiency

Improve reproducibility, resource utilization, experiment management, and delivery speed for training and fine-tuning workloads.

Inference performance

Optimize model-serving latency, throughput, memory utilization, concurrency, and cost per request.

Production reliability

Improve availability and recoverability through automation, observability, capacity management, and resilience testing.

Security and governance

Establish security baselines, access controls, audit records, and a closed-loop vulnerability-remediation process.

Developer experience

Provide standardized SDKs, templates, APIs, pipelines, and documentation to shorten the path from model development to production.

What We Offer

Join our banking AI infrastructure team and contribute to enterprise-grade AI platforms designed for production-scale adoption. You will collaborate with AI, cloud, data, security, risk, and business teams while working on highly reliable, secure, and performant systems. The role provides opportunities to influence critical technical decisions, platform evolution, and engineering capability development. Compensation, benefits, training, and career development are subject to the bank’s internal policies and applicable local regulations.

Suggested Level Naming

Level

Suggested title

Mid-level

AI Infrastructure Engineer

Senior

Senior AI Architecture & Infrastructure Engineer

Principal

AI Platform Architect / AI Infrastructure Principal

Technical leadership

Head of AI Infrastructure / AI Infrastructure Engineering Lead

Before publishing: Complete the seniority level, compensation range, location, reporting line, on-call requirements, primary cloud environment, main models and frameworks, and applicable local compliance requirements.

Document author: Manus AI

References

This job description was prepared based on the hiring requirements provided and does not cite external sources.

More Job Information
Job Function
  • 銀行 · 金融服務 - 一般 · 其他
Location
  • 旺角
Location
  • 油尖旺區
Employment Term
  • 全職
Education
  • 學士
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior System Analyst / Vice President
Senior System Analyst / Vice President

China CITIC Bank International Limited • Hong Kong

On-site
HKD 900,000 - 1,500,000
AI Data Architecture
AI Data Architecture

AIA Hong Kong and Macau • Hong Kong Island

On-site
HKD 800,000 - 1,100,000
Platform Engineer (Data Platform, AI, 25-35K)
Platform Engineer (Data Platform, AI, 25-35K)

Swing Consulting Ltd. • Hong Kong

On-site
10-20 Days Annual Leave
5-day Work Week
AI Engineer Lead | Team Leader | 70K
AI Engineer Lead | Team Leader | 70K

Adecco • Hong Kong

On-site
HKD 900,000 - 1,800,000
金融科技应用岗FinTech Application Specialist,FTO,公司金融方向
金融科技应用岗FinTech Application Specialist,FTO,公司金融方向

CMB Wing Lung Bank Limited • Hong Kong

On-site
HKD 600,000 - 900,000
AI Engineering Lead (GenAI / Platform)
AI Engineering Lead (GenAI / Platform)

Michael Page International (Hong Kong) Limited • Hong Kong

On-site
HKD 900,000 - 1,200,000
AI and Digital Solutions Specialist (Risk Management)
AI and Digital Solutions Specialist (Risk Management)

Leadingnation • Hong Kong

On-site
Medical insurance
Life insurance
Various allowances
AI Engineer
AI Engineer

TEKsystems • Hong Kong

On-site
HKD 600,000 - 900,000
AI Engineer (Internal Automation)
AI Engineer (Internal Automation)

South China Morning Post SCMP • Hong Kong

On-site
HKD 480,000 - 680,000
AI Infrastructure Engineer- Insurance
AI Infrastructure Engineer- Insurance

Michael Page International (Hong Kong) Limited • Hong Kong

On-site
HKD 400,000 - 600,000
Opportunity to work on a renewable contract
Exposure to cutting-edge AI technologies
Collaborative and innovative environment