Staff AI Model Optimization Architect — Scalable Inference

Qualcomm

Cork

On-site

EUR 150,000 - 190,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Salary and stock bonus
Relocation assistance
Education assistance
Pension matching
Insurance (life/medical)
Well-being subsidies
Employee stock purchase

Job summary

QT Technologies Ireland Limited is seeking a Staff Engineer to lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm accelerators.

You will collaborate with compiler, performance, and accuracy teams to transform models into accelerator-efficient execution, balancing throughput, latency, and memory while ensuring production readiness across Day0 to deployment.

Qualifications

  • Expert level expertise in PyTorch and inference-focused model optimization.

Responsibilities

  • Architect and deliver model optimization strategies for PyTorch models on Qualcomm accelerators.
  • Drive graph capture and deployment with PyTorch, ONNX, torch.compile.
  • Design and implement fused kernels using DSL approaches (e.g., Triton).
  • Partner with compiler, performance, and accuracy teams on lowering strategies and runtime integration.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency across configurations.
  • Own transformer optimizations including KVcache management and long context performance.
  • Enable continuous batching and scheduling to optimize memory and tail latency.
  • Architect and scale distributed inference strategies across multi-core/multi-device systems.
  • Establish reusable patterns and tooling for scaling model optimizations to new hardware.
  • Debug complex performance or stability issues to production-ready solutions.

Skills

PyTorch
Python
TorchDynamo
Graph compilation
Transformer optimization
KVcache
Distributed systems
C++/Java
CPU/GPU architectures

Education

Bachelor's degree in Engineering/CS/EE
Master's degree in Engineering/CS/EE
PhD in Engineering/CS/EE

Tools

PyTorch
ONNX
Triton
TorchDynamo

Job description

QT Technologies Ireland Limited is seeking a Staff Engineer to lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm accelerators.

You will collaborate with compiler, performance, and accuracy teams to transform models into accelerator-efficient execution, balancing throughput, latency, and memory while ensuring production readiness across Day0 to deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Model Optimization Architect for Inference
Senior AI Model Optimization Architect for Inference

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Salary, stock and performance related—
Relocation and immigration support
Education Assistance
+1
AI Model Optimization Architect for LLMs & Multimodal
AI Model Optimization Architect for LLMs & Multimodal

Qualcomm • Cork

Hybrid
EUR 120,000 - 180,000
Salary and equity package
Relocation support
Education Assistance
+3
AI Model Optimization Architect - Cork, Ireland
AI Model Optimization Architect - Cork, Ireland

Qualcomm • Cork

On-site
EUR 150,000 - 190,000
Salary and stock bonus
Relocation assistance
Education assistance
+4
AI Model Optimization Architect - Cork, Ireland
AI Model Optimization Architect - Cork, Ireland

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Salary, stock and performance related—
Relocation and immigration support
Education Assistance
+1
AI Model Optimization Architect - Cork, Ireland Cork, Ireland Software Engineering Posted a day ago
AI Model Optimization Architect - Cork, Ireland Cork, Ireland Software Engineering Posted a day ago

Qualcomm • Cork

Hybrid
EUR 120,000 - 180,000
Salary and equity package
Relocation support
Education Assistance
+3
Cloud AI Inference Performance Engineer
Cloud AI Inference Performance Engineer

Qualcomm • Cork

On-site
EUR 90,000 - 150,000
Stock options
Performance bonus
Relocation assistance
+4
Principal AI Software Engineer - Inference Acceleration
Principal AI Software Engineer - Inference Acceleration

Qualcomm • Cork

On-site
EUR 110,000 - 165,000
Salary and stock options
Relocation support
Comprehensive benefits
AI Inference Performance Engineer
AI Inference Performance Engineer

Qualcomm • Cork

Hybrid
EUR 90,000 - 130,000
Salary review and performance bonus
Relocation support
Education Assistance
+2
Senior AI Systems Software Engineer
Senior AI Systems Software Engineer

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Stock bonus
Maternity/Paternity Leave
Employee stock purchase scheme
+1
AI Inference Performance Engineer
AI Inference Performance Engineer

Nutanix • Cork

On-site
EUR 90,000 - 150,000
Stock options
Performance bonus
Relocation support
+1