AI Researcher (KV cache optimization)

Glasswing

Australia

On-site

AUD 120,000 - 160,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Glasswing, an Australian startup, is seeking an AI Researcher to drive innovations in neural network and LLM optimization. You will design and implement optimized inference pipelines and techniques to enhance model efficiency and cost.

The role demands deep expertise in high-performance computing, LLM optimization, and significant experience with AI systems. Join us to forge the future of AI and contribute to our fast-growing company backed by U.S. venture capital.

Qualifications

  • 8+ years of experience in high-performance computing or AI systems.
  • Proven experience in LLM and neural network optimization.
  • Strong communication skills to convey technical details.

Responsibilities

  • Design and implement optimized inference pipelines for LLM workloads.
  • Research and apply techniques like pruning and quantization.
  • Build advanced profiling tools to find bottlenecks.
  • Maintain frameworks for model evaluation and benchmarking.
  • Mentor engineers and researchers in optimization techniques.

Skills

Deep Systems Expertise
LLM & NN Optimization
Communication Skills
Experience with Evaluation Frameworks
Neural Network Compression
C/C++/CUDA Experience

Education

8+ years of experience in high-performance computing or AI systems

Tools

CUDA

Job description

We’re hiring an AI Researcher to join us in tackling one of the most important and fast-moving areas in AI today: driving advanced research in neural network and LLM optimization, identify the most promising opportunities, and translate them into production-ready innovations. You will evaluate and apply emerging techniques: pruning, quantization, inference acceleration, memory optimization, scheduling, runtime tuning and determine what will materially improve LLM inference and KV cache performance, model efficiency, and cost, then partner with engineering to ship the changes that make a measurable difference.

Responsibilities:
  • Inference & Compute Optimization: Design and implement highly optimized inference pipelines and computational kernels to accelerate LLM and neural network workloads, leveraging low-level techniques such as SIMD vectorization, cache-aware memory access patterns, and hardware-specific tuning.
  • Neural Network Compression & Model Optimization: Research and implement pruning, quantization, and other compression techniques to reduce model size and accelerate inference while preserving accuracy. Apply both in-training and post-training optimization methods across LLM and vision model workloads.
  • Profiling & Observability: Build and utilize advanced profiling tools to identify bottlenecks across the inference and training stack—from memory bandwidth and cache utilization to CPU-side data preprocessing stalls and end-to-end pipeline throughput.
  • Evaluation & Benchmarking: Design and maintain rigorous evaluation and benchmarking frameworks for systematic model comparison across optimization configurations. Develop automated pipelines (e.g., LLM-as-a-judge) to measure the impact of optimization techniques on model quality and performance.
  • Mentorship: Act as a technical lead for engineers and researchers, fostering a culture of high-performance code, rigorous benchmarking, and research-to-production excellence. Drive team growth, technical interviews, and cross-functional collaboration.
Required Qualifications:
  • Deep Systems Expertise: 8+ years of experience in high-performance computing, AI systems, or low-level software optimization. Deep familiarity with performance‑critical development including CPU/GPU architecture, memory hierarchies, SIMD/vectorization, and profiling-driven tuning, CUDA
  • LLM & NN Optimization Track Record: Proven experience optimizing neural networks and LLMs through techniques such as pruning, quantization, and inference acceleration, with a demonstrated path from research to production deployment.
  • Communication: Ability to translate complex systems-level constraints and optimization trade-offs into actionable research directions for modeling and engineering teams.
  • Experience building evaluation frameworks, ML observability, or developer tools that help researchers understand and compare model performance across optimization configurations.
  • A history of working on neural network compression, inference acceleration, or applied AI research problems that required bridging algorithmic research with high-performance implementation.
  • Patent authorship or published research in AI/ML optimization.
  • Experience with C/C++/ CUDA inference engines, x86 intrinsics, or similar low-level performance work is a strong plus.
About the Company

An Australian startup with an office in Israel is developing a compute-optimization engine. Backed by leading U.S. venture capital , we’re looking for exceptional talent to join us as true partners on our journey.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Researcher (Neural Network & LLM Optimization)
Senior AI Researcher (Neural Network & LLM Optimization)

Glasswing • Council of the City of Sydney

Hybrid
AUD 130,000 - 160,000
Senior AI Systems & Inference Optimization Engineer
Senior AI Systems & Inference Optimization Engineer

Glasswing • Australia

On-site
AUD 120,000 - 160,000
Senior Performance Engineer
Senior Performance Engineer

CommonAI C.I.C. • Town Of Cambridge

On-site
AUD 132,000 - 208,000
Competitive salary package and pension
Professional development opportunities
Networking with tech and academia
+1
Senior AI / ML Engineer
Senior AI / ML Engineer

Artificial Analysis, Inc. • City of Melbourne

On-site
AUD 198,000 - 296,000
Competitive compensation including-equ
Equity
Senior AI / ML Engineer
Senior AI / ML Engineer

Artificial Analysis • City of Melbourne

On-site
AUD 180,000 - 240,000
Equity
On-site work in SF & Melbourne
Technical Lead
Technical Lead

Future Secure AI • City of Brisbane

Hybrid
AUD 120,000 - 160,000
State-of-the-art technology
Competitive salary
Flexible work environment
+1
Senior AI Inference & NN Optimization Lead
Senior AI Inference & NN Optimization Lead

Glasswing • Council of the City of Sydney

Hybrid
AUD 130,000 - 160,000
Ai Engineer | Ai Infrastructure
Ai Engineer | Ai Infrastructure

Pearson Carter • Australia

Hybrid
AUD 100,000 - 140,000
Senior AI Engineer
Senior AI Engineer

Clyphor AI • Sydney

Hybrid
AUD 120,000 - 160,000
Competitive salary package with equity options
Professional development budget for conferences and courses
Health insurance and wellness programs
+2
Machine Learning Systems & Performance Engineer
Machine Learning Systems & Performance Engineer

Westbury Partners • Sydney

On-site
AUD 150,000 - 230,000