Staff Research Software Engineer, AI Inference & Efficiency

Google Inc.

Santo Niño 1st

On-site

PHP 10,880,000 - 16,815,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Google Research Singapore is advancing fundamental AI capabilities, pioneering Generative AI models, and leading a small agile team in a fast-paced environment. The role focuses on building scalable AI inference systems and collaborating across teams to publish research and drive innovations for billions of users.

You will design novel algorithms for efficient LLM inference, optimize low-level software, and work with hardware accelerators to push performance boundaries in AI deployment.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.
  • 5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.
  • 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • Experience integrating generative AI tools or LLM interfaces into workflows.

Responsibilities

  • Design and implement novel algorithms that optimize the inference stack for LLMs to ensure the highest throughput and lowest latency for billions of users including sampling techniques, speculative decoding advancements, model compression, knowledge distillation and quantization strategies.
  • Work on development and optimization of the low-level software that runs AI models efficiently. Implement and optimize custom kernels for inference, focusing on improving latency, throughput, and memory usage, across different hardware and model architectures.
  • Co-design algorithms that exploit the unique architectural advantages of custom hardware and accelerators. Collaborate with hardware architects on next-generation systems and use a deep understanding of both model architecture and hardware to inform algorithm selection.
  • Amplify impact and influence the research ecosystem through publications and scientific dissemination.

Education

Bachelor’s degree or equivalent practical experience.
8 years of experience in software development.
5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.
5 years of experience with one or more of the following: Speech/audio, reinforcement learning, ML infrastructure, or related ML field.
5 years of experience with ML design and ML infrastructure (deployment, evaluation, data processing, debugging, fine tuning).
Experience integrating generative AI tools or LLM interfaces into workflows.

Job description

Google Research Singapore is advancing fundamental AI capabilities, pioneering Generative AI models, and leading a small agile team in a fast-paced environment. The role focuses on building scalable AI inference systems and collaborating across teams to publish research and drive innovations for billions of users.

You will design novel algorithms for efficient LLM inference, optimize low-level software, and work with hardware accelerators to push performance boundaries in AI deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Research Software Engineer, Google Research - Singapore
Staff Research Software Engineer, Google Research - Singapore

Google Inc. • Santo Niño 1st

On-site
PHP 10,880,000 - 16,815,000
Staff Research Scientist, ML Efficiency, Google Research - Singapore
Staff Research Scientist, ML Efficiency, Google Research - Singapore

Google Inc. • Santo Niño 1st

On-site
PHP 1,800,000 - 3,000,000
Staff Research Scientist - Generative AI Efficiency
Staff Research Scientist - Generative AI Efficiency

Google Inc. • Santo Niño 1st

On-site
PHP 1,800,000 - 3,000,000
Software Engineer, Machine Learning, TPU Workload Optimization - Singapore
Software Engineer, Machine Learning, TPU Workload Optimization - Singapore

Google Inc. • Santo Niño 1st

On-site
PHP 1,200,000 - 1,800,000
Senior GenAI Application Engineer (Contract) for Enterprise AI
Senior GenAI Application Engineer (Contract) for Enterprise AI

RiDiK (a Subsidiary of CLPS. Nasdaq: CLPS) • Santo Niño 1st

On-site
PHP 8,824,000 - 11,765,000
Staff Forward Deployed Engineer, Developer AI, Google Cloud
Staff Forward Deployed Engineer, Developer AI, Google Cloud

Google Inc. • Hinoba-an

On-site
PHP 11,257,000 - 16,260,000
ML Software Engineer - TPU Training & Inference Optimizer
ML Software Engineer - TPU Training & Inference Optimizer

Google Inc. • Santo Niño 1st

On-site
PHP 1,200,000 - 1,800,000
Staff AI Platform Engineer
Staff AI Platform Engineer

V2 Solutions • Hinoba-an

On-site
PHP 1,500,000 - 2,500,000
Customer Engineer, AI, Google Cloud - Singapore
Customer Engineer, AI, Google Cloud - Singapore

Google Inc. • Santo Niño 1st

Hybrid
PHP 900,000 - 1,500,000
Senior GenAI Application Engineer
Senior GenAI Application Engineer

RiDiK (a Subsidiary of CLPS. Nasdaq: CLPS) • Santo Niño 1st

On-site
PHP 8,824,000 - 11,765,000