Machine Learning Engineer (Audio & LLM Stack)

Zohorecruit

Singapore

Remote

SGD 90,000 - 130,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Slash Digital PTE. LTD. is seeking a hands-on Machine Learning Engineer to own and expand our AI stack, focusing on a physics-based harmonic vocal state engine and a dual-instance vLLM deployment on GCP.

You will work directly with the founder and lead developer, owning problems end-to-end from data processing to production-grade model deployment and evaluation, with a strong emphasis on SER and audio analytics.

Qualifications

  • Proficiency in FFT and spectral features for audio analysis.
  • Hands-on with wav2vec2, HuBERT, WavLM for SER models.
  • Strong PyTorch and Python production coding skills.
  • Experience with GCP and deploying ML models.
  • Excellent written English for async collaboration.

Responsibilities

  • Own the Harmonic Vocal State Engine from design to production.
  • Build and maintain dual-instance vLLM infrastructure on GCP.
  • Develop evaluation pipelines to detect regressions and monitor quality.

Skills

Audio signal processing
FFT & Fourier analysis
Audio libraries
PyTorch
Python
GCP & MLOps
vLLM deployment
Model evaluation
English communication

Tools

GCP
vLLM
FastAPI
Git

Job description

Machine Learning Engineer (Audio & LLM Stack)

Slash is a hi-tech startup studio with a mission to build tech AI-powered products and scalable digital platforms that create real-world impact. Since 2016, we’ve partnered with ambitious enterprises and government organizations to design, engineer, and launch cutting-edge solutions — with Generative AI at the core of what we do.

We specialize in AI-powered application delivery, from product design and high-performance engineering to DevOps and AI operations. Headquartered in Singapore, our global team and clients operate R&D hubs across Southeast Asia. We are a team of entrepreneurs, engineers, and product builders dedicated to solving complex technical challenges and turning bold ideas into impactful technology.

About our Client

Our client’s product is an AI voice-sensing device, a breakthrough wearable that detects the gap between what someone says and how their voice actually sounds. Rooted in Pythagorean acoustic physics and the Navarasa framework, the system functions as a state detector rather than a conventional emotion labeler. The team is a small, fast-moving team building at the intersection of emotion labeling, ancient wisdom, and measurable science.

About the Role

We are looking for a hands-on Machine Learning Engineer to take complete ownership of our full AI stack. Your primary responsibility will be expanding our speech emotion recognition (SER) model into a physics-based harmonic vocal state engine, alongside building and maintaining our dual-instance production LLM infrastructure. In this role, you will work directly with the Founder and Lead Developer with zero bureaucracy or committee oversight. We need an engineer who excels at owning problems end-to-end.

Key Responsibilities

1. Harmonic Vocal State Engine (Audio & Physics)

Extend our inherited 30-class speech emotion recognition (SER) model, dataset, checkpoints, and pipeline into a physics-based harmonic detection system using Fourier-derived acoustic analysis and the Navarasa framework.

Design and implement an in-house model validation methodology from scratch using approaches like Gemini's emotion labeling API, cross-validation on open datasets (IEMOCAP, RAVDESS), or custom ground-truth evaluation pipelines.

2. LLM Infrastructure & Operations

Deploy and manage a dual-instance vLLM setup on GCP (g2-standard-24 instance): GPU 0: Llama 3.1 8B for fast-lane prompts; GPU 1: Qwen 2.5 32B for reasoning-heavy prompts.

Own prompt engineering, output validation, Pydantic schema enforcement, retry logic, and quality monitoring across 42 production prompts.

Handle infrastructure scaling and migrations independently without reliance on third-party vendors.

3. Model Quality & Continuous Improvement

Build evaluation pipelines to detect and catch regressions before users experience them.

Continuously optimize and refine model performance as real-world audio accumulates from our device.

Requirements
Requirements & Qualifications

Audio Signal Processing: Strong proficiency in the frequency domain (FFT, Fourier analysis, spectral features) and practical experience with audio processing libraries like librosa or torchaudio.

Audio ML Frameworks: Hands-on experience with modern audio ML models (wav2vec2, HuBERT, WavLM) and a deep understanding of training, evaluating, and fine-tuning SER models.

Core Tech Stack:

PyTorch & Python: Direct experience handling model checkpoints and training loops; ability to write clean, modular, maintainable production-level code.

Infrastructure & MLOps:

GCP & vLLM: Experience managing cloud infrastructure and direct experience deploying and serving models via vLLM.

Methodology & Communication: Proven ability to design rigorous evaluation frameworks and strong written English skills for an async-first team.

Nice to Have

Experience designing or building physics-based audio analysis systems.

Familiarity with the Navarasa framework, acoustic physics, or classical music/acoustic theory systems.

Experience using FastAPI for inference API layers.

Prior experience replacing or migrating away from third-party AI benchmarks or providers.

Genuine personal interest in sound, acoustic physics, consciousness, or the intersection of emotions and modern ML.

Hiring Process (Take-Home Assessment)

We do not conduct whiteboard interviews. Our technical evaluation is a 48-Hour Practical Assessment consisting of two deliverables:

SER Model Assessment: Review our inherited SER validation report and dataset documentation, providing your strategy for validation.

Architecture Assessment: Provide a 1-page technical evaluation of a Python file from our harmonic engine, detailing its architecture and roadmap.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer - Audio & LLM Stack (Hands-On)
ML Engineer - Audio & LLM Stack (Hands-On)

Zohorecruit • Singapore

Remote
SGD 90,000 - 130,000
(Senior) Research Engineer, Hub of the Future, IAIC
(Senior) Research Engineer, Hub of the Future, IAIC

A*STAR RESEARCH ENTITIES • Singapore

On-site
SGD 120,000 - 170,000
Senior AI Full-Stack Engineer — LLMs & Cloud
Senior AI Full-Stack Engineer — LLMs & Cloud

KEYSIGHT TECHNOLOGIES SINGAPORE (SALES) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Research Engineer / Research Engineer, Hub of the Future, IAIC
Senior Research Engineer / Research Engineer, Hub of the Future, IAIC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 120,000 - 180,000
Speech / Applied ML Engineer
Speech / Applied ML Engineer

VALSEA • Singapore

On-site
SGD 90,000 - 130,000
#EG AI Engineer
#EG AI Engineer

NCS Group • Singapore

On-site
SGD 80,000 - 120,000
Senior AI Engineer (Xora Portfolio Company)
Senior AI Engineer (Xora Portfolio Company)

Xora Innovation • Singapore

On-site
SGD 120,000 - 180,000
Machine Learning Engineer (Speech/Audio) - Singapore
Machine Learning Engineer (Speech/Audio) - Singapore

Plaud • Singapore

On-site
SGD 120,000 - 180,000
ESOP
High-Impact Environment
AI tools access
+1
AI Engineer - Agentic & GenAI Systems
AI Engineer - Agentic & GenAI Systems

JOY CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Machine Learning Engineer, Speech and Audio
Machine Learning Engineer, Speech and Audio

Plaud • Singapore

On-site
SGD 90,000 - 130,000
ESOP (Employee Stock Ownership Plan)
Top-spec laptops and equipment
Medical & Insurance Coverage