Embedded Systems Engineer, On-Device Inference

attention labs

Memphis, Northern (TN, KY)

Hybrid

USD 110,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

attention labs in Memphis is hiring an Embedded Systems Engineer to bring on-device inference to regulated hardware without a GPU, running efficiently on constrained CPUs.

You will optimize with ONNX, quantization, and profiling, work near the metal on audio pipelines, and collaborate with research and OEM teams to tailor the model to real devices.

Qualifications

  • Strong systems and embedded experience; proficient in C/C++.
  • Willingness to use Python for tooling.
  • Track record of running models or signal code on constrained hardware.
  • Familiar with inference runtimes and CPU-focused optimization.
  • Care for correctness in environments with limited update capability.

Responsibilities

  • Port models from server to device, with CPU efficiency.
  • Optimize inference using ONNX, quantization, profiling.
  • Work close to the metal on audio pipelines, memory and latency.
  • Partner with research and OEM customers to fit models to hardware.
  • Build tooling and tests for on-device behavior.

Skills

Embedded systems
C/C++
Python for tooling
Performance optimization
Audio DSP / RT constraints

Tools

ONNX
Quantization
Profiling

Job description

Embedded Systems Engineer, On-Device Inference

You will bring the addressee-detection model on-device, building efficient inference for regulated and OEM hardware where a GPU is not an option. This is the work that lets our models run where the audio actually happens.

  • Take our models from the server to the device, running efficiently on constrained CPUs.
  • Optimize inference with tools like ONNX, quantization, and careful profiling, without a GPU to lean on.
  • Work close to the metal on audio pipelines, memory, and latency budgets.
  • Partner with research and OEM customers to fit the model to real hardware and real constraints.
  • Build the tooling and tests that keep on-device behavior honest across devices.

What we are looking for

  • Strong systems and embedded experience: C or C++, plus comfort in Python for tooling.
  • A track record of making models or signal code run fast on constrained hardware.
  • Familiarity with inference runtimes, quantization, and profiling for CPU targets.
  • Care for correctness and reliability in environments you cannot easily update.
  • Comfort owning a hard, specific problem end to end.
  • Bonus: experience with audio DSP, real time constraints, or shipping software into physical devices.

About attention labs

attention labs is early: a small team defining a new category at the intersection of speech, cognitive neuroscience, and machine learning. We work from San Francisco, Toronto, and Memphis, and we are remote-friendly for the right person.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Speech/Audio
Machine Learning Engineer, Speech/Audio

attention labs • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Audio Modeling Engineer — Real-Time Production Inference
Audio Modeling Engineer — Real-Time Production Inference

attention labs • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Embedded AI Engineer, On-Device Models
Embedded AI Engineer, On-Device Models

Deepgram • United States

On-site
USD 140,000 - 190,000
Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
On-Device AI Inference Engineer San Jose
On-Device AI Inference Engineer San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
Embedded AI Engineer: On-Device Speech & Edge Inference
Embedded AI Engineer: On-Device Speech & Edge Inference

Deepgram, Inc. • United States

On-site
USD 140,000 - 190,000
Remote Audio Inference Engineer — High-Performance Serving
Remote Audio Inference Engineer — High-Performance Serving

Cohere • United States

Remote
USD 110,000 - 180,000
A weekly lunch stipend of $75/£75 or"
GPU Optimization Engineer
GPU Optimization Engineer

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Machine Learning Intern - KWS/AED
Machine Learning Intern - KWS/AED

Syntiant • Redwood City (CA)

Hybrid
USD 42,000 - 62,000
Research Scientist, Addressee Detection / Speech Separation
Research Scientist, Addressee Detection / Speech Separation

attention labs • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000