Audio Software Engineer San Jose

Hark, Inc.

San Jose (CA)

On-site

USD 170,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hark, Inc. is looking for a Member of Technical Staff specializing in Real-Time Audio to enhance its voice agent technology. The role demands strong expertise in audio quality management, DSP fundamentals, and real-time audio coding, as well as end-to-end feature ownership from prototyping to production.

Located in San Jose, California, this full-time role offers a competitive salary ranging from $170,000 to $400,000 annually. The ideal candidate will have a passion for improving audio interactions and collaborating cross-functionally within the team.

Qualifications

  • 5+ years of software engineering experience.
  • Experience shipping real-time audio products.
  • Strong DSP fundamentals required.

Responsibilities

  • Own audio quality on client applications.
  • Work the WebRTC audio path from end to end.
  • Manage features from prototyping to production.

Skills

Software engineering experience
Real-time audio systems
WebRTC
Digital Signal Processing (DSP)
C++
Rust
TypeScript

Job description

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

We're hiring a Member of Technical Staff (Real-Time Audio) to join our Product Engineering team. Hark’s voice agent holds real-time, full-duplex conversations with people in homes, cars, and noisy rooms. That experience is only as good as the audio underneath it.

This role owns the real-time audio that makes conversations feel natural (echo cancellation, noise suppression, and voice activity detection) as production code in our live client. This is not a research role and not a DSP theory role. We're looking for someone who can do both: understand the signal processing and ship the code.

Responsibilities
  • Own audio quality on the client: echo, self-interruption, dropouts, and clipping
  • Work the WebRTC audio path end to end: AEC, noise suppression, and VAD
  • Ship DSP to the client as C++/Rust compiled to WebAssembly, and as TypeScript in the audio pipeline
  • Tune endpointing, interruption, and turn-taking so the agent listens like a person
  • Reduce conversational latency and artifacts across the streaming pipeline
  • Work in our React/TypeScript client where audio meets the UI
  • Manage features end-to-end from prototyping through production
  • Collaborate with designers, platform engineers, and our speech team.
Requirements
  • 5+ years of software engineering experience
  • Shipped real-time audio into a product used by real users
  • Hands‑on experience with WebRTC, AEC (echo cancellation), noise suppression, and VAD
  • Strong DSP fundamentals: adaptive filtering, STFT, resampling, and gain control
  • Comfort with latency, buffering, and sample rates in a streaming audio pipeline
  • Owns features end-to-end and works comfortably in a shared production codebase.
Bonus Qualifications
  • Experience working at a voice, speech, or video‑conferencing company
  • ML for audio: noise suppression, VAD, or source separation (e.g. RNNoise, DeepFilterNet, Silero VAD), and on‑device inference (ONNX Runtime, Core ML)
  • Familiarity with WebRTC internals (the Audio Processing Module, AEC3, Opus) and voice‑agent frameworks (LiveKit, Pipecat)
  • TypeScript and React, and comfort working across the product frontend
  • Experience with target-speaker isolation, diarization, or barge‑in and turn‑detection systems for conversational AI.
Compensation

The US base salary range for this full‑time position is between $170,000–$400,000 annually.

The pay offered for this position may vary based on several individual factors, including job‑related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Audio Software Engineer
Audio Software Engineer

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000
Audio Quality and Data Engineer San Jose
Audio Quality and Data Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 250,000 - 300,000
Audio Firmware Engineer San Jose
Audio Firmware Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 120,000 - 300,000
Audio Quality and Data Engineer
Audio Quality and Data Engineer

Hark • San Jose (CA)

Hybrid
USD 250,000 - 300,000
Lead Audio ML Engineer
Lead Audio ML Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Audio Firmware Engineer
Audio Firmware Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Real-Time Audio Engineer - DSP, WebAssembly & WebRTC
Real-Time Audio Engineer - DSP, WebAssembly & WebRTC

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000
Real-Time Audio Engineer (WebRTC DSP, Production)
Real-Time Audio Engineer (WebRTC DSP, Production)

Hark, Inc. • San Jose (CA)

On-site
USD 170,000 - 400,000
Full-Stack Engineer San Jose
Full-Stack Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 170,000 - 400,000
Frontend Engineer San Jose
Frontend Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 170,000 - 400,000