Remote AI Research Engineer: Kernel & Inference

Visa Hunt

Germany (OH)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tether is seeking a specialist to advance AI model serving and inference for edge and on-device applications. You will design pipelines that deliver high throughput with minimal memory footprints across diverse platforms.

You will optimize latency, throughput, and energy use while developing robust evaluation frameworks and collaborating with cross-functional teams to deploy production-ready solutions.

Qualifications

  • PhD in NLP or related field with strong AI R&D track record.
  • Experience in low-latency model serving and inference.
  • Knowledge of Metal Shading Language (MSL) and GPU kernels.
  • Background in edge/mobile deployment and performance optimization.

Responsibilities

  • Design and deploy state-of-the-art model serving architectures with high throughput and low latency.
  • Build and monitor inference tests in simulated and live environments.
  • Prepare test datasets for real-world deployment on low-resource devices.
  • Analyze bottlenecks in serving pipelines and optimize memory usage.
  • Collaborate with cross-functional teams to integrate optimized serving in production.

Skills

NLP R&D
Model serving
Inference optimization
GPU kernels
Edge devices
MSL
Mobile GPUs

Education

PhD in NLP
CS degree

Tools

Metal
Shaders
TensorRT
ONNX

Job description

Tether is seeking a specialist to advance AI model serving and inference for edge and on-device applications. You will design pipelines that deliver high throughput with minimal memory footprints across diverse platforms.

You will optimize latency, throughput, and energy use while developing robust evaluation frameworks and collaborating with cross-functional teams to deploy production-ready solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Edge-Optimized AI Inference Kernel Engineer
Edge-Optimized AI Inference Kernel Engineer

Framework Ventures • United States

Remote
USD 180,000 - 260,000
Remote AI Model Compression & Quantization Engineer
Remote AI Model Compression & Quantization Engineer

Framework Ventures • United States

Remote
USD 180,000 - 260,000
Remote AI Model Inference Engineer - GPU & LoRA
Remote AI Model Inference Engineer - GPU & LoRA

Framework Ventures • United States

Remote
USD 150,000 - 210,000
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

Tether.io • Town of Italy (NY)

Hybrid
USD 120,000 - 150,000
Remote AI Research Engineer: LLM Architect & Training
Remote AI Research Engineer: LLM Architect & Training

Framework Ventures • United States

Remote
USD 150,000 - 230,000
Remote AI Research Engineer - Quantization & Compression
Remote AI Research Engineer - Quantization & Compression

Tether.io • Town of Italy (NY)

Hybrid
USD 120,000 - 150,000
Remote Product Engineer: AI On-Device, Lead Cross-Stack
Remote Product Engineer: AI On-Device, Lead Cross-Stack

Framework Ventures • United States

Remote
USD 120,000 - 190,000
Senior AI Kernel Engineer — Edge Inference & Optimization
Senior AI Kernel Engineer — Edge Inference & Optimization

quadric.io, Inc • Burlingame (CA)

On-site
USD 120,000 - 150,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Inference Engineer — GPU & Edge Optimization
Senior AI Inference Engineer — GPU & Edge Optimization

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000