Remote AI Model Compression & Quantization Engineer

Framework Ventures

United States

Remote

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tether's AI model team seeks an engineer to drive model serving and inference architectures for advanced AI systems. You will optimize deployment across resource-constrained devices and edge platforms to deliver high throughput and low latency in real-world scenarios.

Collaborating with cross-functional teams, you will design robust inference pipelines, establish performance metrics, and push innovations in diffusion models, vision transformers, pruning, and quantization.

Qualifications

  • A degree in Computer Science or related field; PhD in NLP/ML preferred.
  • Knowledge of Metal Shading Language (MSL) and custom compute shaders.
  • Experience in low-level kernel and inference optimization on mobile and edge devices.
  • Understanding of modern model serving architectures and low-latency techniques.
  • Proven ability to design end-to-end inference pipelines and evaluation frameworks.
  • Experience with diffusion models and vision transformers.
  • Familiarity with pruning, quantization, flash attention and KV cache.

Responsibilities

  • Design and deploy state-of-the-art model serving architectures with high throughput and low latency across diverse environments.
  • Build, run, and monitor controlled inference tests in simulated and live environments.
  • Prepare high-quality test datasets and scenarios for real-world deployment challenges.
  • Analyze bottlenecks and optimize serving pipelines for scalability on resource-constrained systems.
  • Collaborate with cross-functional teams to integrate optimized frameworks into edge-ready production pipelines.

Skills

Model Serving
Inference Optimization
Edge Deployment
GPU Kernels
NLP & ML
Distributed Inference
Tensor Parallelism
Diffusion Models
Vision Transformers
Eagle Speculative Decoding

Education

PhD in NLP/ML
Degree in Computer Science or related field

Tools

Metal Shading Language (MSL)

Job description

Tether's AI model team seeks an engineer to drive model serving and inference architectures for advanced AI systems. You will optimize deployment across resource-constrained devices and edge platforms to deliver high throughput and low latency in real-world scenarios.

Collaborating with cross-functional teams, you will design robust inference pipelines, establish performance metrics, and push innovations in diffusion models, vision transformers, pruning, and quantization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Research Engineer - Quantization & Compression
Remote AI Research Engineer - Quantization & Compression

Tether.io • Town of Italy (NY)

Hybrid
USD 120,000 - 150,000
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

Tether.io • Town of Italy (NY)

Hybrid
USD 120,000 - 150,000
Remote AI Research Engineer: Kernel & Inference
Remote AI Research Engineer: Kernel & Inference

Visa Hunt • Germany (OH)

On-site
USD 150,000 - 230,000
Remote AI Research Engineer: LLM Architect & Training
Remote AI Research Engineer: LLM Architect & Training

Framework Ventures • United States

Remote
USD 150,000 - 230,000
Remote AI Model Inference Engineer - GPU & LoRA
Remote AI Model Inference Engineer - GPU & LoRA

Framework Ventures • United States

Remote
USD 150,000 - 210,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • Oregon (WI)

On-site
USD 100,000 - 130,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • San Francisco (CA)

On-site
USD 120,000 - 160,000
Remote AI Pre-Training Research Engineer
Remote AI Pre-Training Research Engineer

Framework Ventures • United States

Remote
USD 150,000 - 230,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • Boston (MA)

On-site
USD 100,000 - 130,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • Seattle (WA)

On-site
USD 100,000 - 140,000