AI Research Engineer, Model Compression & Quantization

Tether.io

España

On-site

PHP 4,914,004 - 7,371,007

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Tether.io seeks an AI Research Member to innovate in model compression and deployment for advanced multimodal AI systems, including large language models. Candidates with experience in quantization, knowledge distillation, and pruning are encouraged to apply.

A degree in Computer Science or related field is preferred. This role requires staying updated with the latest research and authoring papers to contribute to the field. Join Tether.io and enhance the efficiency of AI across edge devices.

Qualifications

  • Degree in Computer Science or related field; Ph.D. in NLP, Machine Learning, or a related field preferred.
  • Solid track record in AI R&D with publications in A* conferences.
  • Experience with PyTorch deep learning frameworks or equivalent.
  • Hands-on experience with model quantization (both QAT and PTQ).
  • Hands-on experience with knowledge distillation for compressing large models into smaller, efficient ones.
  • Hands-on experience with model pruning for compressing large models into smaller, efficient ones.
  • Solid understanding of neural network architectures and training processes, including transformers (LLMs, VLMs), backpropagation, optimization, and fine-tuning.
  • Familiarity with C++ is a plus, especially for implementing low-level quantization kernels or inference optimizations.

Responsibilities

  • Apply low-bit quantization to reduce model size and inference latency for generative AI models while maintaining accuracy.
  • Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models.
  • Implement pruning techniques to remove redundant parameters and attention heads.
  • Analyze trade-offs between model efficiency and accuracy across techniques.
  • Research and apply mixed-precision quantization and other advanced compression strategies.
  • Stay current with the latest research in model compression for multimodal architectures.
  • Document methodologies, experiments, and results clearly.
  • Author technical papers to advance the field of model compression.

Job description

Tether.io seeks an AI Research Member to innovate in model compression and deployment for advanced multimodal AI systems, including large language models. Candidates with experience in quantization, knowledge distillation, and pruning are encouraged to apply.

A degree in Computer Science or related field is preferred. This role requires staying updated with the latest research and authoring papers to contribute to the field. Join Tether.io and enhance the efficiency of AI across edge devices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Engineer (Model Compression & Quantization)
AI Research Engineer (Model Compression & Quantization)

Tether.io • España

On-site
PHP 4,914,000 - 7,372,000
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)

NWU • Bacoor, Hinoba-an

On-site
PHP 89,000 - 201,000
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)

Centro Escolar University • Bacoor

On-site
Hands-on learning opportunities
Collaboration with industry experts
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)

Uc Bcf • Bacoor

On-site
PHP 223,000 - 357,000
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)
Research Engineer Intern (VideoMultimodal LLM) (Start ASAP)

Stfrancis • Bacoor

On-site
PHP 2,376,000 - 3,566,000
Research Engineer Intern — LLM & Multimodal Models
Research Engineer Intern — LLM & Multimodal Models

NWU • Bacoor, Hinoba-an

On-site
PHP 89,000 - 201,000
MTS, Research Engineer
MTS, Research Engineer

Fireworks • San Mateo

On-site
PHP 800,000 - 1,200,000
Remote AI Speech & Audio Specialist (ASR/TTS)
Remote AI Speech & Audio Specialist (ASR/TTS)

TELUS Digital AI Data Solutions • Mexico

On-site
MXN 140,000 - 934,000
Remote work
Global community network
Research Engineer - I
Research Engineer - I

CodeRound • Hinoba-an

On-site
PHP 1,200,000 - 1,800,000
Research Engineer - II
Research Engineer - II

CodeRound • Hinoba-an

On-site
PHP 600,000 - 1,200,000