Edge AI Inference Engineer (C++/Llama ggml)

ITRex Group

United States

Remoto

USD 120.000 - 180.000

Tempo pieno

14 giorni+
Generatore di candidature

Una candidatura fatta su misura per questo lavoro — un curriculum e una lettera di presentazione personalizzati, perfettamente in linea con l'annuncio.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Remote flexibility
Medical benefits
Learning budget

Descrizione del lavoro

ITRex Group is seeking a hands-on C++ engineer to develop and optimize the on-device AI stack. You will port and enhance inference engines such as llama.cpp or ggml to run efficiently on edge devices and across different hardware, focusing on runtime, model loading, and performance.

This role suits engineers who enjoy low-level work and private, fast AI without cloud dependence. You’ll collaborate with researchers to translate research into production and enrich products with the latest ML

Competenze

  • Proficiency in C++ with strong systems programming background.
  • Experience with AI/ML concepts, transformers, and LLMs is preferred.
  • Familiarity with edge devices and on-device inference frameworks.

Mansioni

  • Deploy ML models to edge devices using llama.cpp and ggml frameworks.
  • Collaborate with researchers to move models from research to production environments.
  • Integrate AI features into existing products and optimize runtime performance.
  • Ensure stability and performance of the inference layer across hardware platforms.
  • Demonstrate ability to assimilate new technologies quickly.

Conoscenze

C++
JavaScript
Llama.cpp
ggml
Deep learning
Transformers/LLMs
Edge AI

Formazione

B.Sc. in Computer Science/AI/ML

Strumenti

llama.cpp
ggml

Descrizione del lavoro

ITRex Group is seeking a hands-on C++ engineer to develop and optimize the on-device AI stack. You will port and enhance inference engines such as llama.cpp or ggml to run efficiently on edge devices and across different hardware, focusing on runtime, model loading, and performance.

This role suits engineers who enjoy low-level work and private, fast AI without cloud dependence. You’ll collaborate with researchers to translate research into production and enrich products with the latest ML

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Remote AI Inference Engineer: Edge & Performance
Remote AI Inference Engineer: Edge & Performance

ITRex Group • Stati Uniti

Remoto
USD 120.000 - 170.000
Remote flexibility
Competitive salary + medical benefits
Learning opportunities
Edge AI Inference Engineer — C++ & Production ML
Edge AI Inference Engineer — C++ & Production ML

ITRex Group • Town of Poland (NY)

Ibrido
USD 100.000 - 130.000
Remote flexibility
Competitive salary and medical benefits
Career progression opportunities
Senior GPU AI Platforms Engineer - Edge LLM Inference
Senior GPU AI Platforms Engineer - Edge LLM Inference

NVIDIA • Durham (NC)

In loco
USD 224.000 - 356.500
Equity
Benefits
AI Inference Engineer
AI Inference Engineer

ITRex Group • Stati Uniti

Remoto
USD 120.000 - 180.000
Remote flexibility
Medical benefits
Learning budget
Senior Edge AI GPU & LLM Inference Engineer
Senior Edge AI GPU & LLM Inference Engineer

NVIDIA Corporation • Santa Clara (CA)

In loco
USD 224.000 - 431.000
AI Inference Engineer - Kernel & Edge Optimization
AI Inference Engineer - Kernel & Edge Optimization

Jobgether SRL • Stati Uniti

Remoto
USD 82.000 - 177.000
Remote-first team
Exposure to cutting-edge AI research
Collaborative engineering environment
Senior AI Systems Engineer - LLMs, Edge & On-Prem
Senior AI Systems Engineer - LLMs, Edge & On-Prem

Lattice Semiconductor Corp • Nellis Air Force Base Census-Designated Place (NV)

In loco
USD 199.000 - 243.000
Equity compensation
Healthcare and retirement plans
Paid time off
AI Inference Engineer QVAC
AI Inference Engineer QVAC

ITRex Group • Stati Uniti

In loco
USD 120.000 - 170.000
Remote flexibility
Competitive salary + medical benefits
Learning opportunities
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

In loco
USD 150.000 - 230.000
Equity grants
Medical plan
Vision plan
+5
Elite Senior Principal ML Inference Engineer - Edge & CUDA
Elite Senior Principal ML Inference Engineer - Edge & CUDA

Cerence AI • Stati Uniti

Ibrido
USD 185.000 - 280.000
Annual bonus
Insurance coverage
Paid time off
+3