AI Inference Engineer — GPU-Optimized Rust/Python (Equity)

CVFine by Instrovate Technologies

Greater London

On-site

GBP 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CVFine by Instrovate Technologies in Greater London is seeking an AI Inference Engineer. The role involves supporting transformer-based models, migrating GPU kernels, and developing a Rust-based inference server.

Candidates should have extensive experience in GPU programming and ML systems. The position offers opportunities for performance optimization and reliability improvements. Additional work in ML compilers and frameworks is a plus, making this an ideal role for proficient software engineers.

Qualifications

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques.

Responsibilities

  • Support transformer-based retrieval, text-generation, and multimodal models.
  • Port CUDA kernels to NVIDIA's CuTe DSL.
  • Develop Rust-based inference server.
  • Profile and fix bottlenecks from network ingress.
  • Build dashboards, alerts, and automated remediation.

Job description

CVFine by Instrovate Technologies in Greater London is seeking an AI Inference Engineer. The role involves supporting transformer-based models, migrating GPU kernels, and developing a Rust-based inference server.

Candidates should have extensive experience in GPU programming and ML systems. The position offers opportunities for performance optimization and reliability improvements. Additional work in ML compilers and frameworks is a plus, making this an ideal role for proficient software engineers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer | GPU-Scale Rust/Python | Equity
AI Inference Engineer | GPU-Scale Rust/Python | Equity

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
GPU Infrastructure Engineer for Scalable AI Inference
GPU Infrastructure Engineer for Scalable AI Inference

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 180,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
Equity options
Competitive compensation
GenAI Infra Engineer: Real-Time GPU Serving & Fine-Tuning
GenAI Infra Engineer: Real-Time GPU Serving & Fine-Tuning

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
AI Infrastructure Engineer: GPU & Performance Optimisation
AI Infrastructure Engineer: GPU & Performance Optimisation

twentyAI • Greater London

On-site
GBP 90,000 - 140,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
AI Hardware Acceleration Engineer for ML Performance
AI Hardware Acceleration Engineer for ML Performance

XTX Markets • Greater London

On-site
GBP 60,000 - 100,000
Onsite gym
Extensive medical benefits
Daily breakfast and lunch
+2
Senior ML Infra Architect for Large-Scale AI Simulations
Senior ML Infra Architect for Large-Scale AI Simulations

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Senior AI Infra & DS Engineer — London
Senior AI Infra & DS Engineer — London

LinuxRecruit • Greater London

On-site
GBP 100,000 - 140,000
AI Inference Compiler Engineer for Next-Gen GPUs
AI Inference Compiler Engineer for Next-Gen GPUs

NVIDIA • United Kingdom

On-site
GBP 90,000 - 150,000