Runtime Engineer (C/C++)

Evollabs Tech

Dubai

On-site

AED 400,000 - 650,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evollabs Tech is seeking an experienced Senior Runtime Software Engineer to own the core runtime architecture that bridges compiler frameworks and AI hardware. You will drive runtime performance and stability, and shape memory management, scheduling, and observability—collaborating with firmware and driver teams.

You will work on concurrency models for model execution, profiling, and low-latency optimizations to ensure scalable, maintainable solutions for real-world AI workloads on our NPUs.

Qualifications

  • 5–8+ years in systems software, runtime or performance-critical infra.
  • Proficiency in C/C++ and Python.
  • Strong knowledge of OS fundamentals, multithreading, synchronization and memory management.
  • Experience near hardware (GPU/accelerator/driver/firmware).
  • Familiarity with command queues, execution engines and DMA-based systems.
  • Understanding of computational graphs, tensor execution and memory layouts.
  • Experience profiling/optimizing latency and throughput in performance-sensitive systems.

Responsibilities

  • Own core runtime architecture bridging compiler and hardware.
  • Improve runtime execution engine performance and stability.
  • Design device memory management and efficiency improvements.
  • Collaborate on command submission and hardware interaction layer.

Skills

C/C++
Python
Operating systems
Multi-threading
Memory management
Profiling/Optimization
GPU/driver/firmware familiarity

Tools

CUDA
MLIR
LLVM
ONNX Runtime
TensorRT

Job description

About Us

We are a tech company specializing in the design and development of cutting-edge, customized server hardware solutions optimized for artificial intelligence and machine learning applications. Our mission is to empower businesses and researchers to accelerate their AI initiatives by providing them with high-performance, scalable, and energy-efficient hardware infrastructure.

As a rapidly growing company at the forefront of AI hardware innovation, we are constantly seeking talented and motivated individuals to join our team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology.

You’ll Collaborate With

Compiler engineers, firmware & driver engineers, hardware architects, ML framework engineers and performance & validation teams

What You’ll Own
  • The core runtime architecture that bridges the compiler and the hardware
  • Runtime execution engine performance and stability
  • Device memory management design and efficiency
  • Command submission and hardware interaction layer (in collaboration with firmware/driver teams)
  • Concurrency and scheduling model for model execution
  • Observability tooling (profiling, logging, tracing) within the runtime
  • Technical design decisions that ensure long-term scalability and maintainability
Minimum Qualifications
  • 5–8+ years of experience in systems software, runtime systems, or performance-critical infrastructure
  • Strong proficiency in C/C++and Python
  • Solid understanding of fundamentals of operating systems, multi-threading, synchronization, memory management, and cache behavior
  • Experience working close to hardware (GPU, accelerator, driver, or firmware environments)
  • Familiarity with command queues, execution engines, and DMA-based systems
  • Understanding of computational graphs, tensor execution, and memory layouts
  • Experience profiling and optimizing latency and throughput in performance-sensitive systems
Preferred Qualifications
  • Experience building or architecting a runtime from scratch
  • Familiarity with MLIR, LLVM, or compiler-runtime interfaces
  • Experience with one or many from CUDA, ROCm, TensorRT, TVM, XLA, ONNX Runtime or similar runtimes
  • Experience integrating custom hardware backends into PyTorch, ONNX or OpenXLA
  • Knowledge of quantization and mixed-precision inference
  • Experience with multi-device or distributed execution
  • Background in AI accelerator or semiconductor environments
What Success Looks Like
  • You enable reliable end-to-end execution of compiled models on our NPU
  • Runtime overhead is minimal and performance targets are consistently met
  • Memory management and scheduling are efficient and scalable
  • The runtime architecture is robust, maintainable, and extensible
  • Compiler, firmware, and hardware teams can build confidently on top of your execution layer
  • The system supports real-world AI workloads with stability and observability

Join us in our mission to democratize AI compute — where your firmware expertise becomes the bedrock of tomorrow's AI breakthroughs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Runtime Engineer - C/C++ for AI Hardware
Senior Runtime Engineer - C/C++ for AI Hardware

Evollabs Tech • Dubai

On-site
AED 400,000 - 650,000
Network CCL Engineer
Network CCL Engineer

Evollabs Tech • Dubai

On-site
AED 300,000 - 540,000
Network CCL Engineer
Network CCL Engineer

Gateworth Group • Dubai

On-site
AED 350,000 - 700,000
Competitive salary
Bonus
Benefits
High Performance Computing Software Engineer - Supercomputing
High Performance Computing Software Engineer - Supercomputing

Institute of Foundation Models • Abu Dhabi

On-site
AED 350,000 - 700,000
Senior AI Engineer – LLM Systems
Senior AI Engineer – LLM Systems

Evollabs Tech • Dubai

On-site
AED 420,000 - 660,000
RISC-V/NPU Architect
RISC-V/NPU Architect

Evollabs Tech • Dubai

On-site
AED 500,000 - 900,000
Staff Python / PyTorch Developer — Frontend Inference Compiler - Dubai
Staff Python / PyTorch Developer — Frontend Inference Compiler - Dubai

Cerebras • United Arab Emirates

On-site
Diverse and inclusive work environment
Opportunities for continuous learning and growth
Non-corporate work culture
Staff Python / PyTorch Developer — Frontend Inference Compiler – Dubai
Staff Python / PyTorch Developer — Frontend Inference Compiler – Dubai

Cerebras Systems, Inc. • Dubai

On-site
AED 367,000 - 515,000
Equal opportunity employer
Simple, non-corporate work culture
Support for continuous learning and growth
Software Engineer
Software Engineer

CASABOT Holdings Limited • Dubai

On-site
AED 257,000 - 368,000
Competitive compensation
Stock options
Growth potential
ML Engineer
ML Engineer

Flashmind Labs • Dubai

On-site
AED 300,000 - 540,000