AI Runtime Engineer

Insider, Inc.

United States

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Insider, Inc. is looking for an AI Runtime Engineer in the United States. The position requires expertise in developing low-latency, high-performance runtime software for AI accelerators. Candidates should have a Bachelor’s or Master’s degree in a relevant field and at least 3 years of experience in low-level runtime software development. Responsibilities include optimizing the execution stack and collaborating with hardware teams. The role is an opportunity to work on cutting-edge AI technologies in a diverse and dynamic team.

Qualifications

  • 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems.
  • Deep understanding of task scheduling, concurrency, and memory hierarchy.
  • Experience with hardware-aware optimizations and dataflow architectures.

Responsibilities

  • Develop and optimize the AI runtime software stack for executing deep learning workloads on AI accelerators.
  • Implement task scheduling, memory management, and kernel execution strategies for efficient computation.
  • Optimize data movement between host and device using PCIe, DMA, shared memory.

Skills

C/C++ programming
Low-level systems programming
Task scheduling
Memory management
Debugging and profiling

Education

Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field

Tools

ONNX Runtime
TensorRT
TVM
OpenVINO

Job description

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next‑generation in‑memory computing technology provides orders‑of‑magnitude higher compute efficiency and density compared to today’s best‑in‑class solutions. The high‑performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems.

About the Role

EnCharge AI is seeking an AI Runtime Engineer to develop and optimize the execution stack for our next‑generation AI accelerator. In this role, you will work on low‑latency, high‑performance runtime software that enables efficient execution of deep learning models on specialized hardware. You will collaborate with hardware, compiler, and AI framework teams to deliver optimized AI inference and training performance across cloud and edge environments.

Responsibilities
  • Develop and optimize the AI runtime software stack for executing deep learning workloads on AI accelerators.
  • Implement task scheduling, memory management, and kernel execution strategies for efficient computation.
  • Optimize data movement between host and device using PCIe, DMA, shared memory.
  • Design and implement high‑performance APIs for AI inference frameworks such as OpenVino, ONNX Runtime, vLLM.
  • Work on graph execution optimizations, including kernel fusion, pipelining, tensor tiling, and caching.
  • Integrate runtime components with AI compilers (LLVM, MLIR, XLA, TVM) for optimized execution.
  • Ensure scalability and reliability of the AI runtime for cloud‑based and edge AI deployments.
Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • 3+ years of experience in developing low‑level runtime software for AI accelerators, GPUs, or HPC systems.
  • Strong proficiency in C/C++ and low‑level systems programming.
  • Deep understanding of task scheduling, concurrency, and memory hierarchy.
  • Experience with hardware‑aware optimizations and dataflow architectures.
  • Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO).
  • Experience with low‑latency, high‑throughput workload execution for AI models.
  • Strong debugging and profiling skills for optimizing AI execution performance.
  • Exposure to AI model deployment pipelines (Triton, TensorFlow Serving).

EnchargeAI is an equal employment opportunity employer in the United States.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Engineer
AI Research Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 150,000
Device Driver Engineer
Device Driver Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
Embedded SW Engineer
Embedded SW Engineer

Insider, Inc. • United States

On-site
USD 90,000 - 130,000
Runtime Engineer
Runtime Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 240,000
Runtime Engineer
Runtime Engineer

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity opportunities
Medical, dental, and vision coverage
Retirement savings plan
+2
Edge AI Runtime Engineer
Edge AI Runtime Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
LLM Inference Deployment Engineer
LLM Inference Deployment Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
Senior Runtime Engineer
Senior Runtime Engineer

Cerebras Systems • United States

On-site
USD 100,000 - 130,000
Equal opportunity employer
Continuous learning and growth opportunities
Senior Runtime Engineer
Senior Runtime Engineer

Cerebras • Raleigh (NC)

On-site
USD 140,000 - 230,000
Senior AI Software Architect - Runtime
Senior AI Software Architect - Runtime

Intel Corporation • Santa Clara (CA)

Hybrid
USD 195,200 - 361,200
Stock bonuses
Health benefits
Retirement plan
+1