Principal AI SoC Runtime Software Architect

Velaura

Santa Clara (CA)

On-site

USD 200,000 - 300,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Velaura is seeking a Principal AI SoC Runtime Software Architect to own the software architecture that turns Velaura's heterogeneous AI SoC into a coherent, high-performance execution platform. The role leads end-to-end runtime spanning sensor ingest, preprocessing, AI inference, postprocessing, and delivery of results to robotics and edge AI applications.

You will direct engineers across runtime, kernel, driver, firmware, multimedia, and SDK development with a strong focus on low-latency,

Qualifications

  • Extensive experience designing and building production runtime systems, embedded middleware, multimedia frameworks, or other performance-critical systems software.
  • Strong C/C++ programming skills and demonstrated ability to architect and contribute hands-on to production runtime software spanning APIs, user-space libraries, and low-level driver interfaces.
  • Strong understanding of heterogeneous and asynchronous execution, including command submission, queues, events, dependencies, synchronization, concurrency, scheduling, and resource management.
  • Strong understanding of device memory, DMA, IOMMU/SMMU, cache coherency, memory mapping, shared buffers, buffer lifetimes, and kernel/user-space memory interfaces.
  • Experience optimizing end-to-end data movement and execution across multiple hardware engines rather than focusing solely on individual kernels or accelerator performance.
  • Experience designing stable runtime APIs with compatibility, versioning, error handling, diagnostics, and recovery behavior.
  • Demonstrated ability to debug complex cross-layer correctness and performance problems using disciplined, data-driven methods, profiling, and tracing.
  • Demonstrated technical leadership across component and organizational boundaries, translating system requirements into clear architectures, interfaces, implementation guidance, and validation strategies.]
  • job_responsibilities_description []

Responsibilities

  • Define the SoC-wide execution model for coordinating workloads across heterogeneous compute and media engines, including dependency management, resource arbitration, priority and QoS, and concurrent pipeline behavior.
  • Set the architecture and technical direction for the AI inference runtime, including compiled-model execution, integration with industry-standard AI execution frameworks, and interfaces to applications and the broader Velaura SDK.
  • Own the end-to-end dataflow architecture for sensor-to-application pipelines, ensuring camera, media, preprocessing, inference, and postprocessing components operate as an integrated system.
  • Define the SoC-wide memory and buffer-sharing architecture across user space, the kernel, and heterogeneous hardware engines, establishing ownership, coherency, isolation, synchronization, and lifecycle semantics.
  • Define the division of responsibility and interface contracts among the runtime, kernel drivers, firmware, and hardware engines, including execution, completion, telemetry, fault management, and recovery semantics.
  • Partner with the compiler team to define the compiler-runtime contract, ensuring artifacts contain metadata for runtime load, validate, execute, profile, and compatibility across releases.
  • Establish a system-wide observability and performance architecture that correlates behavior across software and hardware, targeting latency, throughput, bandwidth, power, utilization, and predictability.
  • Define the runtime resilience and validation architecture, including fault-containment and recovery policies, acceptance criteria, and qualification across correctness, concurrency, and performance.

Skills

C/C++
Runtime systems
Heterogeneous compute
Linux kernel
Performance optimization
Profiling & debugging
Technical leadership

Tools

ONNX Runtime
TensorRT
OpenVINO
GStreamer
V4L2

Job description

Role Overview

We are looking for a Principal AI SoC Runtime Software Architect to own the software architecture that turns Velaura's heterogeneous AI SoC into a coherent, high-performance execution platform.

This role will define and lead development of the end-to-end runtime spanning sensor ingest, preprocessing, AI inference, postprocessing, and delivery of results to robotics and other physical and edge AI applications. The runtime will coordinate execution and data movement across the SoC's CPU cores, AI accelerator, vision and multimedia engines, and other embedded processors. The complete runtime must coordinate heterogeneous workloads, manage ownership, synchronization, and safe reuse of shared data buffers, minimize data movement, provide predictable low-latency execution, recover from failures, and expose cohesive APIs and observability to applications and SDK components.

The ideal candidate combines deep runtime and systems-software expertise with a strong understanding of heterogeneous compute, Linux kernel and driver interfaces, DMA and shared-memory architectures, and performance-sensitive AI or multimedia pipelines. This will be a hands-on principal architect and technical lead who directs engineers across runtime, kernel, driver, firmware, multimedia, and SDK development.

Responsibilities
  • Define the SoC-wide execution model for coordinating workloads across heterogeneous compute and media engines, including dependency management, resource arbitration, priority and QoS, and concurrent pipeline behavior.

  • Set the architecture and technical direction for the AI inference runtime, including compiled-model execution, integration with industry-standard AI execution frameworks, and interfaces to applications and the broader Velaura SDK.

  • Own the end-to-end dataflow architecture for sensor-to-application pipelines, ensuring that camera, media, preprocessing, inference, and postprocessing components operate as an efficient and coherent system.

  • Define the SoC-wide memory and buffer-sharing architecture across user space, the kernel, and heterogeneous hardware engines, establishing clear ownership, coherency, isolation, synchronization, and lifecycle semantics.

  • Define the division of responsibility and interface contracts among the runtime, kernel drivers, firmware, and hardware engines, including execution, completion, telemetry, fault management, and recovery semantics.

  • Partner with the compiler team to define the compiler-runtime contract, ensuring compiled artifacts contain the metadata and execution information needed for the runtime to load, validate, execute, profile, and maintain compatibility across releases.

  • Establish a system-wide observability and performance architecture that correlates behavior across software and hardware layers and enables optimization against latency, throughput, bandwidth, power, utilization, and predictability goals.

  • Define the runtime resilience and validation architecture, including fault-containment and recovery policies, architecture-level acceptance criteria, and qualification across correctness, concurrency, compatibility, performance, and sustained workloads.

Required Qualifications
  • Extensive experience designing and building production runtime systems, embedded middleware, multimedia frameworks, or other performance-critical systems software.

  • Strong C/C++ programming skills and demonstrated ability to architect and contribute hands-on to production runtime software spanning application-facing APIs, user-space libraries, and low-level driver, firmware, and hardware interfaces.

  • Strong understanding of heterogeneous and asynchronous execution, including command submission, queues, events, dependencies, synchronization, concurrency, scheduling, and resource management.

  • Strong understanding of device memory, DMA, IOMMU/SMMU, cache coherency, memory mapping, shared buffers, buffer lifetimes, and kernel/user-space memory interfaces.

  • Experience optimizing end-to-end data movement and execution across multiple hardware engines rather than focusing solely on individual kernels or accelerator performance.

  • Experience designing stable runtime APIs with well-defined compatibility, versioning, error handling, diagnostics, and recovery behavior.

  • Demonstrated ability to debug complex cross-layer correctness and performance problems using disciplined, data-driven methods, profiling, and tracing.

  • Demonstrated technical leadership across component and organizational boundaries, including translating system requirements into clear architectures, interfaces, implementation guidance, and validation strategies.

Preferred Qualifications
  • Direct experience developing or extending AI inference runtimes or execution providers, such as ONNX Runtime, TensorRT-like runtimes, OpenVINO, TensorFlow Lite delegates, Qualcomm QNN/SNPE, TVM runtimes, or comparable systems for NPUs, GPUs, DSPs, or other accelerators.

  • Linux kernel development or upstream contribution experience involving device, accelerator, media, or shared-memory subsystems.

  • Experience with robotics, autonomous systems, edge AI, ROS 2, camera pipelines, ISP integration, V4L2/media, GStreamer, or other sensor-driven workloads.

  • Familiarity with compiled-model artifacts, quantized execution, tensor layouts, graph partitioning, memory planning, and compiler/runtime integration.

  • Experience with embedded Linux SDKs, production deployment, long-term runtime/API compatibility, or the security and isolation requirements of multi-process accelerator systems.

$200,000 - $300,000 a year

Compensation & Benefits

At Velaura, we believe exceptional talent deserves exceptional rewards. Compensation for this role includes a base salary and very competitive equity compensation, allowing team members to share in the company's long-term success.

The base pay listed represents a good faith estimate that the Company reasonably expects to pay for this position at the time of hire. Actual compensation will depend on multiple factors, including the candidate's skills, qualifications, relevant experience, and geographic location.

In addition to base salary and equity compensation, Velaura offers a comprehensive benefits package that may include medical, dental, and vision coverage; paid time off; flexible work arrangements; professional development opportunities; and other benefits designed to support the well-being and growth of our team.

Velaura is committed to pay equity and transparency and regularly benchmarks compensation to ensure we remain competitive in the market.

Why Velaura?

Velaura is building next-generation compute technology for cloud, edge, and PhysicalAI. Our solutions will enable robots, autonomous systems, drones, and other intelligentmachines to operate efficiently in the physical world.
This is an opportunity to help build foundational technology at a time when the industry is undergoing fundamental change. You will work alongside experienced leaders, architects, engineers, and operators who have delivered industry-defining products across mobile, cloud, semiconductor, and AI platforms. If you enjoy solving difficult problems, working across disciplines, and helping shape thefuture of Physical AI.

Equal Employment Opportunity and Accommodations

Velaura is an Equal Opportunity Employer that is committed to inclusion and diversity.Qualified applicants will receive consideration for employment without regard to race,color, religion, national origin, gender, sexual orientation, gender identity, disability orprotected veteran status. We also take affirmative action

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Software Lead – Physical AI
Platform Software Lead – Physical AI

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Equity participation
Medical/dental/vision
Paid time off
+1
Senior RTL Engineer
Senior RTL Engineer

Velaura • California (MO)

On-site
USD 200,000 - 500,000
Senior Emulation Engineer
Senior Emulation Engineer

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
AI Systems Architect (Models & Hardware Co-Design)
AI Systems Architect (Models & Hardware Co-Design)

Velaura • Austin (TX)

On-site
USD 200,000 - 500,000
Medical coverage
Dental & Vision
Paid time off
+3
Design Verification Engineer- AI Accelerator Lead
Design Verification Engineer- AI Accelerator Lead

Velaura AI, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
Design Verification Engineer- AI Accelerator Lead
Design Verification Engineer- AI Accelerator Lead

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 350,000
Design Verification Engineer- CPU
Design Verification Engineer- CPU

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 350,000
Equity participation
Competitive base salary
Performance incentives
Senior Formal Verification Engineer
Senior Formal Verification Engineer

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Equity participation
Competitive base salary
Medical, dental, vision
Senior Systems Administrator/ IT Engineer
Senior Systems Administrator/ IT Engineer

Velaura • Santa Clara (CA)

On-site
USD 145,000 - 200,000
Equity participation
Medical, dental, and vision coverage
Flexible work arrangements
RTL Power Engineer
RTL Power Engineer

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 350,000