Position: Principal Engineer/Architect – Generative AI & Edge Systems, Automotive
Location: San Jose CA (4 days onsite)
Employment Type: Full-Time
Overview
We are partnering with a top-five semiconductor company to hire Principal Engineer / Architect - Generative AI & Edge Systems (Automotive).
This is an exciting opportunity to join a world-class automotive engineering team developing next-generation Generative AI and deep learning solutions for edge computing platforms. Working across silicon, software, and automotive customer teams, you will help bring cutting-edge AI technologies into production for future in-vehicle systems.
Key Responsibilities
- Research and develop Generative AI and deep learning-based systems for in-vehicle domains - Infotainment, Speech, UI, and multimodal ADAS - with a focus on architectures deployable on resource-constrained edge hardware.
- Lead integration of GenAI frameworks and toolchains optimized for automotive edge compute, driving adoption of cutting-edge on-device AI techniques from research into production.
- Partner with Silicon and Software engineering teams and automotive customers to define and deliver technical solutions for GenAI and LLM requirements, with edge feasibility as a core design constraint from the start.
- Own end-to-end LLM optimization for on-device deployment on GPUs and NPUs - profiling models, identifying bottlenecks, and driving improvements at both the model level (quantization, distillation, pruning) and framework level to maximize inference performance within embedded compute and power budgets.
- Track industry and academic advances in edge AI and efficient LLM techniques, translating relevant breakthroughs into practical, deployable improvements for in-vehicle systems.
Required Skills & Experience
- Master's degree in Computer Science, Electrical Engineering, Mathematics, or related field, with 10+ years of overall software engineering experience, including 4+ years directly relevant to GenAI/LLM systems.
- Deep hands-on experience with ML/DL frameworks and tooling: TensorFlow, PyTorch, NeMo, TAO, TensorRT, CUDA, and vLLM, with working knowledge of edge-inference frameworks (e.g., TensorRT-LLM, ONNX Runtime, GGUF/llama.cpp).
- Proven experience optimizing state-of-the-art models for edge deployment - quantization, pruning, distillation - with practical tradeoff analysis across latency, memory bandwidth, and CPU vs. NPU vs. GPU compute.
- Hands-on experience deploying and optimizing multiple LLMs and multimodal models concurrently on edge devices - managing shared compute/memory budgets, model orchestration, and runtime tradeoffs when running several models side-by-side on constrained hardware.
- Practical experience with speech and multimodal AI pipelines - ASR, TTS, speech-to-intent, or vision-language fusion - in resource-constrained, real-time environments.
- Strong programming proficiency in C, C++, and Python, including performance-critical and memory-constrained code.
- Hands-on experience with embedded software, RTOS, and microcontroller-based platforms.
- Familiarity with training data preparation, curation, and open-source datasets for fine-tuning or evaluation.
- Working knowledge of in-vehicle AI technical stacks - Speech, Voice, or ADAS/AD systems.
- Exposure to automotive functional safety (ISO 26262) and cybersecurity (ISO 21434) standards.
- Demonstrated ability to move quickly in agile environments, translating research and prototypes into shippable, production-grade systems.
Additional Information
- This is a hands-on Individual Contributor role at the Principal Engineer / Architect level.
- Candidates should be willing to work onsite in San Jose, CA.
- Relocation support is available for this role