AI Accelerator Software Senior Principal Engineer- Framework Integration

Ampere Computing

Warszawa

Hybrid

PLN 417,000 - 625,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Premium health care
Paid time off
Office amenities
Remote work support
Sports card

Job summary

Ampere Computing is seeking an AI Accelerator Software Senior Principal Engineer – Framework Integration to lead strategy for high-performance deep learning inference across Ampere accelerator platforms. You will set direction for enabling major ML frameworks and optimizing for throughput, latency, and efficient compute across data centers to edge.

This role requires deep architectural ownership, cross-team influence, and mentorship to raise the bar for quality, maintainability, and performance

Qualifications

  • BS/MS/PhD with 12/8/5+ years experience respectively, in relevant field.
  • Proven experience with integrating PyTorch, ONNX, and llama.cpp into accelerators.
  • Linux accelerator runtime/driver experience is a plus.

Responsibilities

  • Lead end-to-end framework integration for AI accelerators.
  • Drive full-stack SW/HW acceleration from inference to kernel optimization.
  • Improve model performance and correctness for PyTorch/ONNX/llama.cpp ecosystems.
  • Provide HW/SW co-design directions to maximize efficiency and throughput.
  • Build and evolve software components for AI accelerators.
  • Collaborate across compiler/runtime/kernel and systems teams from cloud to edge.
  • Mentor and raise standards for quality and performance.

Skills

Python
C/C++
Performance engineering
Linux experience
ML concepts
Framework integration

Education

BS in CS/CE/EE or related field
MS in CS/CE/EE
PhD in related field

Tools

PyTorch
ONNX
llama.cpp
vLLM
SGLang

Job description

AI Accelerator Software Senior Principal Engineer- Framework Integration

Ampere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.

As a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.

Join us at Ampere and work alongside a passionate and growing team-we’d love to have you apply!

About the Role

As an AI Accelerator Software Senior Principal Engineer – Framework Integration, you will lead end-to-end technical strategy and delivery for high-performance deep learning inference across Ampere accelerator platforms. You will set direction for how major ML frameworks are enabled and optimized for our hardware, ensuring high throughput, low latency, and efficient compute/memory utilization for current and next-generation AI workloads spanning data centers to edge.

This role is distinguished by deep technical ownership, architecture leadership, and cross-team influence—driving outcomes from performance modeling and integration strategy through production-ready runtime and kernel behavior.

What You’ll Achieve
  • Framework integration leadership (PyTorch / ONNX / llama.cpp)
    Own and advance integration of major deep learning frameworks—PyTorch, ONNX, llama.cpp, and related tooling—into the Ampere deep learning accelerator backend, enabling robust execution of real-world model graphs and operators.
  • Full-stack acceleration across the SW/HW execution path
    Drive acceleration across the end-to-end stack, including (as applicable):
    • inference serving and orchestration enablement
    • compiler/graph lowering and optimization
    • runtime library and execution management
    • user-mode execution paths and performance-critical interfaces
    • compute kernel development and micro-optimizations
    • profiling, benchmarking, and continuous performance tuning
  • Model enablement: performance + accuracy
    Lead efforts to improve performance and correctness for models using popular frameworks and serving stacks such as vLLM and SGLang, ensuring stable behavior under production inference patterns (prefill/decode, batching, KV cache behavior, scheduling, etc.).
  • Hardware/software co-design and optimization
    Provide technical direction for HW/SW co-optimization of existing and evolving AI architectures to:
    • maximize computational efficiency
    • increase sustained throughput
    • reduce latency and variance
    • improve scalability across cores, memory hierarchies, and system configurations
    • raise the ceiling on what Ampere platforms can deliver
  • Build and evolve state-of-the‑art AI accelerator software
    Contribute to and shape the architecture of software/hardware AI co‑processors and accelerators, defining reusable components, reference implementations, and performance guardrails.
  • Cross‑functional technical collaboration and influence
    Partner with compiler/runtime/kernel, platform, and systems teams to integrate and validate AI capabilities in Ampere’s platforms and accelerators from cloud to edge.
  • Mentorship and technical excellence
    Set engineering standards through code reviews, design reviews, benchmark methodologies, and mentorship—raising the overall bar for quality, maintainability, and performance.
About You
  • Education & experience: BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or a related technical field & 12 years of related experience; or MS degree & 8 years; or PhD & 5 years
  • Deep framework expertise: Proven experience with software development focused on PyTorch, ONNX, and llama.cpp, including integration, graph/operator enablement, and performance‑focused engineering.
  • Linux accelerator runtime / driver experience (preferred): Experience building or extending user‑mode drivers and/or runtime libraries for GPUs or deep learning accelerators on Linux is a plus.
  • Strong systems programming + performance tuning: Deep expertise in Python and C/C++, with a strong track record in performance engineering (profiling, optimization, throughput/latency analysis, memory behavior).
  • Solid ML/AI fundamentals: Strong understanding of AI/ML concepts (neural networks, data processing frameworks), and familiarity with modern model families including Transformers and Diffusion architectures.
  • Fluent with modern AI development tools (preferred): Comfortable using modern AI programming tools and workflows such as Codex or Claude Code.
What We’ll Offer

At Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long‑term incentive, and comprehensive benefits. The full base pay range for this role is between 416,500 PLN and 624,500 PLN.

  • Premium medical health care, so that you and your family members can feel secure in your health.
  • A generous paid time off policy so that you can embrace a healthy work‑life balance.
  • A wide variety of office amenities including nutritious snacks and refreshing drinks, free gym and sauna access to keep you fueled and healthy.
  • Flexible working hours and a remote work policy that includes reimbursement of connectivity costs and equipment to work from home.
  • Sports card fully financed by Ampere.

#LI-Hybrid

Ampere is an inclusive and equal opportunity employer and welcomes applicants from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, national origin, citizenship, religion, age, veteran and/or military status, sex, sexual orientation, gender, gender identity, gender expression, physical or mental disability, or any other basis protected by federal, state or local law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Accelerator Software Senior Principal Engineer- Framework Integration
AI Accelerator Software Senior Principal Engineer- Framework Integration

Ampere • Warszawa

Hybrid
PLN 416,000 - 625,000
Premium medical health care
Generous paid time off
Office amenities
+2
AI Accelerator Software Principal Engineer- Framework Integration
AI Accelerator Software Principal Engineer- Framework Integration

Ampere • Warszawa

On-site
PLN 302,000 - 453,000
Premium medical health care
Generous paid time off
Office amenities and snacks
+1
AI Accelerator Software Principal Engineer – Runtime Library
AI Accelerator Software Principal Engineer – Runtime Library

Ampere • Warszawa

Hybrid
PLN 301,000 - 453,000
Premium medical health care
Generous paid time off
Office amenities: snacks, gym access
+2
AI Accelerator Software Principal Engineer – Runtime Library
AI Accelerator Software Principal Engineer – Runtime Library

Ampere Computing • Warszawa

On-site
PLN 302,000 - 453,000
Premium medical care
Generous PTO
Gym and sauna access
+3
Remote AI Accelerator Framework Engineer - Principal
Remote AI Accelerator Framework Engineer - Principal

Ampere • Warszawa

On-site
PLN 302,000 - 453,000
Premium medical health care
Generous paid time off
Office amenities and snacks
+1
Remote AI Accelerator Software Lead — Framework & Performance
Remote AI Accelerator Software Lead — Framework & Performance

Ampere • Warszawa

Hybrid
PLN 416,000 - 625,000
Premium medical health care
Generous paid time off
Office amenities
+2
Senior AI Accelerator Software Architect (Framework & Perf) - Remote
Senior AI Accelerator Software Architect (Framework & Perf) - Remote

Ampere Computing • Warszawa

Hybrid
PLN 417,000 - 625,000
Premium health care
Paid time off
Office amenities
+2
Senior C++ Software Engineer with CUDA/GPU/TPU
Senior C++ Software Engineer with CUDA/GPU/TPU

EPAM Systems • Poland

Hybrid
PLN 240,000 - 360,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+4
Lead C++ Developer
Lead C++ Developer

EPAM Systems • Poland

Hybrid
PLN 320,000 - 520,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+4
Software Engineer Poland
Software Engineer Poland

amp • Warszawa

On-site