Research Engineer, AI Systems

CV in

United Kingdom

Hybrid

GBP 88,000 - 184,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lightning AI is seeking a Research Engineer to enhance AI systems across its platform, turning research ideas into robust, scalable production components. You will own experiments validating architectural choices for large workloads and diagnose inefficiencies to propose concrete improvements that boost speed, stability, and resource utilization.

The role bridges cutting-edge research and high-availability infrastructure, with a focus on delivering production-grade solutions and open-source

Qualifications

  • Bachelor's degree in CS, Engineering, or related field.
  • Strong PyTorch experience and DL workflows.
  • Experience with large-scale training or inference in production.
  • Strong distributed systems knowledge.
  • Ability to deliver production-grade code and APIs.

Responsibilities

  • Oversee optimization of large-scale training and inference performance across distributed systems.
  • Conduct deep analysis with customers to identify bottlenecks and design resilient production AI systems.
  • Architect and maintain high-throughput inference pipelines and model serving systems.
  • Develop observability, profiling, and debugging utilities for model execution and system behavior.
  • Integrate performance improvements into the Lightning ecosystem via stable APIs and workflows.
  • Collaborate with hardware partners (NVIDIA, TPU, etc.) to optimize execution strategies.
  • Contribute open-source efforts and high-quality technical documentation.

Skills

Deep learning (PyTorch)
Distributed systems
Software engineering

Education

Bachelor's degree in Computer Science or Engineering

Tools

CUDA
Triton
TensorRT
vLLM
SGLang
Dynamo

Job description

Research Engineer at Lightning AI.

About the role Lightning AI is seeking a Research Engineer to improve the performance, reliability, and scalability of AI systems on our platform. You will own the design and execution of experiments that validate architectural choices for large-scale AI workloads. This role requires you to translate novel research ideas into robust, production-grade components that operate reliably at scale. You will be responsible for diagnosing systemic inefficiencies and proposing concrete engineering solutions that impact the entire product stack. The ideal candidate thrives on turning ambiguous problems into measurable improvements in speed, stability, and resource utilization. You will act as a technical bridge between cutting-edge research prototypes and the high-availability infrastructure serving enterprise customers. Your work will directly influence how the Lightning platform scales to meet the demands of the most challenging AI workloads.

Key facts

Location: London, New York, San Francisco, Seattle (Hybrid: 2 days/week in-office) Engagement: Full-time Salary: $120,000 to $250,000 USD base salary

What you'll do
  • Oversee the optimization of large-scale training and inference performance across distributed systems and heterogeneous accelerators.
  • Conduct deep analysis alongside customers to pinpoint workload bottlenecks and architect resilient production AI systems.
  • Architect, build, and maintain high-throughput inference pipelines, scalable model serving systems, and performance-centric engineering tools.
  • Design and deliver observability, profiling, and debugging utilities that provide deep insight into model execution and system behavior.
  • Integrate performance enhancements directly into the Lightning ecosystem through clean, maintainable APIs and automated workflows.
  • Work closely with hardware partners from NVIDIA, TPU, and other accelerator ecosystems to co-develop efficient execution strategies.
  • Drive the creation of open-source contributions and high-quality technical documentation that clarify system capabilities and limitations.
  • Evaluate and prototype inference optimization methods such as quantization, speculative decoding, mixed precision, and memory-efficient training techniques.
  • Implement low-level kernels and integrations using CUDA, Triton, TensorRT, vLLM, SGLang, or Dynamo to unlock maximum hardware throughput.
  • Champion best practices in software engineering by focusing on API consistency, debuggability, and reliable production code delivery.
Requirements
  • Demonstrate advanced proficiency with deep learning frameworks, with a core focus on PyTorch and its ecosystem.
  • Bring hands-on experience managing large-scale training or inference workloads in demanding production environments.
  • Show a firm grasp of distributed systems concepts, including data, model, and pipeline parallelism, as well as elastic scaling and checkpointing strategies.
  • Exhibit strong software engineering capabilities in API design, debugging complex systems, and delivering production-grade code.
  • Prove a track record of identifying and resolving performance bottlenecks within machine learning infrastructure stacks.
  • Hold a Bachelor degree in Computer Science, Engineering, or a closely related technical field.
  • Display the ability to operate effectively in a fast-paced, cross-functional environment with shifting priorities and tight deadlines.
  • Commit to working in a hybrid model with a minimum of two days in the office per week across London, New York, San Francisco, or Seattle locations.
  • Meet compensation bands that align with local market standards and internal equity considerations for the listed regions.
  • Adhere to company policies regarding employment eligibility and legal work authorization in your country of residence.
Nice to have
  • Hands-on familiarity with inference optimization methods like quantization, speculative decoding, mixed precision, or memory-efficient training.
  • Direct experience with performance-critical frameworks and tools such as CUDA, Triton, TensorRT, vLLM, SGLang, or Dynamo.
  • A history of meaningful contributions to open-source machine learning or infrastructure projects with public repositories.
  • Previous experience within startups or highly collaborative technical environments where impact was measured by execution speed.
  • Possession of an advanced degree, such as a Master or PhD, in artificial intelligence, machine learning, or computer systems.
Practical notes

This is a full-time position with total compensation that includes a discretionary bonus and equity components. The salary range specified for this role is $120,000 to $250,000 USD base salary, adjusted for geographic location and individual qualifications. Benefits coverage includes medical, dental, and vision insurance for eligible employees. Retirement savings options include 401(k) matching for United States-based staff or pension contributions for United Kingdom-based staff. You will enjoy unlimited PTO, a two-week winter break, paid parental leave, and a four-week paid sabbatical after four years of service. Additional support is provided through in-office meals, a professional development allowance, wellness stipends, and WFH equipment stipends. Travel requirements are minimal, and the role operates under a hybrid schedule that requires two days of in-office presence per week. Employment is contingent on successful completion of any applicable visa authorization or work eligibility checks where required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Research Engineer, LLM Training & Post-Training
Senior Research Engineer, LLM Training & Post-Training

Lightning-Ai • Greater London

Hybrid
GBP 122,000 - 229,000
Health coverage
Equity
401(k) matching
+8
Senior Software Engineer, Core Platform
Senior Software Engineer, Core Platform

Lightning AI • Greater London

Hybrid
GBP 134,000 - 186,000
Health coverage
Equity
Retirement benefits
+3
Senior Software Engineer, Agents
Senior Software Engineer, Agents

Lightning-Ai • Greater London

Hybrid
GBP 134,000 - 186,000
Health coverage
Equity
Retirement plan
+7
Frontend Engineer
Frontend Engineer

Jackalope Digital LLC • Greater London

Hybrid
GBP 89,000 - 184,000
Comprehensive Health Coverage
Meaningful Equity
401(k) matching (US) and pension (UK)
+2
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • United Kingdom

On-site
GBP 140,000 - 200,000
Senior Software Engineer, Training & Experimentation
Senior Software Engineer, Training & Experimentation

United States Digital Space LLC • Greater London

Hybrid
GBP 90,000 - 140,000
Comprehensive health coverage
Meaningful equity (RSUs)
401(k) matching (US) and UK pension
+2
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
ML Research Scientist - Member of Technical Staff
ML Research Scientist - Member of Technical Staff

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Competitive salary
Equity ownership
Private healthcare
+2
Performance Engineer
Performance Engineer

CommonAI CIC • Cambridge

On-site
GBP 65,000 - 90,000
Collaborative environment
High impact in growing org
Competitive salary and pension
+3
Principal Machine Learning Infrastructure Engineer London, United Kingdom
Principal Machine Learning Infrastructure Engineer London, United Kingdom

PhysicsX Ltd • Greater London

On-site
GBP 80,000 - 100,000
Equity options
10% employer pension contribution
Free office lunches
+6