Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation

Ambient.ai

San Francisco (CA)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options
Comprehensive benefits
Flexible time off

Job summary

Ambient.ai seeks an experienced AI Infrastructure Engineer to design and scale real‑time AI infrastructure for multimodal models and video data. You will partner with researchers and product teams to deploy cutting-edge ML systems in production.

The role emphasizes LLM/LVM inference, model evaluation, and performance tuning, with a focus on reliability and cost efficiency in an enterprise setting.

Qualifications

  • BS/MS in CS or related field or equivalent practical experience.
  • 4+ years building infrastructure, distributed systems, ML platforms, or production AI systems.
  • Strong Python programming and software engineering fundamentals.

Responsibilities

  • Design, build, and maintain AI infrastructure for real-time computer vision, LLM, LVM, and multimodal workloads.
  • Scale systems to process terabytes of video and sensor data in real time.
  • Optimize inference latency, throughput, GPU utilization, reliability, and cost.
  • Develop robust evaluation harnesses and benchmarking for model quality and system performance.
  • Build infrastructure for continuous model evaluation, experimentation, and deployment.
  • Collaborate with researchers to productionize latest AI advances.

Skills

Python
Machine learning
Distributed systems
GPU utilization
Cloud infrastructure
Model serving
LLM/LVM inference
Observability

Education

BS/MS in Computer Science or related field

Tools

vLLM
Triton Inference Server
CUDA
PyTorch
TensorRT
ONNX

Job description

Build a safer world with us, one incident at a time. Ambient.ai is the category creator and leader in Agentic Physical Security. Powered by Ambient Pulsar, the first reasoning Vision‑Language Model purpose‑built for physical security, our platform seamlessly integrates with existing security cameras and physical access control systems to unify monitoring, access control, threat assessment, response, and investigations through an always‑on reasoning layer that augments security operators with superhuman capabilities. The results: 95% fewer false alarms, investigations 20× faster, and 10× faster response. The momentum speaks for itself: we doubled new ARR in FY26 and have delivered results for world‑class customers including Cisco, ServiceNow, SentinelOne, TikTok, Bayer, and MoMA. Founded in 2017 and backed by Andreessen Horowitz, Y Combinator, and Allegion Ventures, Ambient.ai is on a fast‑paced journey to fulfill our mission: prevent every security incident possible.

About The Role

Reporting to Raghu Nallamothu, you will design, build, and optimise the AI infrastructure that powers Ambient.ai’s real‑time intelligence platform. In this role, you will work on the systems required to run state‑of‑the‑art deep learning models across many terabytes of video data in real time. You will help build and scale infrastructure for inference, evaluation, and continuous model improvement across computer vision models, large language models, large vision models, and multimodal AI systems. This role is ideal for someone with a strong blend of infrastructure engineering, production ML systems, LLM/LVM inference, evaluation harnesses, and inference optimisation experience. You will partner closely with research scientists and product engineering teams to bring the latest AI advancements into production for our customers.

What You'll Do
  • Design, build, and maintain cutting‑edge AI infrastructure for real‑time computer vision, LLM, LVM, and multimodal inference workloads.
  • Build scalable systems for running state‑of‑the‑art models across large volumes of video and sensor data.
  • Optimise inference performance across latency, throughput, GPU utilisation, reliability, and cost.
  • Develop robust evaluation harnesses and benchmarking systems to measure model quality, system performance, regressions, and production readiness.
  • Build infrastructure for continuous model evaluation, experimentation, and deployment.
  • Partner with research scientists to productionise the latest advances in computer vision, LLMs, LVMs, RAG, and multimodal AI.
  • Improve model‑serving architecture, including batching, caching, routing, quantisation, model parallelism, and hardware utilisation.
  • Develop data engines and feedback loops for collecting training data, evaluating model behaviour, and continuously improving AI performance.
  • Create reliable observability, monitoring, and debugging tools for production AI systems.
  • Help define best practices for deploying, evaluating, and operating AI systems in real‑world enterprise environments.
What You'll Bring
  • 4+ years of industry experience building infrastructure, distributed systems, machine learning platforms, or production AI systems.
  • BS/MS in Computer Science or a related technical field, or equivalent practical experience.
  • Strong programming background, especially in Python, with solid software engineering fundamentals.
  • Experience designing and building scalable machine learning infrastructure for training, inference, evaluation, and deployment.
  • Hands‑on experience running deep learning models in production, ideally including LLMs, LVMs, vision‑language models, or multimodal models.
  • Strong understanding of inference optimisation techniques, including batching, caching, quantisation, parallelism, memory optimisation, GPU utilisation, and latency reduction.
  • Experience with model‑serving frameworks or systems such as vLLM, Triton Inference Server or similar technologies.
  • Experience building evaluation frameworks, test harnesses, benchmarks, regression tests, or model‑quality measurement systems.
  • Strong background in machine learning and deep learning; computer vision experience is a strong plus.
  • Experience designing data engines or pipelines for collecting, managing, and curating training and evaluation data.
  • Familiarity with integrating advanced AI systems such as LLMs, LVMs, RAG pipelines, embedding models, or multimodal models into production applications.
  • Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU‑based workloads.
  • Strong collaboration and communication skills, with the ability to work effectively with research scientists, product teams, infrastructure teams, and stakeholders.
  • Proactive problem‑solving ability, a strong ownership mindset, and adaptability to incorporate new AI technologies and methodologies.
Nice to Have
  • Experience operating large‑scale GPU infrastructure or distributed inference systems.
  • Experience with CUDA, NCCL, PyTorch, TensorRT, ONNX, or similar ML systems technologies.
  • Experience with video understanding, real‑time computer vision, multimodal AI, or physical‑world AI systems.
  • Experience with model compression, speculative decoding, distillation, pruning, or low‑latency serving techniques.
  • Experience with prompt evaluation, model regression testing, human‑in‑the‑loop evaluation, or automated quality gates.
  • Familiarity with retrieval‑augmented generation, vector databases, embedding models, re‑rankers, or search infrastructure.
  • Experience building internal ML platforms or tools used by researchers and applied ML teams.
What Success Looks Like

You will be successful in this role if you can build practical, scalable infrastructure that helps Ambient.ai deploy better AI models faster and more reliably. You should be comfortable working across the full stack of production AI systems, from model behaviour and evaluation to serving architecture, GPU performance, observability, and customer‑facing reliability. This is a hands‑on engineering role for someone excited to help bring the next generation of AI, computer vision, LLMs, and LVMs into real‑world production environments.

Why Join Us
  • We are creating an entirely new category within a 180+ billion‑dollar physical security industry and are passionate about our mission to prevent every security incident possible.
  • We partner with an incredible customer roster of F500 companies, including Adobe, TikTok, Gap and SentinelOne.
  • Regular full‑time employees receive stock options for the opportunity to share ownership in the success of our company.
  • Comprehensive health and welfare package (Medical, Dental, Vision, Life, EAP, Legal Services, 401(k) plan).
  • We offer flexible time off to rest and recharge, including Winter Break.
  • The latest tech and awesome swag will be delivered to your door.
  • Enjoy a full range of opportunities to connect with your awesome co‑workers.
  • We love to hike, are foodies, and love music! Check out our most recent Ambient Spotify playlist.

We’ve found that in‑person time meaningfully supports collaboration, creativity, and team alignment. Our talent, engineering, product, design, and marketing teams work from our Redwood City office three days a week. All other Bay Area employees join on Fridays to stay connected and close out the week together.

Ready to learn more? Connect with us on LinkedIn or YouTube.

Ambient.ai is proud to be an Equal Opportunity Employer. Ambient does not unlawfully discriminate on the basis of race, color, religion, sex (including pregnancy, childbirth, breastfeeding, or related medical conditions), gender identity, gender expression, national origin, ancestry, citizenship, age, physical or mental disability, legally protected medical condition, family care status, military or veteran status, marital status, registered domestic partner status, sexual orientation, genetic information, or any other basis protected by local, state, or federal laws. Ambient is an E‑Verify participant.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation

Ambient AI, Inc. • Redwood City (CA)

Hybrid
USD 190,000 - 270,000
Stock options
Health, dental, vision
401(k)
+1
Senior Software Engineer, AI Data Systems & Database Infrastructure
Senior Software Engineer, AI Data Systems & Database Infrastructure

Ambient.ai • San Francisco (CA)

On-site
USD 180,000 - 240,000
Stock options
Software Engineer - Backend (Product)
Software Engineer - Backend (Product)

Ambient.ai • Redwood City (CA)

On-site
USD 168,000 - 205,000
Senior Sales Engineer
Senior Sales Engineer

Ambient.ai • San Francisco (CA)

On-site
USD 140,000 - 200,000
Stock options
Health + welfare package
Winter break
Senior Sales Engineer
Senior Sales Engineer

Ambient.ai • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Stock options
Health & welfare package
Winter Break off
+1
Senior Software Engineer, AI Data Systems & Database Infrastructure
Senior Software Engineer, AI Data Systems & Database Infrastructure

Socket.dev • Redwood City (CA)

On-site
USD 190,000 - 270,000
Stock options
Health benefits
Flexible time off
+1
Sr. Software Engineer, Fullstack
Sr. Software Engineer, Fullstack

Ambient.ai • San Francisco (CA)

On-site
USD 100,000 - 140,000
Stock options
Flexible time off
Latest tech and swag
Senior Sales Engineer
Senior Sales Engineer

Ambient • Redwood City (CA), Northern (KY)

On-site
USD 140,000 - 190,000
Stock options
Comprehensive health plan
Flexible time off
Senior Software Engineer, AI Data Systems & Database Infrastructure
Senior Software Engineer, AI Data Systems & Database Infrastructure

Ambient AI, Inc. • Redwood City (CA)

On-site
USD 180,000 - 260,000
Stock options
Comprehensive health & welfare
Flexible time off
+1
Regional Sales Manager, Strategic Accounts
Regional Sales Manager, Strategic Accounts

Ambient • Virginia (IL), Northern (KY)

Hybrid
USD 120,000 - 170,000
Stock options
Health benefits
Flexible time off
+1