AI Engineer / LLM Systems Engineer

Sphere Software

United States

Remote

USD 120,000 - 170,000

Full time

35 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Sphere Software seeks an AI Engineer to lead AI-first product development and optimize production-grade AI systems. You will own complex initiatives, work with senior stakeholders, and drive LLM inference infrastructure from concept to production.

The role emphasizes hands-on experience with LLM engineering, RAG pipelines, vector search, and multi-agentic workflows, plus scalable AI architecture and performance tuning.

Qualifications

  • Proven experience building production AI/LLM systems.
  • Hands-on with LLM engineering, RAG pipelines, and multi-agent workflows.
  • Ability to design scalable AI system architectures.

Responsibilities

  • Lead development of AI-first features and products under deadlines.
  • Implement AI-first solutions with partners and clients.
  • Design architecture for AI-powered products and LLM apps.
  • Validate and improve AI products from feedback and reqs.
  • Define technical scope with execs and stakeholders.
  • Benchmark and optimize LLM inference performance.
  • Tune inference for high throughput and low latency.
  • Monitor performance metrics and identify bottlenecks.

Skills

LLM engineering
RAG pipelines
Vector search
Multi-agentic workflows
Production AI systems
System design
Performance tuning
Inference optimization
Communication with stakeholders

Tools

vLLM
TensorRT-LLM
Triton Inference Server

Job description

We are looking for an AI Engineer to lead the development of AI-first products and solutions. This role combines advanced LLM engineering, system architecture, and high-performance inference optimization.

The ideal candidate has strong software engineering fundamentals and hands-on experience building production-grade AI systems, including advanced RAG pipelines, multi-agentic workflows, and LLM inference infrastructure. This person will take ownership of complex technical initiatives, work directly with senior stakeholders and strategic partners, and drive AI solutions from concept to production.

Responsibilities
  • Lead the development and delivery of high-priority AI-first features and products, often working under tight deadlines.
  • Drive technical implementation of AI-first solutions in collaboration with major partners and clients.
  • Design system architecture for AI-powered products, including advanced LLM applications, RAG pipelines, and multi-agentic workflows.
  • Validate and improve existing AI products based on user feedback and evolving business requirements.
  • Define and manage technical scope for AI initiatives in collaboration with C-level executives, marketing, directors, and other senior stakeholders.
  • Validate, benchmark, and optimize LLM inference infrastructure.
  • Conduct load and performance testing, identify bottlenecks, and implement inference optimization improvements.
  • Optimize model serving performance through batching, quantization, and inference engine configuration.
  • Monitor and analyze inference performance using metrics such as TTFT, TPOT, and ITL.
Requirements
  • Strong software engineering background with substantial hands-on experience building production AI/LLM systems.
  • Practical experience with LLM engineering, RAG, vector search, and multi-agentic workflows.
  • Experience designing and implementing scalable AI system architectures.
  • Experience with LLM inference frameworks such as vLLM, TensorRT-LLM, and/or Triton Inference Server.
  • Understanding of LLM inference optimization, including batching, quantization, and performance tuning.
  • Experience benchmarking and troubleshooting LLM inference performance.
  • Understanding of inference performance metrics, including TTFT, TPOT, and ITL.
  • Strong system design and problem-solving skills.
  • Ability to independently drive complex technical initiatives and communicate effectively with senior stakeholders.
Nice to Have
  • Experience with Golang and Bash scripting.
  • Experience with GPU-based deployments and cloud infrastructure.
  • Experience working with self-hosted and open-weight models.
  • Experience optimizing inference infrastructure for high-throughput and low-latency workloads.
  • Previous experience leading AI-first product development or strategic technical collaborations.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Systems Architect & LLM Engineer
AI Systems Architect & LLM Engineer

Sphere Software • United States

Remote
USD 120,000 - 170,000
AI Engineer
AI Engineer

Compunnel, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Engineer
Senior AI Engineer

Empathy Talent • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Engineer/Architect
AI Engineer/Architect

Empiric • Indianapolis (IN)

On-site
USD 100,000 - 150,000
AI Engineer
AI Engineer

LeoTechnologies • Boca Raton (FL), Northern (KY)

On-site
USD 140,000 - 190,000
Senior AI Engineer: Scalable LLMs & MLOps Leader
Senior AI Engineer: Scalable LLMs & MLOps Leader

Compunnel, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Engineer - FL
AI Engineer - FL

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
AI Engineer (Applied LLM Systems)
AI Engineer (Applied LLM Systems)

Discernis • New York (NY)

On-site
USD 140,000 - 200,000
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Ai Software Architect Llm Agentic Systems Msys Tech India Pvt Ltd Chennai
Ai Software Architect Llm Agentic Systems Msys Tech India Pvt Ltd Chennai

Vibehackers • United States

On-site
USD 180,000 - 260,000