Senior AI Engineer (Xora Portfolio Company)

Xora Innovation

San Diego (CA)

Hybrid

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

ELEMYNT, built by Xora Innovation, seeks a seasoned engineer to own LLM-powered systems across production environments. You will design agents, implement tooling, and ensure safe, scalable, and observable services deployed in customer environments.

You will build end-to-end pipelines for model ingestion, evaluation, and deployment, including fine-tuning, retrieval, and robust monitoring in hybrid or on-site settings.

Qualifications

  • Bachelor's or Master's in Computer Science or related engineering field and 5+ years shipping production software.
  • Strong Python and reliability in async, HTTP and streaming APIs.
  • Production experience with prompting, context engineering, tool calling, and structured outputs.
  • Hands-on design and shipping of agents: loop, tools, memory, failure handling; LangGraph or equivalent.
  • Experience building RAG systems: embeddings, chunking, hybrid search, reranking.
  • Experience fine-tuning open-weight models (LoRA/QLoRA) on multi-GPU with curated data.
  • Experience with LLM evaluation, guardrails, tracing, and production scoring.

Responsibilities

  • Build and ship LLM-powered capabilities end to end from prototype to production.
  • Design agents that plan multi-step work with tool calls and structured outputs.
  • Create retrieval with embeddings, chunking, and hybrid search for context.
  • Fine-tune open-weight models with multi-GPU setups and curated data.
  • Develop evaluation loops with automated scoring and regression tracking.
  • Instrument model calls and tooling with tracing for cost and reliability.
  • Turn LLM capabilities into clean APIs and reusable engineering tooling.

Skills

Python
LLM systems in production
Async APIs
HTTP API design
Testing & code review
LangGraph
JSON Schema / Pydantic
RAG systems
Model fine-tuning (LoRA/QLoRA)
Tracing / observability

Education

Bachelor's or Master's in Computer Science or related engineering

Tools

LangGraph
Pydantic / JSON Schema
vLLM
TGI
SGLang
RAG tooling

Job description

About Elemynt

ELEMYNT is an early-stage startup built by Xora Innovation. We develop applied intelligence that brings AI into the real world. Our platform combines advanced machine learning, high-performance simulation, and modern software engineering to accelerate the design, validation, and deployment of new materials. Our work sits at the intersection of AI, physics, and large-scale computation. The problems are hard, the stakes are high, and the impact is tangible.

About Elemynt

ELEMYNT is an early-stage startup built by Xora Innovation. We develop applied intelligence that brings AI into the real world. Our platform combines advanced machine learning, high-performance simulation, and modern software engineering to accelerate the design, validation, and deployment of new materials. Our work sits at the intersection of AI, physics, and large-scale computation. The problems are hard, the stakes are high, and the impact is tangible.

About The Role

This role owns the LLM systems behind our platform: the agents and fine-tuned models that ship as product, and the engineering that keeps them reliable – evaluation, tracing, and production-quality services. It's deeply hands-on, from model internals to shipped software.

The platform runs inside our customers’ own secure environments: their compute, their cloud, or a hybrid. So the LLM layer has to work with commercial APIs and self-hosted models alike, and carry its own safeguards wherever it lands. Every LLM capability we ship stands on this work.

What You Will Do
  • Build and ship LLM-powered capabilities end to end: prototype, evaluate, deploy, and iterate them into production services users rely on.
  • Design agents that plan and carry out multi-step work: tool calling, structured outputs, durable state, and the judgment to know when an agent is the wrong tool.
  • Build retrieval that gives models the right context: ingestion, chunking, embeddings, hybrid search, reranking.
  • Fine-tune open-weight models with LoRA, QLoRA, or full-parameter tuning on multi-GPU, curating the training data and choosing the method by task, compute budget, and target.
  • Build evaluation loops that gate what ships: automated scoring, LLM-as-judge, and regression tracking against curated test sets.
  • Instrument model calls and tool use with tracing, so quality, cost, and failures stay debuggable in production.
  • Turn LLM capabilities into clean APIs and reusable tooling that other engineers build on.
What We Are Looking For
  • Bachelor's or Master's degree in Computer Science or a related engineering field, and 5+ years building and shipping production software, including deep hands-on work building LLM-powered systems in production.
  • Strong Python and a track record of shipping reliable services: async, HTTP and streaming APIs, testing, code review.
  • Production experience with LLMs: prompting and context engineering, tool calling, structured output, and the latency and cost work that keeps them usable.
  • Hands-on experience designing and shipping agents: the loop, the tools, context, memory, and where they fail. A framework such as LangGraph or equivalent; structured outputs in Pydantic or JSON Schema.
  • Experience building RAG systems: embeddings, chunking, hybrid search, reranking, and a feel for what actually moves retrieval quality.
  • Direct experience fine-tuning open-weight models (LoRA, QLoRA, or full-parameter) on multi-GPU, including curating and formatting the training data.
  • Experience with LLM evaluation and guardrails: LLM-as-judge or automated scoring, regression tracking, and tracing over agent runs.
  • Experience building shared LLM tooling or platform components that other engineers build on, and comfort owning ambiguous systems end to end in an early-stage environment.
NICE TO HAVE
  • Self-hosted inference with vLLM, TGI, or SGLang, served behind an OpenAI-compatible interface.
  • Interoperability standards for tools and agents, such as MCP.
  • Retrieval over structured data: knowledge graphs, hybrid search, reranking at scale.
  • LLMs applied to scientific or other technical data; experience making APIs and tool surfaces easy for agents to call reliably.
  • Contributions to open-source AI/ML: agent frameworks, eval tooling, RAG, fine-tuned models.
LOCATION

Singapore or United States. We’re hiring in both to reach the right person. Work model is on-site or hybrid, set per location.

CLOSING NOTE

If you don't tick every box but this is clearly your kind of work, get in touch.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agentic AI Engineer (Xora Portfolio Company)
Agentic AI Engineer (Xora Portfolio Company)

Xora Innovation • San Diego (CA)

Hybrid
USD 150,000 - 190,000
Senior Data & ML Infrastructure Engineer (Xora Portfolio Company)
Senior Data & ML Infrastructure Engineer (Xora Portfolio Company)

Xora Innovation • San Diego (CA)

Hybrid
USD 160,000 - 260,000
On-site or hybrid work model
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • United States

On-site
USD 180,000 - 280,000
Application Engineer – LLM
Application Engineer – LLM

Salt Digital Recruitment • United States

On-site
USD 120,000 - 180,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallelwireless • United States

On-site
USD 140,000 - 210,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

Remote
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

Hybrid
USD 140,000 - 190,000
AI Engineer
AI Engineer

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
US$1,000 per month food and commuting allowance
Laptop of choice
ESOP available
ML/AI Research Engineer — Agentic AI Lab (Founding Team)
ML/AI Research Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity