Software Engineer, AI Systems (Haven Safety)

AI Fund

United States

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Haven Safety AI is building the enterprise learning intelligence layer for safety and is seeking an experienced AI/LLM engineer to design, ship, and operate production-grade reasoning systems that connect evidence bases to product experiences used by investigators and safety leaders.

In this build-and-operate role reporting to the CTO, you will own multi-step LLM pipelines, agent orchestration, grounding of evidence, evaluation pipelines, and production AI operations in a regulated enterprise

Qualifications

  • Production LLM systems shipped in production environments.
  • Agentic workflows: building multi-step, tool-calling workflows.
  • Evaluation discipline with datasets and regression testing.
  • Retrieval judgment for context selection.
  • Production ownership from deployment to monitoring and incident response.
  • Model judgment across providers and tradeoffs in quality, latency and cost.
  • Graph reasoning with Neo4j and Cypher.
  • Strong Python production habits (FastAPI, async, testing).
  • Deep graph experience with embeddings and knowledge graphs.
  • Enterprise AI security and policy controls.
  • Azure and hybrid search experience.

Responsibilities

  • Production reasoning systems: Build and operate multi-step LLM pipelines coordinating model calls, tool calls, graph queries, retrieval, quality gates, and agent handoffs.
  • Agent orchestration: Extend Haven’s coordinated agent team and orchestration layer for incident progression from evidence to enterprise learning.
  • Grounding and retrieval: Design the context layer across Neo4j traversal, vector search, and hybrid retrieval for accurate evidence grounding.
  • Evaluation: Build datasets, scoring, regression suites, model comparisons, human-label loops, and per-stage quality attribution.
  • Production AI operations: Implement tracing, tool-call audits, cost and latency monitoring, failure handling, and quality dashboards; catch loops and drift.
  • Technical direction: Select models by task across OpenAI, Anthropic, and Google; shape AI roadmap.

Skills

Production LLM systems
Agentic workflows
Evaluation discipline
Retrieval judgment
Production ownership
Model judgment
Graph reasoning
Strong Python
Deep graph experience
Enterprise AI security
Azure and hybrid search

Tools

Neo4j
Cypher
LangGraph
LangChain
Azure AI Search
Pinecone
MongoDB Atlas
pgvector
Elasticsearch

Job description

Haven Safety AI is building the enterprise learning intelligence layer for safety. Co-founded with The AES Corporation and AI Fund, the venture studio founded by Andrew Ng, Haven helps high-risk organizations learn faster from what goes wrong so they can prevent what comes next.

Haven works alongside existing EHS enterprise systems to improve how organizations investigate, assess, and learn from incidents. INVESTIGATE guides evidence synthesis, timeline development, multi-threaded causal analysis, and corrective actions. ASSURE continuously reviews completed investigations for evidence quality, causal coverage, guideline adherence, and CAPA strength. LEARN reasons across incident history to surface recurring control failures, repeated corrective-action patterns, CAPA debt, and emerging signals.

The platform combines current incident evidence, company knowledge, historical cases, and an industry knowledge graph. A coordinated team of specialized AI agents examines evidence, controls, engineering factors, procedures, regulations, training, and organizational history, then produces one traceable assessment for human review. Customers have reported an 80% reduction in root cause analysis labor time using Haven.

www.havensafety.com

About the role:

You will work on the AI and LLM engineering layer that connects Haven’s evidence base and knowledge graph to the product experiences investigators and safety leaders use. This is a build-and-operate role reporting to the CTO. You will design the reasoning, ship it, instrument it, and improve it using production evidence.

The work is technically demanding and operationally consequential. Haven serves enterprise customers in regulated, safety-critical industries through multi-tenant and dedicated deployments. A plausible answer and a correct answer can look the same until someone acts on it, so precision, traceability, evaluation, and human oversight are core product requirements.

Haven Safety is a VC-backed pre-seed venture. This role will be a important member of the founding team and will require wearing multiple hats. This role is based in the US and relocation will not be considered.

What you will own:
  • Production reasoning systems. Build and operate multi-step LLM pipelines that coordinate model calls, tool calls, graph queries, retrieval, quality gates, and specialist-agent handoffs.
  • Agent orchestration. Extend Haven’s coordinated agent team and the orchestration layer that carries an incident from evidence through analysis, review, and enterprise learning.
  • Grounding and retrieval. Design the context layer across Neo4j graph traversal, vector search, and hybrid retrieval so every model call receives the right evidence and organizational knowledge.
  • Evaluation. Build datasets, scoring, regression suites, model comparisons, human-label loops, and per-stage quality attribution for extraction and reasoning tasks.
  • Production AI operations. Implement tracing, tool-call audits, cost and latency monitoring, failure handling, and quality dashboards; catch loops, hallucinations, and silent drift before customers do.
  • Technical direction. Select models by task across OpenAI, Anthropic, and Google; partner with product and knowledge engineering; and help shape the AI roadmap.
What we are looking for:

We care more about depth, judgment, and ownership than a skills checklist. Prior safety experience is not required, but experience operating in an enterprise-level environment is strongly preferred. If you meet most of these requirements and are excited by the problem, we would like to hear from you.

Core experience:
  • Production LLM systems. AI or ML engineering, including shipping LLM systems that real users depend on.
  • Agentic workflows. Hands-on experience building and debugging multi-step, tool-calling workflows with LangGraph, LangChain, or an equivalent framework.
  • Evaluation discipline. A repeatable approach to LLM evaluation, including representative datasets, regression testing, LLM-as-judge techniques, or human review loops.
  • Retrieval judgment. Experience assembling context for LLMs and a clear point of view on what to retrieve, how much, and why.
  • Production ownership. A track record of owning systems from deployment through monitoring and incident response, including a strong story about a failure or regression you diagnosed and fixed.
  • Model judgment. Comfort working across model providers and explaining tradeoffs in quality, latency, cost, context, and operational risk.
  • Graph reasoning. Comfort with Neo4j and Cypher, or a comparable graph store, and the ability to ramp quickly on graph data modeling. Strongly preferred.
  • Strong Python. Production habits around FastAPI, asynchronous services, testing, observability, and maintainable interfaces are required.
  • Deep graph experience. Cypher fluency, schema evolution, MERGE patterns, embeddings, and operating a live knowledge graph.
  • Enterprise AI security. Prompt-injection awareness, context-leak prevention, tenant isolation, role-based access, and policy-layer separation.
  • Azure and hybrid search. Experience running production AI services in Azure and using Azure AI Search, Pinecone, MongoDB Atlas, pgvector, Elasticsearch, or a similar platform.
Why join Haven:
  • Meaningful reasoning problems. Work across evidence, causal pathways, controls, organizational history, and corrective actions in a domain where correctness matters.
  • An evaluation-first culture. Make quality measurable, observable, and improvable instead of relying on demos or intuition.
  • Visible customer impact. Build for safety teams in energy, utilities, infrastructure, construction, and manufacturing, with direct feedback from the people using the output.
  • Small team, high ownership. Work closely with the CTO, product, and knowledge engineering, make consequential technical decisions, and see your work reach customers quickly.

Sponsorship will not be provided for this role.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, AI Systems (United States)
Software Engineer, AI Systems (United States)

Lever, Inc. • United States

On-site
USD 120,000 - 180,000
Software Engineer, AI Systems (United States)
Software Engineer, AI Systems (United States)

Aifund • United States

On-site
USD 180,000 - 300,000
AI Systems Engineer for Enterprise Safety Platform
AI Systems Engineer for Enterprise Safety Platform

Aifund • United States

On-site
USD 180,000 - 300,000
AI Engineer
AI Engineer

Pinpoint Global Communications • United States

On-site
USD 120,000 - 180,000
Enterprise Account Executive (AI for Industrial Safety)
Enterprise Account Executive (AI for Industrial Safety)

AI Fund • United States

Hybrid
USD 120,000 - 170,000
Competitive compensation
Uncapped commissions
Equity participation
+2
Founding Member of Technical Staff
Founding Member of Technical Staff

Clera • United States

On-site
USD 150,000 - 200,000
Founding team equity
Early-stage upside
AI Engineer
AI Engineer

Valsoft Corporation • United States

On-site
USD 140,000 - 230,000
AI Solutions Engineer
AI Solutions Engineer

Havenparkcommunities • Orem (UT)

On-site
USD 120,000 - 180,000
AI Engineer
AI Engineer

Valsoft Corporation • Northern (KY)

On-site
USD 120,000 - 180,000
Staff AI Engineer
Staff AI Engineer

Harnham • San Francisco (CA)

On-site
USD 180,000 - 240,000