Senior Software Engineer, Artificial Intelligence/LLM

Beacon AI

San Carlos

Híbrido

ARS 228.464.000 - 319.849.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura completa en un minuto — currículum adaptado y carta de presentación, listos para enviar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Healthcare coverage
Paid time off
401(k)

Descripción de la vacante

Beacon AI in San Carlos, CA is seeking a senior ML/LLM engineer to own features end-to-end, from retrieval-augmented generation to tool-calling and production deployment.

You will collaborate with ML/infra and product teams to design scalable, low-latency, and cost-aware solutions in a safety-critical aviation domain. The role is hybrid with onsite and remote work options within the US.

Formación

  • 5-8 years experience in ML/LLM, production ML/LLM systems preferred.
  • Ability to own features end-to-end from design to production.
  • Strong emphasis on evaluation, guardrails, and cost/latency controls.

Responsabilidades

  • Build user-facing LLM features with robust outputs and schema validation.
  • Own the service layer: APIs, workers, and contracts in Python or TypeScript.
  • Handle retrieval, embeddings, and indexing for documents and time series.
  • Create offline evals and online metrics for task success and latency.
  • Implement safety, privacy, and compliance measures across systems.
  • Operate and monitor the deployed services with tracing and dashboards.

Conocimientos

LLM development
RAG systems
Tool-calling
Python
TypeScript
ML evals
Latency optimization

Herramientas

LangChain
OpenAI
Bedrock
Anthropic
OpenSearch
Pinecone

Descripción del empleo

About Beacon AI

We’re a fast-moving team of aviators, engineers, and operators building an AI platform to make flying safer, more efficient, and more capable. Backed by top investors, we’ve secured a dozen Department of Defense contracts and partnered with major airlines to deliver mission-critical systems. We operate without silos or heavy processes. Small, focused teams own what they build, ship quickly, and learn fast, pushing the boundaries of how humans and AI work together in aviation.

You will ship LLM-powered product features end-to-end. That means designing retrieval and tool-calling flows, writing the services that run them, building evals and guardrails, and watching cost, latency, and quality in production. You’ll partner with the ML/infra teammates on embeddings, indexing, and model hosting, and with the product teammates on user experience and outcomes. We move fast, and we care about reliability in a safety-critical domain.

This role is for senior engineers who can own a feature or service end-to-end (5-8 years experience, including some in production ML/LLM systems). You'll make independent architecture and eval decisions within your service's scope and partner directly with product and infra.

What you’ll do

Build user-facing LLM features

  • Design and implement retrieval-augmented generation and tool-calling flows using frameworks like LangChain or equivalent primitives, where simpler is better.

  • Deliver robust JSON and schema-bound outputs with validation, retries, and fallbacks.

  • Add function calling to integrate with internal tools, search, routing, and data services.

Own the service layer

  • Ship APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff.

  • Add caching, request shaping, prompt templates, and context packing to control latency and cost.

  • Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints as needed.

Retrieval and data prep

  • Collaborate with infrastructure teammates to develop chunking, embeddings, and indexing capabilities for documents, time series, and multimedia.

  • Choose and tune vector backends such as OpenSearch, pgvector, or Pinecone.

  • Keep knowledge bases fresh with data syncs from S3, Aurora, DynamoDB, and external sources.

Evaluation and quality

  • Create offline evals and golden sets for prompts, retrievers, and tools.

  • Stand up online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request.

  • Run A/B tests and prompt/version rollouts with guardrails and canaries.

Safety, privacy, and compliance

  • Implement content and policy checks, PII detection and redaction, access controls, and auditing.

  • Design human-in-the-loop paths for sensitive actions.

  • Handle aviation data with care and follow internal security standards.

Operate what you build

  • Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation.

  • Debug tricky failures across retrieval, prompts, tools, and providers.

What will make you successful
  • Shipped LLM apps: You’ve put LLM features in front of users and improved them with data.

  • Strong builder: Comfortable writing production code, tests, and docs. You keep things simple and observable.

  • RAG and tools depth: You understand embeddings, chunking, vector search tradeoffs, and function calling.

  • Quality mindset: You design evals, define success metrics, and iterate based on evidence.

  • Cost and latency aware: You track p95, hit SLAs, and reduce cost without hurting quality.

  • Clear communicator: You explain tradeoffs and align partners across product, infra, and security.

  • Ownership: You can take a feature from design through production with minimal oversight.

Nice to have
  • Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.

  • Prompt versioning, guardrails, and provider routing in production.

  • Multimodal work with time series or video.

  • Familiarity with GPU inference, Triton, or TensorRT-LLM.

  • Aviation or other safety-critical domain exposure.

  • DevOps basics for CI/CD, IaC, and secure secrets handling.

Example problems you might tackle in month one
  • Transform an internal knowledge base into a low-latency RAG service, complete with explicit schemas and evaluations.

  • Add tool-calling to automate a repetitive cockpit or ops workflow with guardrails and audit trails.

  • Reduce the cost per request through improved chunking, caching, and prompt refactoring, while maintaining task success rates.

Work Location
This is a hybrid role based in San Carlos, CA, with 3+ days per week onsite and the option to work remotely on remaining days.

Perks & Benefits (Full-Time Employees)

  • Healthcare: 100%* of employee medical premiums covered; 25% for dependents

  • Time Off: 3 weeks PTO plus 13+ paid company holidays

  • 401(k): Offered (no current employer match, but we are committed to enhancing this benefit in the future)

Due to U.S. export control regulations, we can only hire U.S. Persons (U.S. citizens, Green Card holders, lawful permanent residents, or individuals granted asylum or refugee status). We are unable to provide visa sponsorship or support visa transfers. All work must be performed in the United States.

Beacon AI is an equal opportunity employer and does not discriminate based on race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected characteristic. We prohibit harassment or discrimination of any kind in the workplace and comply with all applicable federal, state, and local employment laws.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Staff Software Engineer, Artificial Intelligence/LLM
Staff Software Engineer, Artificial Intelligence/LLM

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 365.542.000
Healthcare coverages
Time off (PTO)
401(k) plan
Senior Software Engineer, Cloud Infrastructure
Senior Software Engineer, Cloud Infrastructure

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 365.542.000
Healthcare coverage
PTO and holidays
401(k) plan
Staff Software Engineer, Cloud Infrastructure
Staff Software Engineer, Cloud Infrastructure

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 380.773.000
Healthcare
PTO
401(k)
Software Engineer, Cloud Infrastructure (Multiple Seniority Levels)
Software Engineer, Cloud Infrastructure (Multiple Seniority Levels)

Engg • San Carlos

Presencial
ARS 257.209.000 - 317.729.000
Healthcare coverage
3 weeks PTO
401(k)
Lead Backend Software Engineer
Lead Backend Software Engineer

Engg • San Carlos

Presencial
ARS 272.339.000 - 363.119.000
Healthcare premiums covered
3 weeks PTO + 13+ holidays
401(k) plan
Lead Software Engineer, Frontend/Web App
Lead Software Engineer, Frontend/Web App

Engg • San Carlos

Presencial
ARS 210.935.000 - 316.403.000
Healthcare
PTO
401(k)
Lead Software Engineer, Advanced Pilot Assistant Software (Autonomy/Robotics)
Lead Software Engineer, Advanced Pilot Assistant Software (Autonomy/Robotics)

Engg • San Carlos

Presencial
ARS 211.820.000 - 287.469.000
Healthcare: 100% employee premiums
3 weeks PTO + 13 holidays
401(k) plan
Software Engineer, Frontend/Web App
Software Engineer, Frontend/Web App

Engg • San Carlos

Presencial
ARS 166.430.000 - 257.209.000
Healthcare
PTO & Holidays
401(k)
Staff Software Engineer, Frontend/Web App
Staff Software Engineer, Frontend/Web App

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 350.311.000
Healthcare: 100% employee premiums
3 weeks PTO + 13+ holidays
401(k): Planned improvement in future
Senior Software Engineer, iOS/Mobile
Senior Software Engineer, iOS/Mobile

Fab2 • San Carlos

Híbrido
ARS 182.771.000 - 243.694.000
Healthcare: 100%* of employee medical
Time Off: 3 weeks PTO + 13+ holidays
401(k): Offered