Staff Software Engineer, Cloud Infrastructure

Engg

San Carlos

Híbrido

ARS 289.387.000 - 357.926.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Healthcare
Time Off
401(k)

Descripción de la vacante

Beacon AI is seeking a Staff Software Engineer, Cloud Infrastructure to lead our AWS foundation and LLM platform. You will design, implement, and operate scalable, secure services. The role emphasizes ML infra, IoT infra, and cost governance, with a hands‑on, first-principles approach.

The position involves building retrieval pipelines, model endpoints, and secure data flows, collaborating across engineering and product teams in a fast-moving environment.

Formación

  • Hands-on, multi-cloud capable with strong AWS expertise.
  • Experience designing scalable, secure cloud infra and ML pipelines.
  • Familiarity with LLM infra, retrieval pipelines, and LangChain is a plus.

Responsabilidades

  • Design, provision, and maintain AWS infrastructure with IaC.
  • Build CI/CD and testing for apps, infra, and ML pipelines.
  • Operate secure networking, IAM, and audit logging.
  • Stand up model endpoints and manage LLM services.
  • Develop RAG data systems, vector search, and data pipelines.
  • Implement monitoring, cost governance, and guardrails.
  • Ensure safety, privacy, and regulatory compliance.
  • Collaborate with product and application teams on features.
  • Optimize performance and scaling for ML/inference workloads.

Conocimientos

AWS
IaC
CDK
Terraform
CI/CD
GitHub Actions
SageMaker
Bedrock
LangChain
Security

Herramientas

OpenSearch
Aurora
Glue
Lake Formation

Descripción del empleo

ABOUT BEACON AI

We’re a fast-moving team of aviators, engineers, and operators building an AI platform to make flying safer, more efficient, and more capable. Backed by top investors, we’ve secured a dozen Department of Defense contracts and partnered with major airlines to deliver mission-critical systems. We operate without silos or heavy processes. Small, focused teams own what they build, ship quickly, and learn fast, pushing the boundaries of how humans and AI work together in aviation.

ROLE OVERVIEW

We are seeking skilled Staff Software Engineers, Cloud Infrastructure to lead the buildout of our AWS foundation and our LLM platform. You will design, implement, and operate services that are scalable, reliable, and secure. The broad scope means focus areas in LLM/ML Infra and IoT infra are strong bonus points. For ML Infra, build the stack that powers retrieval-augmented generation and application workflows built with frameworks like LangChain. Experience with IoT AWS services is a plus. You will work closely with other engineers and product management. The ideal candidate is hands‑on, comfortable with ambiguity, and excited to build from first principles.

KEY RESPONSIBILITIES
  • Cloud Infrastructure Setup and Maintenance
    • Design, provision, and maintain AWS infrastructure using IaC tools such as AWS CDK or Terraform.
    • Build CI/CD and testing for apps, infra, and ML pipelines using GitHub Actions, CodeBuild, and CodePipeline.
    • Operate secure networking with VPCs, PrivateLink, and VPC endpoints. Manage IAM, KMS, Secrets Manager, and audit logging.
  • LLM Platform and Runtime
    • Stand up and operate model endpoints using AWS Bedrock and/or SageMaker; evaluate when to use ECS/EKS, Lambda, or Batch for inference jobs.
    • Build and maintain application services that call LLMs through clean APIs, with streaming, batching, and backoff strategies.
    • Implement prompt and tool execution flows with LangChain or similar, including agent tools and function calling.
  • RAG Data Systems and Vector Search
    • Design chunking and embedding pipelines for documents, time series, and multimedia. Orchestrate with Step Functions or Airflow.
    • Operate vector search using OpenSearch Serverless, Aurora PostgreSQL with pgvector, or Pinecone. Tune recall, latency, and cost.
    • Build and maintain knowledge bases and data syncs from S3, Aurora, DynamoDB, and external sources.
  • Evaluation, Observability, and Cost Governance
    • Create offline and online eval harnesses for prompts, retrievers, and chains.
    • Track quality, latency, and regression risk.
    • Instrument model and app telemetry with CloudWatch and OpenTelemetry. Build token usage and cost dashboards with budgets and alerts.
    • Add guardrails, rate limits, fallbacks, and provider routing for resilience.
  • Safety, Privacy, and Compliance
    • Implement PII detection and redaction, access controls, content filters, and human-in-the-loop review where needed.
    • Use Bedrock Guardrails or policy services to enforce safety standards.
    • Maintain audit trails for regulated environments.
  • Data Pipeline Construction
    • Build ingestion and processing pipelines for structured, unstructured, and multimedia data. Ensure integrity, lineage, and cataloging with Glue and Lake Formation.
    • Optimize bulk data movement and storage in S3, Glacier, and tiered storage.
    • Use Athena for ad-hoc analysis.
  • IoT Deployment Management
    • Manage infrastructure that deploys to and communicates with edge devices. Support secure messaging, identity, and over-the-air updates.
  • Analytics and Application Support
    • Partner with product and application teams to integrate retrieval services, embeddings, and LLM chains into user-facing features.
  • Performance Optimization
    • Tune retrieval quality, context window use, and caching with Redis or Bedrock Knowledge Bases.
    • Optimize inference with model selection, quantization where applicable, GPU/CPU instance choices, and autoscaling strategies.
WHAT WILL MAKE YOU SUCCESSFUL
  • End-to-End Ownership: Drives work from design through production, including on-call and continuous improvement.
  • LLM Systems Experience: Shipped or operated LLM-powered applications in production. Familiar with RAG design, prompt versioning, and chain orchestration using LangChain or similar.
  • AWS Depth: Strong with core AWS services such as VPC, IAM, KMS, CloudWatch, S3, ECS/EKS, Lambda, Step Functions, Bedrock, and SageMaker.
  • Data Engineering Skills: Comfortable building ingestion and transformation pipelines in Python. Familiar with Glue, Athena, and event-driven patterns using EventBridge and SQS.
  • Security Mindset: Applies least privilege, secrets management, network isolation, and compliance practices appropriate to sensitive data.
  • Evaluation and Metrics: Uses quantitative evals, A/B testing, and live metrics to guide improvements.
  • Clear Communication: Explains tradeoffs and aligns partners across product, security, and application engineering.
BONUS POINTS
  • 9+ years working with serverless or container platforms on AWS.
  • Experience with vector databases, OpenSearch, or pgvector at scale.
  • Hands‑on with Bedrock Guardrails, Knowledge Bases, or custom policy engines.
  • Familiarity with GPU workloads, Triton Inference Server, or TensorRT-LLM.
  • Experience with big data tools for large-scale processing and search.
  • Background in aviation data or other safety‑critical domains.
  • DevOps or DevSecOps experience automating CI/CD for ML and app services.
WORK LOCATION

This is a hybrid role based in San Carlos, CA, with 3+ days per week onsite and the option to work remotely on remaining days.

PERKS & BENEFITS (FULL-TIME EMPLOYEES)
  • Healthcare: 100%* of employee medical premiums covered; 25% for dependents
  • Time Off: 3 weeks PTO plus 13+ paid company holidays
  • 401(k): Offered (no current employer match, but we are committed to enhancing this benefit in the future)

Due to U.S. ex

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Software Engineer, Cloud Infrastructure
Senior Software Engineer, Cloud Infrastructure

Engg • San Carlos

Híbrido
ARS 228.464.000 - 289.387.000
Healthcare benefits
3 weeks PTO + holidays
401(k) plan
Software Engineer, Cloud Infrastructure (Multiple Seniority Levels)
Software Engineer, Cloud Infrastructure (Multiple Seniority Levels)

Engg • San Carlos

Presencial
ARS 257.209.000 - 317.729.000
Healthcare coverage
3 weeks PTO
401(k)
Staff Software Engineer, Cloud Infrastructure
Staff Software Engineer, Cloud Infrastructure

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 380.773.000
Healthcare
PTO
401(k)
Senior Software Engineer, Cloud Infrastructure
Senior Software Engineer, Cloud Infrastructure

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 365.542.000
Healthcare coverage
PTO and holidays
401(k) plan
Staff Software Engineer, Artificial Intelligence/LLM
Staff Software Engineer, Artificial Intelligence/LLM

Engg • San Carlos

Híbrido
ARS 228.464.000 - 319.849.000
Healthcare coverage
PTO 3 weeks
Company holidays
Senior Software Engineer, Artificial Intelligence/LLM
Senior Software Engineer, Artificial Intelligence/LLM

Engg • San Carlos

Híbrido
ARS 274.156.000 - 350.311.000
Healthcare coverage
PTO
401(k) plan
Staff Software Engineer, Artificial Intelligence/LLM
Staff Software Engineer, Artificial Intelligence/LLM

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 365.542.000
Healthcare coverages
Time off (PTO)
401(k) plan
Software Engineer, Artificial Intelligence/LLM
Software Engineer, Artificial Intelligence/LLM

Engg • San Carlos

Híbrido
ARS 274.156.000 - 396.003.000
Healthcare 100% coverage for employees
3 weeks PTO + 13+ holidays
401(k) plan
Senior Software Engineer, Artificial Intelligence/LLM
Senior Software Engineer, Artificial Intelligence/LLM

Beacon AI • San Carlos

Híbrido
ARS 228.464.000 - 319.849.000
Healthcare coverage
Paid time off
401(k)
Lead Backend Software Engineer
Lead Backend Software Engineer

Engg • San Carlos

Presencial
ARS 272.339.000 - 363.119.000
Healthcare premiums covered
3 weeks PTO + 13+ holidays
401(k) plan