Llm Application Engineer

Bolder Apps

Argentina

Remote

ARS 152,979,000 - 254,965,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote work
Autonomy
Professional growth

Job summary

Bolder Apps seeks a mid-senior LLM Application Engineer to design, ship, and harden production AI pipelines for client products. You will own structured extraction across HTML, PDFs, and images, with end-to-end orchestration, confidence flags, retries, and cost-aware optimization.

You will work with Google Gemini and other major providers, building scalable backends in Firebase/GCP and collaborating with product, mobile, and QA teams to meet strict SLAs and quality targets.

Qualifications

  • Shipped LLM apps to production with real users.
  • Strong experience with Google Gemini and multimodal workflows.
  • Proficient in Python backend and cloud functions.

Responsibilities

  • Own production LLM pipelines end to end: ingestion, multimodal model calls, structured records, and storage, including confidence flags and retries.
  • Design prompt and schema strategies for consistent, product-ready outputs.
  • Build classification and filtering layers on top of extraction.

Skills

LLM pipelines
Google Gemini
Multimodal prompts
Python backend
Serverless cloud
Firestore
OpenAI
English proficiency

Tools

GCP
Firebase
Document AI
Gmail API

Job description

We're hiring a mid-senior LLM Application Engineer on a remote monthly retainer to design, ship, and harden production AI pipelines for client products at Bolder Apps. You'll own structured extraction and classification systems that turn messy real-world inputs (email, HTML, PDFs, images) into reliable product data, with measurable quality gates, evals, and cost control. You'll work on Firebase / GCP-style backends with product, mobile, and QA. We want someone who has shipped LLM apps for real users, not demos. You should be strong across modern LLMs and especially fluent with Google Gemini (multimodal prompts, structured outputs, failure modes, and cost/latency tradeoffs), with solid experience on other major providers too. If you can hit hard quality targets, keep dollars-per-run honest, and leave runbooks another engineer can pick up, we want to talk.

About Us

Bolder Apps is a product development studio that partners with US-based startups and established companies to build and scale innovative digital products. We specialize in AI-powered development, full-cycle product creation, and engineering team augmentation. Our mission is simple: build bolder, faster, and smarter.

Our Culture & Values (read before applying)

We move fast. We take ownership. We work with AI, not against it. And we expect everyone to bring ideas, not wait for instructions.

There are no daily checklists, no micromanagement, and no corporate politics. Instead, you'll have autonomy, trust, and a team that's always ready to help you grow. At Bolder Apps, impact matters more than titles, and curiosity matters more than seniority.

If you want a place where you can level up fast and actually see your work making a difference - welcome aboard.

Requirements
Responsibilities
  • Own production LLM pipelines end to end: ingestion, multimodal model calls, structured records, and storage, including confidence flags, retries, and idempotent rescans
  • Design prompt and schema strategies (including schema-aligned or constrained outputs) so results are consistent and product-ready
  • Build classification and filtering layers on top of extraction (taxonomy mapping, demographic or audience filters, deduplication, and related cleanup logic)
  • Define and run evaluation harnesses (golden sets, regression suites, online metrics) so quality does not regress when prompts, models, or parsers change
  • Hit and report against hard quality targets (precision-style gates for completeness, duplicates, incorrect inclusions, image presence, and similar product SLAs)
  • Optimize token usage, model tiering, caching, and batching to keep dollars-per-run and latency under control
  • Harden reliability for long-running async jobs (timeouts, partial recovery, memory limits, safe production deploys)
  • Partner with Flutter / mobile and QA on field contracts, review queues, and incident debugging
  • Document architecture and runbooks so ownership is shared, not a single point of failure
  • Stay current on Gemini and peer LLM APIs; recommend when to swap models, add fallbacks (e.g. document AI), or tighten schemas
  • Shipped LLM applications in production (not demos only): prompts, structured outputs, retries, observability, and real failure handling
  • Strong hands-on experience with Google Gemini, including multimodal (text + image / document-style) workflows and structured extraction
  • Practical experience with at least one other major LLM stack (OpenAI, Anthropic, or similar) and good judgment on when to use which
  • Structured extraction from messy inputs: HTML, PDFs, images, and mixed email-like content
  • Classification / taxonomy systems on top of LLM outputs
  • Evaluation discipline: offline evals, regression suites, and production quality metrics tied to clear acceptance criteria
  • Cost and latency awareness: token budgeting, cheaper tiers, caching, batching; can explain dollars-per-run tradeoffs to a PM
  • Python backend experience on serverless cloud (Cloud Functions or equivalent) and document stores (e.g. Firestore) or similar GCP patterns
  • English at C1 or above for client-adjacent debugging with a PM
  • US hours overlap through roughly 5 PM EST when live coordination is needed
  • Ownership habits: honest estimates, early blockers, finished releases
Nice to have
  • Schema-aligned LLM frameworks (BAML, Instructor, Outlines, or similar)
  • Google Document AI or other OCR / document intelligence as a fallback path
  • Gmail API / OAuth products and restricted-scope compliance familiarity
  • Computer vision for product-image quality checks
  • Building eval corpora from real production data and iterating until contractual SLAs pass
  • Agency or multi-client studio experience
  • Fully remote and async-friendly, with required overlap through ~5 PM EST when client or release coordination needs it
  • Monthly retainer structure with recurring AI pipeline work for engineers who keep production quality and cost honest
  • Real autonomy over how you structure prompts, schemas, evals, and deploys. We do not hand you a rigid playbook
  • Direct line to PMs, mobile engineers, and decision-makers
  • Tooling budget for the LLM and cloud tools you need to move fast
  • A peer network of product-minded builders across overlapping client projects
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote LLM Applications Engineer - Production Pipelines
Remote LLM Applications Engineer - Production Pipelines

Bolder Apps • Argentina

Remote
ARS 152,979,000 - 254,965,000
Remote work
Autonomy
Professional growth
Staff Software Engineer, Artificial Intelligence/LLM
Staff Software Engineer, Artificial Intelligence/LLM

Beacon AI • San Carlos

Hybrid
ARS 274,156,000 - 365,542,000
Healthcare coverages
Time off (PTO)
401(k) plan
Staff Software Engineer, Cloud Infrastructure
Staff Software Engineer, Cloud Infrastructure

Beacon AI • San Carlos

Hybrid
ARS 274,156,000 - 380,773,000
Healthcare
PTO
401(k)
Senior Software Engineer, Artificial Intelligence/LLM
Senior Software Engineer, Artificial Intelligence/LLM

Beacon AI • San Carlos

Hybrid
ARS 228,464,000 - 319,849,000
Healthcare coverage
Paid time off
401(k)
Staff Software Engineer, Artificial Intelligence/LLM
Staff Software Engineer, Artificial Intelligence/LLM

BeaconAI • San Carlos

On-site
ARS 228,464,000 - 319,849,000
Healthcare coverage
PTO 3 weeks
Company holidays
Senior Python Backend Developer / ML Engineer (IR-535)
Senior Python Backend Developer / ML Engineer (IR-535)

Intellectsoft • Argentina

On-site
ARS 1,800,000 - 3,200,000
Awesome projects with an impact
Udemy courses of your choice
Team-building events
+2
Senior Software Engineer, Cloud Infrastructure
Senior Software Engineer, Cloud Infrastructure

Beacon AI • San Carlos

Hybrid
ARS 274,156,000 - 365,542,000
Healthcare coverage
PTO and holidays
401(k) plan
Software Engineer, Artificial Intelligence/LLM
Software Engineer, Artificial Intelligence/LLM

BeaconAI • San Carlos

On-site
ARS 274,156,000 - 396,003,000
Healthcare 100% coverage for employees
3 weeks PTO + 13+ holidays
401(k) plan
Senior Software Engineer, Artificial Intelligence/LLM
Senior Software Engineer, Artificial Intelligence/LLM

BeaconAI • San Carlos

On-site
ARS 274,156,000 - 350,311,000
Healthcare coverage
PTO
401(k) plan
Senior AI and Agentic Engineer (PCS847)
Senior AI and Agentic Engineer (PCS847)

Overseas • Buenos Aires

On-site
ARS 152,323,000 - 213,252,000