AI Engineer - Generative AI and Agents

Azumo

Brasil

Teletrabalho

BRL 624 000 - 936 000

Tempo integral

Há 11 dias
Gerador de candidaturas

Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Remote-first culture (LA)
Paid time off (PTO)
US holidays
AI training & certifications
Mentored career development
Profit sharing
US remuneration

Resumo da oferta

Azumo is hiring an AI Engineer to own production behavior of AI systems, including retrieval pipelines, tool-using agents, evaluation harnesses, and robust guardrails. The role operates fully remotely across Latin America, aligned to client time zones and workflows, with production deployments and measurable reliability goals.

You will work within Azumo's four-lane engineering structure, owning the AI Engineer lane and contributing to secure, scalable, and cost-aware AI systems.

Qualificações

  • 4+ years building and shipping production software with a modern backend language.
  • Proven production experience with LLM-based systems including retrieval-augmented generation and tool calling.
  • Hands-on work with vector and retrieval infrastructure such as pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search.
  • Experience with agent frameworks and tool-using systems and stateful pipelines.
  • Experience designing evaluation suites and metrics for LLM-based systems.
  • Cloud deployment experience with Docker, CI/CD pipelines, and IaC.

Responsabilidades

  • Own production behavior of AI systems in live environments.
  • Build and maintain retrieval pipelines, tool-using agents, and evaluation harnesses.
  • Implement guardrails for reliability, cost, and latency.
  • Deploy containerized services on Azure or AWS with observability.
  • Collaborate with Data/Data Scientist lanes and ship production software for clients.

Conhecimentos

4+ years experience
Backend development (Python)
English communication

Formação académica

Bachelor's degree in CS or related

Ferramentas

pgvector
Pinecone
Qdrant
FAISS
Azure AI Search
LangGraph
LangChain
CrewAI
MCP

Descrição da oferta de emprego

Azumo builds and operates production AI systems for companies ranging from seed-stage startups to Meta. We are hiring an AI Engineer to own what those systems do once they are live: retrieval pipelines, tool-using agents, evaluation harnesses, and the guardrails that keep them dependable in front of real users. The role is fully remote across Latin America, aligned to your client's working day.

You will not be building demos. Azumo has shipped production AI since 2016, and the work here starts where the prototype ends, making a system reliable, measurable, and affordable enough to put in front of customers.

Where this role sits

Azumo's engineering organization is built around four lanes. The Data Engineer lane owns pipelines, storage, and the retrieval layer. The Data Scientist lane owns the question and the method. The Software Engineer lane owns AI-augmented product delivery. This role is the AI Engineer lane, and it owns production behavior.

One question places the boundary: when the output is wrong, whose problem is it? "The method was inappropriate" is a Data Scientist question. "The system did the wrong thing with an appropriate method" is yours.

Not quite your profile? Check our other openings:
  • If you build the pipelines and retrieval layer models depend on — Data Engineer
  • If you decide what to measure and which method answers it — Data Scientist
  • If you ship product software with agents in your toolchain — AI-Augmented Software Engineer
  • If you've done all of the above and answered to the client directly — Forward Deployed Engineer
What you will build
  • Retrieval systems. Chunking and embedding pipelines, hybrid search, reranking, and evaluation of retrieval quality, built on pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search.
  • Agentic workflows. Stateful multi-step execution, tool calling, MCP servers, structured output enforcement, context-window management, deterministic fallbacks, and human-in-the-loop gates for the decisions that need one.
  • Evaluation. Test sets that reflect the decision the system is actually making, model-as-judge scoring, regression tracking across prompt and model changes, and honest error analysis. If a change made the system better, you should be able to prove it.
  • Reliability and safety. Prompt-injection defense, output validation, guardrails, PII handling, and graceful degradation when a model or tool call fails.
  • Production operation. Containerized deployment on Azure or AWS, CI/CD, observability, and explicit latency, cost, and token budgets that you own rather than discover after the invoice.
  • Work inside the client's environment. Their repositories, their standups, sometimes their customer calls. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.
How we work

Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an automated audit across the whole codebase on day one and every day after, grading security, cost, and architecture findings by severity with the exact file and line, so a small team can move quickly without quality drifting. We stay vendor-neutral across OpenAI, Anthropic, and open-weight models, and we run Valkyrie, our own production layer, when a single interface to any model is the right call.

About Azumo

Azumo is a San Francisco based software development company that has been building intelligent applications since 2016. We provide nearshore AI engineering teams to organizations that need production AI faster than they can hire for it: as an embedded engineering team, as AI staff augmentation alongside an existing team, or as a full project build. Our engineers work from Latin America, aligned to United States time zones, and have delivered for Twitter, Meta, Discovery Channel, Omnicom, UnitedHealth, and CENTEGIX.

We hire for seniority and test for it before anyone joins a client team. We support engineers in going deep on the modern AI stack, and we give time back to open-source work, community teaching, and philanthropy.

Requirements
Basic qualifications
  • 4+ years building and shipping production software, with a modern backend language (e.g., Python) as your primary focus, plus the engineering fundamentals that go with it: testing, code review, CI/CD, Git, containers, API design, and async programming.
  • Demonstrable production experience with LLM-based systems: retrieval-augmented generation, function and tool calling, structured output enforcement, and prompt design as an engineering discipline rather than trial and error.
  • Hands-on work with vector and retrieval infrastructure such as pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search, including the retrieval-quality problems that come with it.
  • Experience with agent frameworks and tooling: LangGraph, LangChain, CrewAI, the Model Context Protocol (MCP), or native Python execution loops. We care that you have shipped a stateful, tool-using system, not which framework you used.
  • You have built an evaluation suite for an LLM system. Test-set design, model-as-judge or equivalent scoring, and regression tracking with Langfuse, Ragas, LangSmith, or something you wrote yourself. This is the requirement we screen hardest on.
  • Cloud deployment experience, Azure preferred and AWS acceptable, with Docker, CI/CD pipelines, and infrastructure as code (GitHub Actions, Terraform, or Bicep).
  • Working discipline around latency, token cost, and throughput. You can explain what a feature costs to run and what you did about it.
  • Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in real delivery work.
  • Clear written and spoken English, C1 or above, and the confidence to explain a technical trade-off directly to a client.
  • Bachelor's degree in Computer Science, Data Science, or a related field, or equivalent professional experience.
Preferred qualifications
  • Fine-tuning and adaptation of open-weight models: LoRA, QLoRA, PEFT, and a clear view of when fine-tuning is the wrong answer.
  • Self-hosted or open-weight inference and serving, and the cost and latency trade-offs against hosted APIs.
  • Multimodal systems covering vision, speech, or document understanding alongside text.
  • Security work specific to LLM systems: prompt-injection testing, red-teaming, and output sanitization.
  • Delivery under a compliance regime such as SOC 2 or HIPAA.
  • Streaming, real-time, or high-throughput inference workloads.
  • Contributions to open-source AI libraries, published technical writing, or active participation in the AI engineering community.
Benefits
  • 100% remote-first culture (work anywhere in Latin America)
  • Paid time off (PTO)
  • U.S. Holidays
  • AI Training and certifications
  • Mentored career development
  • Profit sharing
  • $US remuneration.
Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Cloud Devops Engineer
Cloud Devops Engineer

Azumo • São Paulo

Teletrabalho
BRL 467 000 - 674 000
Remote-first culture
Paid time off
US Holidays
+5
AI Software Engineer (Generative AI) - Latin America - Remote
AI Software Engineer (Generative AI) - Latin America - Remote

Azumo • Brasil

Presencial
BRL 180 000 - 280 000
Paid time off (PTO)
U.S. Holidays
Training
+4
Web, Mobile and Desktop Development
Web, Mobile and Desktop Development

Rabelo Digital • Brasil

Teletrabalho
BRL 180 000 - 300 000
AI Engineer
AI Engineer

Solvedex • Brasil

Presencial
BRL 180 000 - 280 000
Senior Deployed AI Engineer – OpenAI
Senior Deployed AI Engineer – OpenAI

LinkedIn Job Wrapping • São Paulo

Presencial
BRL 721 000 - 927 000
Senior ML Specialist
Senior ML Specialist

careers-quartile • Brasil

Presencial
BRL 250 000 - 420 000
AI/ML Engineer
AI/ML Engineer

RemoteJobsOne • São Paulo

Teletrabalho
BRL 1 040 000 - 1 819 000
Senior Python Backend Engineer
Senior Python Backend Engineer

BriteCore • Rio de Janeiro

Presencial
BRL 1 200 000 - 2 400 000
Senior MLOps Engineer - Remote - Latin America
Senior MLOps Engineer - Remote - Latin America

FullStack • Manaus

Teletrabalho
BRL 240 000 - 300 000
Competitive pay
100% remote work
Work with leading startups and Fortune
+2
Senior MLOps Engineer - Remote - Latin America
Senior MLOps Engineer - Remote - Latin America

FullStack • Salvador

Teletrabalho
BRL 618 000 - 1 030 000
Competitive pay
100% remote work
Work with leading startups and Fortune
+2