[VMT] Platform AI Software Engineer

Softwaremind

Buenos Aires

Presencial

ARS 105.515.000 - 165.810.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Educational resources
Flexible schedule
Work From Anywhere
Referral Program
Supportive atmosphere

Descripción de la vacante

Softwaremind is seeking a backend engineer to harden an AI copilot for process engineers in oil refineries and chemical plants. You will own the trust layer and grounding mechanisms within a large Python codebase, ensuring the agent says "I don't know" and surfaces assumptions for correction.

You will diagnose misbehavior, extend evaluation harnesses, improve latency, and apply concurrency fixes, all while supporting SOC-like reliability and staged rollouts across teams in

Formación

  • Strong Python in large, shared backend codebases.
  • Shipped an LLM-based feature to production and built an evaluation that changed a real decision.
  • Production debugging from symptom to confirmed root cause.
  • Prompt work treated as engineering with measured adherence and regression testing.
  • Experience with feature flags and staged rollouts across time zones.

Responsabilidades

  • Investigate agent misbehavior from real customer plants and diagnose, fix, and regression-test.
  • Build and extend evaluation harnesses that gate prompt changes and model migrations.
  • Add reliability mechanisms: assumption surfacing, grounding checks, output-quality reviewers.
  • Diagnose and resolve cross-cutting performance and concurrency issues.
  • Raise code quality within the team's review conventions.

Conocimientos

Strong Python
LLM in production
Prompt engineering
Observability tooling

Herramientas

Matplotlib

Descripción del empleo

Project - the aim you'll have

Our client builds an AI copilot for process engineers in oil refineries and chemical plants: a natural-language interface where engineers ask questions about live plant data - equipment, sensor tags, process trends - and get grounded, chart-backed answers. The users are experienced engineers who are rightly skeptical of AI: in this domain, a fabricated number or a silent wrong assumption has real cost. The product wins or loses on whether the agent can be trusted.

This role owns the trust layer of that agent inside a large, active Python codebase. It is not feature work with an LLM endpoint bolted on. The work is the mechanics of agent reliability: making the agent say "I don't know" instead of inventing, surfacing every assumption it makes so the user can correct it, grounding every claim in actual data, holding output quality through model migrations, and keeping latency acceptable while doing all of the above.

  • An assumption auditor: detects the silent assumptions the agent makes when answering (which equipment, which time window), validates them via multi-draw consensus, and surfaces them in the UI as correctable chips - the engineer can fix an assumption and rerun the analysis.
  • A grounding auditor that catches reports fabricated from empty data feeds before they reach the user.
  • An adversarial reviewer sidecar that critiques generated charts for correctness before display.
  • Successive frontier-model evaluations (loop behavior, directive adherence, regression on a replay harness) that decided when to flip the product's default model - including, twice, deciding NOT to flip.
  • Hardening of a plant-exploration tool against hallucinating structure that the data does not support.
  • A latency fix: a narration side-loop was inflating query response times; capped it and made it best-effort.
  • A concurrency fix making a shared data-reset path atomic, eliminating intermittent production read errors.

If reading that list is more interesting to you than building another CRUD feature, this role is for you.

Expectations - the experience you need
  • Strong Python in large, shared, evolving backend codebases: you will work daily in code you didn't write, alongside people committing to it every day.
  • You have shipped an LLM-based feature to production AND built an evaluation that changed a real decision (a model choice, a prompt rollback, a killed feature).
  • Production debugging from symptom to confirmed root cause: latency spikes, concurrency errors, failures that produce no log line.
  • Prompt work treated as engineering: measured adherence, regression testing against a fixed case set - not vibes.
  • Comfort with feature-flag discipline and staged rollouts (default-off, soak, flip), small PRs, and mostly-async collaboration across US and Australia time zones.
  • High autonomy: problems arrive ambiguous ("the agent feels slow", "the engineers don't trust the numbers") and you turn them into scoped, verifiable fixes without waiting for a spec.
  • Direct, precise written English.
Nice to have
  • GCP (Vertex AI in particular); AWS/Azure acceptable.
  • Observability tooling (tracing, structured logging, latency percentiles you actually watched).
  • Experience with charting/plotting pipelines (matplotlib or similar) feeding a UI.
  • Industrial, process, or time-series data domain experience.
  • Heavy AI-tooling development workflow (Claude Code or similar) - the team works this way.
What you will do
  • Investigate agent misbehavior reported from real customer plants and turn each case into a diagnosis, a fix, and a regression test.
  • Build and extend the evaluation harnesses that gate prompt changes and model migrations.
  • Add reliability mechanisms to the agent: assumption surfacing, grounding checks, output-quality reviewers.
  • Diagnose and resolve cross-cutting performance and concurrency issues.
  • Raise code quality in the areas you touch, within the team's review conventions.
Our Benefits
  • Educational resources
  • Flexible schedule and Work From Anywhere
  • Referral Program
  • Supportive and chill atmosphere

We are accepting applications from LATAM countries

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Applied AI Engineer
Senior Applied AI Engineer

Network Solutions • Argentina

Presencial
ARS 1.800.000 - 3.000.000
Sr Lead Software Engineer - AI Agents
Sr Lead Software Engineer - AI Agents

JPMorgan Chase & Co. • Buenos Aires

Presencial
ARS 181.560.000 - 272.339.000
Senior Software Engineer: Agentic Ai & Production Leader
Senior Software Engineer: Agentic Ai & Production Leader

Agileengine, Llc. • Buenos Aires

Híbrido
ARS 800.000 - 1.600.000
Sr. AI Engineer
Sr. AI Engineer

Promtior • Córdoba

Presencial
ARS 135.892.000 - 226.487.000
20 paid days off per year
Flex Days: monthly team activities
Internal clubs
Software Engineer - AI
Software Engineer - AI

Beyond • Buenos Aires

Presencial
ARS 151.272.000 - 211.781.000
OSDE 210 for family group
Work from Home Allowance
Birthday leave
+2
Software Engineer - AI
Software Engineer - AI

Newfold Digital • Argentina

Presencial
ARS 181.094.000 - 271.641.000
Ai Platform Context Lead Id84742
Ai Platform Context Lead Id84742

Agileengine • Buenos Aires

Híbrido
ARS 211.276.000 - 286.732.000
Professional growth
Competitive compensation
A selection of exciting projects
+1
Software Engineer, GTM AI - Python
Software Engineer, GTM AI - Python

Telnyx • Argentina

Presencial
ARS 900.000 - 1.500.000
Full-Stack Engineer — Systems Conductor
Full-Stack Engineer — Systems Conductor

Bold Business • Municipio de Rincón de los Sauces

Presencial
ARS 135.820.000 - 211.276.000
Siena - Platform Experience Engineer
Siena - Platform Experience Engineer

Silver.dev • Buenos Aires

Presencial
ARS 3.000.000 - 6.000.000
Equity or stock grants
Learning budget
AI-fluency by default