AI/ML Engineer

Katalyst Data Management

Rio de Janeiro

Híbrido

BRL 613 000 - 920 000

Tempo integral

14 dias+
Gerador de candidaturas

Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.

Ultrapassa os filtros ATS

Resumo da oferta

Katalyst Data Management is seeking a Senior AI/ML Engineer to lead production AI capabilities and own architecture for AI projects moving from prototype to production. You will work with a small, high-velocity team to ship scalable AI solutions for global energy clients.

The role covers GPU infrastructure, data ingestion pipelines, and DevOps/platform operations across Rio de Janeiro, Houston, or Calgary. We value curiosity, pragmatic decisions, and delivery of reliable, production-ready

Qualificações

  • Experience with Elasticsearch/OpenSearch in production environments.
  • Strong familiarity with AWS services (EC2, S3) and exposure to AI platforms such as Bedrock.
  • Experience with monitoring/observability tools (OpenTelemetry, Prometheus, Grafana).
  • DevOps CI/CD pipelines with GitHub Actions or Jenkins.
  • Experience supporting NVIDIA GPU environments (CUDA, drivers, nvidia-smi).
  • Familiarity with OCR technologies and document processing pipelines.
  • Proficiency with Git and Node.js/npm ecosystems.
  • Familiarity with agent-assisted development tools.
  • Experience supporting production AI/data/search platforms in high-availability environments.
  • Exposure to large-scale data ingestion or document processing systems.

Responsabilidades

  • GPU Infrastructure & Server Management: monitor, maintain, and optimize ML infra including NVIDIA GPUs.
  • Manage GPU workload allocation across embedding, OCR, and inference pipelines.
  • Establish observability with health metrics, alerts, and uptime targets.
  • Identify improvements for efficiency, resources, and pipeline performance.
  • Diagnose hardware and system issues (thermal, NVIDIA Xid, drivers, power).
  • Data Ingestion & Pipeline Operations: operate and monitor large-scale ingestion (500GB+) pipelines.
  • Manage preprocessing workflows for LAS, DLIS, SEG-Y, NAV, PDF, TIFF formats.
  • Monitor OCR workflows and resolve bottlenecks in legacy document processing.
  • Develop dashboards, logging, and alerting for visibility and reliability.
  • Perform data validation and QA on ingested content.

Conhecimentos

Elasticsearch/OpenSearch
AWS (EC2 S3)
Monitoring/Observability
CI/CD (GitHub Actions/Jenkins)
NVIDIA GPU environments
OCR/document processing
Git/version control
Node.js / npm
Agent-assisted tools
Production AI/data platforms
Large-scale data ingestion

Formação académica

Bachelor's degree in Computer Science or related field

Ferramentas

Docker
GitHub Actions
Jenkins
OpenTelemetry
Prometheus
Grafana
Ansible
PostgreSQL

Descrição da oferta de emprego

Join the dynamic and collaborative team at Katalyst Data Management (KDM)! KDM is seeking an AI/ML Engineer to support the operation, optimization, and reliability of the infrastructure powering our AI innovation platform. This hands‑on role will help maintain GPU environments, monitor large‑scale data ingestion and document processing pipelines, support DevOps and platform operations, and contribute to scalable AI/ML solutions that deliver real value to global energy clients. The ideal candidate brings practical experience in systems administration, DevOps, MLOps, or software engineering; strong problem‑solving skills; and the curiosity and adaptability to learn new tools quickly in a fast‑moving technical environment.

  • Position Located in Houston, TX USA; Calgary, AB Canada; or Rio de Janeiro, Brazil
  • 8:00 a.m. – 5:00 p.m. Monday to Friday
  • Full‑Time position
  • Hybrid Schedule Availability
The Company

Katalyst Data Management (KDM) is a global leader in subsurface data management solutions for the energy industry. For over 30 years, we have helped oil and gas companies, national governments, and energy organizations maximize the value of their data through secure storage, quality management, digital delivery, and online data marketing services. Our industry‑leading iGlass™ platform provides customers with reliable, secure, 24/7 access to critical subsurface information, supported by robust system redundancy and data protection controls. Through innovation, technical expertise, and exceptional customer service, KDM continues to deliver trusted solutions to clients around the world.

Key Responsibilities and Accountabilities

Katalyst Data Management is seeking a Senior AI/ML Engineer to serve as a technical lead for segments of our AI innovation platform. This is a hands‑on, high‑impact role where you will own architecture and day‑to‑day development for AI projects that move from prototype to production.

You will work directly with the Director of Innovation and a small, high‑velocity team to build, ship, and iterate AI capabilities that deliver real value to leading oil and gas clients worldwide. This is not a research‑only role; it requires sound architectural judgment, practical execution, and the ability to deliver production‑ready solutions in a rapidly evolving technology landscape.

The position may be based in Rio de Janeiro, Brazil; Houston, Texas; or Calgary, Alberta. We value adaptability and intellectual curiosity over mastery of any single technology stack, and we are especially interested in candidates who have repeatedly learned new tools quickly, made thoughtful technical decisions under uncertainty, and shipped AI/ML solutions that work in production.

Key Responsibilities:
GPU Infrastructure & Server Management
  • Monitor, maintain, and optimize machine learning infrastructure, including test environments and production‑grade NVIDIA GPU systems (e.g., the Houston DGX environment).
  • Manage GPU workload allocation and balancing across services such as embedding generation, OCR processing, and inference pipelines.
  • Establish and maintain system observability, including health metrics, alerting frameworks, and uptime targets for AI platforms.
  • Identify and implement improvements to system efficiency, resource utilization, and pipeline performance.
  • Diagnose and resolve hardware and system‑level issues, including thermal constraints, NVIDIA Xid errors, driver compatibility, and power optimization.
Data Ingestion & Pipeline Operations
  • Operate and monitor large‑scale ingestion pipelines processing high‑volume datasets (500GB+), supporting thousands of documents in production environments.
  • Manage preprocessing workflows for domain‑specific data formats, including LAS, DLIS, SEG‑Y, NAV, PDF, and TIFF.
  • Monitor, troubleshoot, and optimize OCR workflows, with a focus on resolving bottlenecks in scanned legacy document processing.
  • Develop and maintain dashboards, logging, and alerting mechanisms to ensure pipeline visibility and reliability.
  • Perform data validation and quality assurance checks on ingested content; elevate anomalies and inconsistencies as required.
DevOps & Platform Support
  • Support CI/CD pipelines, automated testing frameworks, and release processes for the Katapult platform.
  • Assist in the administration and optimization of Elasticsearch/OpenSearch clusters, including index management and performance tuning.
  • Develop and maintain infrastructure documentation, including runbooks, troubleshooting guides, and operational playbooks.
  • Support and enhance containerized deployment environments using Docker and configuration management tools such as Ansible.
  • Partner with Senior AI/ML Engineers to support RAG pipeline development, testing, and operationalization.
  • Assist with benchmarking and performance evaluation of embedding models, search configurations, and retrieval workflows.
  • Contribute to internal automation, scripting, and tooling to improve team productivity and platform reliability.
  • Participate in agile ceremonies, including sprint planning, code reviews, and technical design discussions.
Skills & Qualifications Required:
  • Experience with Elasticsearch or OpenSearch administration in production environments.
  • Strong familiarity with AWS services (e.g., EC2, S3) and exposure to AI platforms such as Bedrock or similar.
  • Experience with monitoring and observability tools (e.g., OpenTelemetry, Prometheus, Grafana).
  • Working knowledge of CI/CD pipelines and tools such as GitHub Actions or Jenkins.
  • Experience supporting NVIDIA GPU environments (CUDA, drivers, nvidia‑smi).
  • Familiarity with OCR technologies and document processing pipelines.
  • Proficiency with Git/version control systems and exposure to Node.js / npm ecosystems.
  • Familiarity with agent‑assisted development tools (e.g., Codex, Claude Code).
  • Experience supporting production AI, data, or search platforms in high‑availability environments.
  • Exposure to large‑scale data ingestion or document processing systems.
  • Interest in or familiarity with the oil and gas / energy domain is an asset.
  • Strong English communication skills, both oral and written, are required.
Preferred Qualifications & Skills
  • Oil and gas industry experience or domain knowledge, including well data, seismic data, or regulatory documents.
  • Experience with GPU‑accelerated workloads and the NVIDIA ecosystem.
  • Background in OCR pipelines and large‑scale document processing.
  • Familiarity with vision‑language models for document understanding.
  • Experience with agent‑to‑agent or multi‑agent AI system patterns.
  • Patent, publication, or conference presentation experience in ML/AI.
Required Education and Experience
  • 1–3 years of professional experience in DevOps, systems administration, MLOps, or software engineering (or equivalent hands‑on experience)
  • Experience working in small, fast‑moving teams with a high degree of autonomy.
  • Basic understanding of machine learning concepts, including embeddings, inference, and model serving
  • Hands‑on experience with Linux workstations or servers
  • Experience deploying and managing containerized applications using Docker
  • Working knowledge of PostgreSQL or similar relational databases
  • Proficiency in scripting with Python and/or Bash for automation and monitoring.

This role is based in a professional office setting with routine use of computers and collaboration tools. Work is highly technical and team-oriented, involving close coordination with Cloud, Infrastructure, and Software Engineering teams. The environment emphasizes automation, scalability, and security, with frequent virtual meetings and occasional cross‑functional project work.

Physical Demands:

This is primarily a sedentary role, involving extended periods of computer work. Occasional tasks may include setting up equipment, organizing supplies, or preparing meeting spaces, which could require light lifting, bending, or standing as needed.

Position Type and Expected Hours of Work:

This is a full-time position. As a full-time position, the AI/ML Engineer is eligible to participate in benefits coverage offered to Katalyst employees based on geographical location.

Monday through Friday, 8:00 a.m. to 5:00 p.m., with occasional extended hours to support deployments or maintain service continuity

Travel:

Minimal travel may be required for on‑premises infrastructure support, implementation, or occasional team collaboration. If participation in an offsite meeting is requested, all related travel expenses will be covered or reimbursed, in accordance with the company’s policies and local regulations.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior AI/ML Engineer
Senior AI/ML Engineer

Katalyst Data Management • Rio de Janeiro

Híbrido
BRL 300 000 - 520 000
Software Developer, Full Stack
Software Developer, Full Stack

Katalyst Data Management LP • Rio de Janeiro

Híbrido
BRL 80 000 - 100 000
Senior Data Scientist
Senior Data Scientist

bp • São Paulo

Híbrido
BRL 180 000 - 360 000
Senior Platform Engineer - GCP (Proficiency in Spanish and English) - Brazil
Senior Platform Engineer - GCP (Proficiency in Spanish and English) - Brazil

Thoughtworksreferral • Brasil

Presencial
BRL 280 000 - 360 000
Data Engineering & Infrastructure Lead ID88014
Data Engineering & Infrastructure Lead ID88014

AgileEngine, LLC. • Rio de Janeiro

Presencial
BRL 613 000 - 1 073 000
Professional growth
Competitive USD-based compensation
Projects with Fortune 500 clients
+1
AI Data Engineer III
AI Data Engineer III

Rimini Street • Brasil

Presencial
BRL 180 000 - 240 000
Data Engineering & Infrastructure Lead ID88014
Data Engineering & Infrastructure Lead ID88014

AgileEngine, LLC. • Salvador

Híbrido
BRL 150 000 - 230 000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Senior MLOps Engineer - Remote - Latin America
Senior MLOps Engineer - Remote - Latin America

FullStack • Manaus

Teletrabalho
BRL 240 000 - 300 000
Competitive pay
100% remote work
Work with leading startups and Fortune
+2
Teach Lead Data Engineer
Teach Lead Data Engineer

Jobgether • Brasil

Presencial
BRL 280 000 - 520 000
Career growth opportunities
Collaborative environment
Global data/tech exposure
Technical Lead
Technical Lead

AgileEngine • Brasil

Presencial
BRL 614 000 - 922 000
Professional growth
Competitive compensation
Exciting projects
+1