Principal Data Engineer, LLM/AI Platforms

Jobgether SRL

United States

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flexible remote work opportunities
Health and wellness programs
Equity opportunities
Parental leave

Job summary

Exabyte seeks a Principal Data Engineer to architect and scale data platforms powering next‑gen LLM/AI platforms. You will lead data pipelines, modeling, and MLOps while mentoring engineers in a highly autonomous setting.

The role blends hands‑on engineering with platform architecture, focusing on performance, resilience, and cost efficiency at massive scale. Collaboration with data scientists and product managers is essential.

Qualifications

  • Master’s or PhD in Computer Science, Data Engineering, or related STEM field or equivalent practical experience.
  • 10+ years in Data or Platform Engineering, incl. AI/ML/data science platforms at massive scale.
  • 3+ years in Principal/Staff-level engineering with leadership and mentorship.
  • Hands-on LLM engineering: fine-tuning, prompt engineering, deployment, RAG, agentic workflows.
  • Designing large-scale distributed systems with sharding, partitioning, concurrency, fault tolerance.
  • Proficient in Python or JVM-based tech, writing clean, production-grade code.
  • Experience with Spark, Dask, Flink for distributed data processing.
  • Strong knowledge of AWS/GCP/OCI data services.
  • Docker and Kubernetes expertise.
  • Kafka or Pulsar for messaging/streaming.
  • Snowflake, BigQuery, Airflow, Kubeflow data orchestration.
  • MLOps tools: MLflow, SageMaker, Vertex AI.
  • Familiar with LangChain or LlamaIndex for agentic AI.
  • Solid engineering practices: reviews, testing, secure development.
  • Experience applying AI to automate workflows and drive business value.
  • Excellent communication and cross-team collaboration.
  • Direct experience deploying LLMs in production is a plus.
  • Background in cybersecurity/regulatory industries is a plus.
  • Open-source contributions a plus.

Responsibilities

  • Architect, implement, and optimize data platforms and pipelines for LLMs, RAG, and agentic systems at exabyte scale.
  • Drive adoption of agentic workflows to enable autonomous, data-driven capabilities.
  • Design scalable, fault-tolerant, secure, cost-efficient data solutions for rapid iteration.
  • Develop production-ready, well-tested code with emphasis on performance and reliability.
  • Lead data modeling, semantic cataloging, and architecture for AI/ML workloads.
  • Establish MLOps and DataOps practices for observability and recovery.
  • Own end-to-end lifecycle of critical data services from development to deployment.
  • Collaborate with researchers, product managers, and engineers to productionize research prototypes.
  • Lead workshops, reviews, and mentoring to raise technical standards across AI platforms.
  • Champion DevSecOps and secure development across distributed data environments.
  • Identify opportunities to improve platform performance and developer productivity.

Skills

Python
JVM
Distributed systems
MLOps
Cloud platforms
Kafka/Pulsar
Data modeling
Leadership
Mentorship
DevSecOps

Education

Master’s degree or PhD in CS / related

Tools

Spark
Dask
Flink
Snowflake
BigQuery
Airflow
Kubeflow
MLflow
SageMaker
Vertex AI
Docker
Kubernetes
Kafka

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Data Engineer, LLM/AI Platforms based in United States.

This is a principal-level engineering role focused on building the data infrastructure that powers next-generation LLM and AI platforms at massive scale.

You will architect and optimize data platforms supporting LLMs, Retrieval-Augmented Generation (RAG), and sophisticated agentic systems.

The role combines deep hands‑on engineering with technical leadership, platform architecture, and continuous innovation.

You will work with extremely large datasets and distributed systems while emphasizing scalability, resilience, performance, and cost efficiency.

A major focus is turning advanced AI and data science concepts into reliable, production‑grade services.

You will collaborate with data scientists, product managers, and engineering teams while mentoring engineers and raising technical standards.

This is an opportunity to shape AI platform engineering practices across critical, large‑scale systems in a highly autonomous environment.

Accountabilities
  • Architect, implement, and optimize data platforms and pipelines designed for LLMs, RAG, and advanced AI agentic systems at Exabyte scale.
  • Drive the adoption and deployment of agentic workflows and agent‑harnessing techniques to support autonomous, data‑driven capabilities.
  • Design highly scalable, fault‑tolerant, secure, and cost‑effective data solutions that enable rapid iteration without compromising engineering quality.
  • Develop production‑ready code with strong attention to performance, maintainability, testing, and operational reliability.
  • Provide technical leadership in data modeling, normalization, semantic cataloging, and data architecture for AI and machine learning workloads.
  • Establish MLOps and DataOps best practices for LLM platforms, including monitoring, observability, automated recovery, and service reliability.
  • Own the end‑to‑end lifecycle of critical data services, including development, testing, deployment, monitoring, and continuous optimization.
  • Collaborate with data scientists, product managers, and engineering teams to transform research prototypes into robust, production‑ready services.
  • Lead technical workshops, design reviews, and knowledge‑sharing initiatives while mentoring engineers and strengthening organizational expertise in AI platform technologies.
  • Champion DevSecOps practices and engineering standards across large‑scale distributed data environments.
  • Identify opportunities to improve platform performance, reliability, scalability, and developer productivity through new technologies and engineering practices.
Requirements
  • Master’s degree or PhD in Computer Science, Data Engineering, or a related STEM discipline, or equivalent practical experience.
  • 10+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms at massive scale.
  • 3+ years of experience in a Principal or Staff‑level engineering capacity, with demonstrated technical leadership and mentorship experience.
  • Hands‑on expertise with LLM engineering, including fine‑tuning, prompt engineering, deployment, RAG, and agentic workflow development.
  • Proven experience designing and delivering large‑scale distributed systems, including sharding, partitioning, concurrency, and fault‑tolerant architectures.
  • Expert‑level proficiency in Python or JVM‑based technologies, with a strong ability to write clean, performant, maintainable, and well‑tested production code.
  • Deep experience with distributed data processing frameworks such as Spark, Dask, or Flink.
  • Strong knowledge of cloud platforms such as AWS, GCP, or OCI and their associated data services.
  • Expertise with containerization and orchestration technologies including Docker and Kubernetes.
  • Experience with messaging and streaming technologies such as Kafka or Pulsar.
  • Familiarity with data warehousing and orchestration platforms such as Snowflake, BigQuery, Airflow, and Kubeflow.
  • Experience with MLOps technologies such as MLflow, SageMaker, or Vertex AI.
  • Familiarity with agentic AI frameworks such as LangChain or LlamaIndex.
  • Strong understanding of engineering practices including peer code reviews, resilient architecture, comprehensive testing, and secure development methodologies.
  • Demonstrated ability to use AI technologies to improve decision‑making, automate workflows, increase efficiency, and support measurable business outcomes.
  • Strong communication and collaboration skills, with the ability to influence technical direction and mentor engineers across teams.
  • Direct experience deploying and managing LLMs in production is a plus.
  • Experience in cybersecurity, intelligence, or highly regulated industries is a plus.
  • Contributions to open‑source data or AI/ML projects are a plus.
Benefits
  • CAD $210,000–$320,000 annual base salary for Canadian‑based employment, plus variable/incentive compensation, equity, and benefits.
  • Flexible remote work opportunities.
  • Comprehensive health and wellness programs supporting physical and mental wellbeing.
  • Competitive vacation and holiday programs to support time away and recharge.
  • Paid parental and adoption leave.
  • Professional development and continuous learning opportunities at all career levels.
  • Employee networks, geographic communities, and volunteer opportunities to build professional connections.
  • Opportunities to work on large‑scale AI, data engineering, and cybersecurity technologies.
  • Equity opportunities as part of the overall compensation package.
  • Retirement and financial benefits as applicable.
  • Inclusive workplace practices and support for employees with disabilities.
  • Canadian employment requires legal entitlement to work in Canada and may include applicable background checks.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer – AI/Machine Learning
Lead Data Engineer – AI/Machine Learning

Core Specialty • Cincinnati (OH)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision, and life ins.
Disability insurance
401(k) company-match
+5
Senior Software Engineer - Cloud Platform Engineering
Senior Software Engineer - Cloud Platform Engineering

Jobgether SRL • United States

Remote
USD 141,000 - 185,000
Remote work across United States
Equity opportunities
Health insurance
+5
Data Platform Engineer - 90/HR- REMOTE
Data Platform Engineer - 90/HR- REMOTE

ContractStaffingRecruiters.com • Branford (CT)

Remote
USD 130,000 - 160,000
Principal Data Engineer, LLM/AI Platforms (Remote)
Principal Data Engineer, LLM/AI Platforms (Remote)

CrowdStrike • United States

On-site
CAD 210,000 - 320,000
Market-leading compensation
Comprehensive wellness programs
Professional development opportunities
Lead Data Engineer – AI/Machine Learning
Lead Data Engineer – AI/Machine Learning

Core Specialty Insurance Holdings, Inc. • Cincinnati (OH)

Hybrid
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+5
Staff AI Engineer
Staff AI Engineer

Robots & Pencils • California (MO)

On-site
USD 125,558 - 173,239
[Job-31573] AI Engineer Master
[Job-31573] AI Engineer Master

ciandt • United States

Hybrid
USD 180,000 - 270,000
Health insurance
Dental insurance
Life insurance
+8
Gen AI Engineer - Full Time
Gen AI Engineer - Full Time

NeerInfo Solutions Pvt. Ltd. • Plano (TX)

On-site
USD 120,000 - 160,000
Long-term Disability
Health and Dependent Care Reimb. A/cs
Insurance (Accident, Critical Illness,
+1
AI Engineer
AI Engineer

DataJobs • Washington

On-site
USD 160,000 - 190,000
Hybrid work environment
Tuition reimbursement or training stip
Career development support
Staff Software Engineer, Data (AI)
Staff Software Engineer, Data (AI)

Juniper Square • United States

Remote
USD 210,000 - 260,000
Health, dental, and vision
Flexible time off
Annual development stipend