AI Infrastructure Lead Architect (All Genders)

Accenture DACH

Kronberg im Taunus

Vor Ort

EUR 120.000 - 180.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Accenture DACH seeks a Lead and Principal Infrastructure Architect to own end-to-end compute infrastructure design for large-scale AI/ML systems, including distributed training environments and accelerator-rich stacks. You translate business goals into scalable architectures with cost-efficiency, lead architecture reviews, set standards, and mentor engineers while partnering with clients to shape roadmaps.

You will deploy and operate CI/CD, AI monitoring, security, compliance, and capacity

Qualifikationen

  • Solid background in coding, building, monitoring, and troubleshooting AI/ML deployments.
  • Strong understanding of AI/ML concepts and infrastructure.
  • Experience with on-premise or public cloud AI infra deployment.
  • Proficiency in Python, Java, or C++ for systems work.
  • Ability to work in fast-paced, cross-functional environments.

Aufgaben

  • Own end-to-end architecture and design of optimized compute infra for AI/ML systems, including large-scale distributed training.
  • Develop and evaluate architecture alternatives across compute, networking, storage, orchestration, and model serving.
  • Lead architecture assessments, identify gaps, risks, bottlenecks, and optimization opportunities.
  • Drive architectural decision-making with transparent rationale and alignment to SLAs and standards.
  • Define and maintain the AI infra roadmap, capacity planning, and technology evolution.
  • Architect and optimize the full stack for performance, power, cost, and scalability.
  • Design and tune large-scale GPU clusters and distributed training systems.
  • Serve as authoritative AI infra expert on at least one hyperscaler cloud (AWS/Azure/GCP).
  • Design deployment, automation, and CI/CD strategies for reliable production releases.
  • Establish AI monitoring and observability, SLAs/SLOs, alerting, and cost tracking.
  • Integrate AI/ML systems into enterprise environments with security/compliance.
  • Lead capacity planning and cost modeling to balance performance and efficiency.
  • Collaborate with clients and teams to translate requirements into architecture and standards.
  • Set technical direction, standards, and best practices; mentor engineers and conduct reviews.

Kenntnisse

Python
Java
C++
AI/ML
Infrastructure design
Cloud architectures
Cost optimization

Ausbildung

Bachelor's degree in Computer Science or related field

Tools

Apache Airflow
Kubeflow
GPU clustering
CI/CD tooling

Jobbeschreibung

Job Description

YOU ARE As a Lead and Principal Infrastructure Architect, you own end-to-end responsibility for designing optimized compute infrastructure for large-scale AI and machine learning systems, including large-scale distributed training environments.

As the authority who translates business goals, SLAs, and client standards into infrastructure architectures that perform at scale while being deliberately engineered for cost‑efficiency, you weigh multiple viable solutions for any given problem—across compute, networking, storage, orchestration, and model serving—and make rational, well‑justified architectural decisions tailored to each client’s situation, constraints, and standards. You architect and optimize the full computational stack for performance, power, cost, and scalability; design and tune large‑scale GPU clusters and distributed training systems; and ensure infrastructure meets security, compliance, and regulatory requirements.

As the recognized AI infrastructure expert in at least one hyperscaler cloud (such as AWS, Azure, or Google Cloud), you bring authoritative knowledge of that platform’s AI/ML services, accelerators, networking, and cost levers, and apply it to deliver best‑in‑class solutions. Beyond design, you set technical direction and standards, lead and mentor engineers and architects, partner with clients and stakeholders to shape the infrastructure roadmap, and are ultimately accountable for delivering AI/ML infrastructure that meets business SLAs, controls cost, and scales to enterprise and frontier workloads.

The Work
  • Own the end-to-end architecture and design of optimized compute infrastructure for large-scale AI/ML systems, including large-scale distributed training environments, from concept through delivery.
  • Develop and evaluate architecture alternatives, weighing trade‑offs across compute, networking, storage, orchestration, and model serving to make rational, well‑justified decisions tailored to each client’s situation and standards.
  • Lead architecture assessments and reviews of existing and proposed environments, identifying gaps, risks, bottlenecks, and optimization opportunities, and recommending remediation.
  • Drive architectural decision‑making, documenting rationale, trade‑offs, and assumptions so decisions are transparent, defensible, and aligned with business SLAs and standards.
  • Define and maintain the AI infrastructure roadmap, planning capacity, scaling, and technology evolution in step with business and product goals.
  • Architect and optimize the full computational stack for performance, power, cost, and scalability, ensuring infrastructure meets business SLAs while being deliberately engineered for cost‑efficiency.
  • Design and tune large‑scale GPU clusters and distributed training systems, including accelerator selection, interconnect/networking, and storage for high‑throughput training workloads.
  • Serve as the authoritative AI infrastructure expert in at least one hyperscaler cloud (AWS, Azure, or GCP), applying deep knowledge of its AI/ML services, accelerators, networking, and cost levers.
  • Design deployment, automation, and CI/CD strategies for reliable, repeatable, and scalable releases of AI systems, models, and data pipelines into production.
  • Establish AI monitoring and observability strategy across InfraOps and MLOps, defining SLAs, SLOs, alerting, and performance/cost tracking, and driving continuous optimization.
  • Integrate AI/ML systems into enterprise environments, ensuring interoperability, security, compliance, and adherence to regulatory and client standards.
  • Lead capacity planning and cost modeling, forecasting compute needs and engineering cost‑efficiency into the architecture without compromising performance.
  • Collaborate with clients, stakeholders, and engineering teams to align infrastructure decisions with business outcomes, translating requirements into actionable architecture and standards.
  • Set technical direction, standards, and best practices, mentoring engineers and architects and leading design and code reviews across the team.
Education
  • Bachelor's Degree in Computer Science, Computer Engineering, or a related engineering field.
Basic Qualification
  • Solid background in coding, building, monitoring, troubleshooting applications of AI/ML models; selecting, designing and infrastructure for deploying and running them on premise or on public cloud.
  • Strong understanding of AI and machine learning as a subject.
  • Strong understanding of computing infrastructure as a subject, preferred knowledge of AI infrastructure.
  • Good proficiency in programming languages such as Python, Java, or C++.
  • Experience with data pipeline and workflow management tools (e.g., Apache Airflow, Kubeflow).
  • Strong problem‑solving skills and ability to work in a fast‑paced environment.
  • Excellent communication and collaboration skills.
  • Significant experience in AI/ML infrastructure engineering or related roles on a hyperscaler platform for deploying large‑scale solutions.
  • Proven experience in leading and managing AI projects and teams.
  • Strong project management skills, with the ability to manage multiple projects simultaneously.
  • Demonstrated experience in evaluating and selecting AI technologies and frameworks.
  • Ability to work with cross‑functional teams and drive project alignment.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Infrastructure Principal Architect (All Genders)
AI Infrastructure Principal Architect (All Genders)

Accenture DACH • Kronberg im Taunus

Vor Ort
EUR 180.000 - 240.000
AI Infrastructure Architect (All Genders)
AI Infrastructure Architect (All Genders)

Accenture • Kronberg im Taunus

Vor Ort
EUR 120.000 - 160.000
AI Infrastructure Architect (All Genders)
AI Infrastructure Architect (All Genders)

Accenture DACH • Kronberg im Taunus

Vor Ort
EUR 90.000 - 140.000
Director, AI Engineering
Director, AI Engineering

Jobtailor • Deutschland

Remote
EUR 140.000 - 200.000
Lead AI Architect
Lead AI Architect

Jobtailor • Deutschland

Remote
EUR 140.000 - 190.000
Senior Artificial Intelligence Engineer
Senior Artificial Intelligence Engineer

IPI Technolab • Deutschland

Vor Ort
EUR 90.000 - 130.000
Compute Solution Architect
Compute Solution Architect

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
AI Large Language Model (LLM) Technology Architect (All Genders)
AI Large Language Model (LLM) Technology Architect (All Genders)

Accenture DACH • Würzburg

Vor Ort
EUR 110.000 - 150.000
AI Engineer (all levels)
AI Engineer (all levels)

Secure Systems Engineering GmbH • Berlin

Hybrid
EUR 60.000 - 90.000
Flexible hybrid working
Comfortable travel policy
Continuous training programs
Staff AI Engineer
Staff AI Engineer

Jobtailor • Deutschland

Remote
EUR 120.000 - 180.000