Senior DevOps / Platform Engineer, AI Infrastructure (m/f/x)

Join

United States

Remote

USD 86,000 - 98,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

WellPass fitness membership
VSOP program participation
Travel and accommodation to Leipzig (€

Job summary

MAIA is seeking a Senior DevOps/Platform Engineer to take technical ownership of our production infrastructure, deploying and evolving Linux-based systems, containers, and API gateways across multi-cloud environments. You will drive automation, CI/CD, and observability, while ensuring security and ISO 27001 readiness.

You will work closely with software engineers and leadership in a fully remote German setup, with future expansion into European infrastructure and AI inference options.

Qualifications

  • Several years of experience operating production SaaS systems on Linux servers.
  • Experience with Docker and Docker Compose in production and knowledge of reverse proxies/API gateways like Traefik or Kong.
  • Practical PostgreSQL experience including backups, restores, performance analysis, and pooling.
  • Built and maintained CI/CD pipelines using GitHub Actions, GitLab CI, or similar tools.
  • Managed infrastructure with Terraform, Pulumi, or comparable IaC tooling.
  • Operated an observability stack (Grafana, Loki, Prometheus, Sentry, etc.).
  • Understanding of security foundations: IAM, least privilege, secrets management, vulnerability scanning, patching.
  • Strong English communication; German is a plus but not required.
  • Knowledge of LLM serving trade-offs and AI infrastructure concepts.

Responsibilities

  • Take technical ownership of MAIA’s production infrastructure.
  • Make deployments safe, repeatable, and engineer-friendly with IaC and CI/CD improvements.
  • Operate and optimize PostgreSQL in production, including backups and capacity planning.
  • Develop and improve a self-hosted observability stack with metrics, logs, and traces.
  • Strengthen security across infrastructure, including IAM, secrets management, TLS, and patching.
  • Design ISO 27001 controls with auditable evidence generation.
  • Contribute to AI infrastructure, evaluating model and inference providers, latency, and cost.
  • Explore self-hosted LLM inference options where valuable, including GPU setups.
  • Improve incident response, document playbooks, and enhance operational resilience.
  • Improve internal developer experience with clear interfaces and automated workflows.
  • Identify cost drivers and support data-driven build, buy, and hosting choices.

Skills

Linux administration
Docker in production
API gateways
PostgreSQL
CI/CD pipelines
Terraform / IaC
Observability stack
Security fundamentals
LLM serving
English communication

Tools

Kubernetes
GitHub Actions

Job description

MAIA is the AI platform built for companies where generic AI tools break down because the data is complex, the stakes are high, and precision matters.

We work on the difficult parts of enterprise AI: understanding complex documents, integrating AI into real workflows, making organisational knowledge accessible, and operating the underlying systems reliably and securely.

As our customer base, product capabilities, and engineering team grow, we are investing further in the platform beneath MAIA. We are looking for our first dedicated Senior DevOps / Platform Engineer to take technical ownership of our infrastructure and production operations.

You will develop an existing platform and shape its next stage across architecture, automation, observability, security, and AI infrastructure. Our environment combines self-administered Linux servers, European infrastructure providers, dedicated hardware, and services from AWS, Azure, and GCP. Over time, we will also expand our European infrastructure and inference options as one part of our platform strategy.

This is a hands‑on senior individual contributor role without people management. You will initially be the only dedicated Platform Engineer, working closely with our software engineers and company leadership.

We are opening this role to remote candidates across Germany. For this setup to work, you need a strong track record of independently operating production systems, communicating proactively, and moving complex infrastructure work forward without constant coordination.

Tasks
  • You take technical ownership of MAIA’s production infrastructure. You operate and evolve self-administered Linux systems across virtual machines and dedicated servers, including containers, networks, reverse proxies, and API gateways.
  • You make deployments safe, repeatable, and easy for engineers to use. You improve our Infrastructure as Code, GitHub Actions workflows, deployment processes, automated checks, versioning, and rollback capabilities.
  • You operate and improve PostgreSQL in production. This includes performance analysis, connection pooling, capacity planning, backups, and regularly tested restore procedures.
  • You develop our self-hosted observability stack. You connect metrics, logs, traces, and actionable alerts so that we understand system behaviour and identify problems before they affect more customers.
  • You strengthen security across our infrastructure. You implement and maintain IAM, least privilege, secrets management, TLS, vulnerability scanning, and patch management.
  • You implement technical controls for ISO 27001. You design them so that evidence is generated continuously and remains understandable and auditable.
  • You help shape our AI infrastructure. You integrate model and inference providers and evaluate them based on reliability, latency, throughput, cost, and operational effort.
  • You explore self-hosted LLM inference where it creates real value. Over time, this may include GPU infrastructure and serving technologies such as vLLM. Previous production experience in this area is helpful but not required.
  • You improve incident response and operational resilience. You investigate root causes, establish useful runbooks, document operational knowledge, and turn incidents into lasting improvements.
  • You improve the internal developer experience. You reduce manual work, create clear interfaces and workflows, and help engineers ship changes with confidence.
  • You make infrastructure and inference costs transparent. You identify relevant cost drivers and help us make informed build, buy, and hosting decisions.

You will help determine the initial priorities after assessing the existing platform. We expect you to identify the most relevant risks, explain the available options, and take improvements through to reliable production operation.

Requirements

Your technical experience

  • You have several years of experience operating production SaaS systems on Linux servers you or your team administered directly. Experience limited to fully managed hyperscaler services is not sufficient for this role.
  • You have operated Docker and Docker Compose in production and understand reverse proxies or API gateways such as Traefik or Kong.
  • You have practical experience operating PostgreSQL, including backups and restores you have personally configured and tested, performance analysis, and connection pooling.
  • You have built and maintained CI/CD pipelines using GitHub Actions, GitLab CI, or comparable systems.
  • You have managed infrastructure through Terraform, Pulumi, or comparable Infrastructure as Code tooling.
  • You have operated an observability stack using tools such as Grafana, Loki, Prometheus, Sentry, or comparable technologies.
  • You understand the security foundations of production infrastructure, including IAM, least privilege, secrets management, vulnerability scanning, and patching.
  • You can design infrastructure that is reliable for customers and straightforward for engineers to use.
  • You communicate fluently in English. German is helpful but not required.

Your AI foundation

You understand how RAG systems work and where they commonly fail in production. You can discuss the basic trade‑offs involved in LLM serving, including latency, throughput, reliability, cost, and operational complexity.

You follow developments in models, inference providers, and the wider GenAI market closely enough to evaluate new options critically.

We do not expect you to arrive as an expert in GPU operations or self-hosted inference. We expect a strong platform engineering foundation and the technical curiosity to develop deeper expertise in AI infrastructure.

How you work
  1. You take responsibility from initial investigation through production operation.
  2. You investigate unfamiliar systems until you understand the relevant behaviour and failure modes.
  3. You make architecture decisions deliberately and document the reasoning behind them.
  4. You value maintainability, automation, tested recovery, and clear operational processes.
  5. You prioritise according to risk and impact and communicate trade‑offs openly.
  6. You remain structured during incidents and communicate clearly with technical and non‑technical stakeholders.
  7. You use AI tools productively while remaining accountable for every architecture decision and production change.
  8. You communicate proactively in a distributed team and make your progress, decisions, risks, and blockers visible.
The bar for this remote role

This position combines broad technical responsibility with a high degree of autonomy. For the remote setup, we are therefore looking for candidates who have already demonstrated that they can:

  • independently operate business‑critical production systems;
  • identify and prioritise infrastructure work without waiting for detailed instructions;
  • lead technical improvements across team boundaries;
  • communicate reliably in writing and remotely; and
  • remain accountable after a change has been deployed.

Interest in growing into these responsibilities is valuable, but for this particular opening, we need evidence that you have already carried a substantial part of them in practice.

Helpful, but not required
  • Experience with Hetzner, NixOS, or self-hosted Supabase.
  • Experience implementing technical ISO 27001 controls and audit evidence.
  • Experience with SRE practices such as SLOs, error budgets, and incident reviews.
  • Experience operating GPUs or serving LLMs with technologies such as vLLM.
Benefits
  • A substantial technical area to shape and own in a well‑funded, growing AI startup.
  • Direct influence on MAIA’s architecture, reliability, security, and engineering productivity.
  • The opportunity to deepen your expertise in AI infrastructure and the operation of LLM workloads.
  • Close collaboration with our engineering team, CTO, and company leadership.
  • Short decision paths and the authority to move important infrastructure work forward.
  • €75,000–€85,000 gross annual salary, depending on experience and scope.
  • The opportunity to participate in our VSOP.
  • A permanent, full‑time position with flexible working hours.
  • Fully remote work from anywhere within Germany.
  • Regular opportunities to meet and work with the team in Leipzig, with travel and accommodation covered.
  • Access to a WellPass fitness membership.

This role is open exclusively to candidates currently based in Germany. Applications from outside Germany cannot be considered for this position.

Our hiring process consists of three stages:

  1. A 30‑minute introductory conversation.
  2. A 90‑minute technical interview.
  3. A final conversation with company leadership.

We move quickly and keep you informed throughout the process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Engineer (f/m/x)
Senior AI Engineer (f/m/x)

Lever, Inc. • Germany (OH)

Remote
USD 136,000 - 181,000
Fully remote-first environment
Vienna office option
Flexible hours
+5
Senior AI Engineer (f/m/x)
Senior AI Engineer (f/m/x)

Lever, Inc. • Town of Italy (NY)

Remote
USD 136,000 - 204,000
Fully remote-first
Vienna office option
Flexible hours
+2
Senior Backend Engineer: Machine Learning Infrastructure
Senior Backend Engineer: Machine Learning Infrastructure

Lever, Inc. • Town of Italy (NY)

On-site
EUR 71,000 - 106,000
Fully remote work environment
Unlimited vacation
Home-office stipend
+3
Senior Forward Deployed Platform Engineer (m/w/d)
Senior Forward Deployed Platform Engineer (m/w/d)

United States Digital Space LLC • United States

Hybrid
USD 180,000 - 240,000
Edenred card
Extra vacation day
Applied AI Engineer – Systems & Reliability (remote/Berlin-based)
Applied AI Engineer – Systems & Reliability (remote/Berlin-based)

HiPeople • United States

Hybrid
USD 102,000 - 159,000
Stock options
Educational stipend
Remote/On-site in Berlin
+2
Senior Backend Engineer: Machine Learning Infrastructure
Senior Backend Engineer: Machine Learning Infrastructure

Lever, Inc. • Germany (OH)

Remote
EUR 71,000 - 106,000
Fully remote work
Unlimited vacation
Home-office stipend
+5
AI Engineer (m/f/d)
AI Engineer (m/f/d)

United States Digital Space LLC • United States

On-site
USD 140,000 - 210,000
Senior AI Engineer (f/m/x)
Senior AI Engineer (f/m/x)

Almaz Capital • United States

Remote
USD 113,000 - 169,000
Learning budget
Private coaching sessions with experts
Office weeks in Vienna
Senior DevOps & Platform Engineer — Remote (Germany)
Senior DevOps & Platform Engineer — Remote (Germany)

Join • United States

Remote
USD 86,000 - 98,000
WellPass fitness membership
VSOP program participation
Travel and accommodation to Leipzig (€
Staff Platform Engineer – Infrastructure
Staff Platform Engineer – Infrastructure

Recruiting from Scratch • United States

On-site
USD 200,000 - 300,000
20 days PTO annually
100% covered health insurance
401(k) with employer contribution
+1