Erhalte mehr Antworten von Arbeitgebern
Versende in nur wenigen Minuten einen passgenauen Lebenslauf.
Cortea AI in Berlin is seeking a Senior/Staff Platform Engineer to own the platform architecture and core stack, choosing technologies and guiding implementation from design to production.
You will work across product engineers, DevOps, and SRE to build reliable, observable systems, with emphasis on security, compliance, and scalable AI tooling.
We ship fast, and we intend to keep shipping fast. The consequence is that our system is outgrowing the amount of deliberate design that went into it. We run AI agents over customer documents in a domain where being wrong is expensive, and a lot of the load-bearing decisions about how those fit together are still implicit. You will be making those decisions explicit and then making them hold. That means codifying golden paths, building the shared primitives and APIs that product engineers work on top of, owning the architecture of the system as a whole, and establishing what our reliability actually needs to be before an incident establishes it for us. Platform at Cortea sits between software engineering, DevOps and SRE. Its customers are other engineers. This is not a DevOps role under a different name: you will spend more time in application code than in YAML, and the reason we want infrastructure experience is that we don't believe you can design a system well without being able to run it. The bar we are hiring against is someone who has taken a system from nothing to production and owned it end to end. You picked the technology, argued the tradeoffs in writing, provisioned the infrastructure, and were responsible for it when it broke.
We use AI heavily and we want you to. What we are not looking for is someone who outsources their judgment to it. The mental model of this system has to live in your head, not in a context window. Use agents to move quickly on implementation. Do the design, the reasoning and the writing yourself, and be able to defend every decision you ship.
Four areas, roughly in the order you will spend time on them. You won't work on all of them at once, but you should be open to any of them.
A system of record for managing and distributing constantly evolving agent configurations, under strict auditability and tenant isolation requirements.
Abstracting over use cases that change faster than the code.
That holds up across a growing number of parallel agent executions.
With quotas and rate limiting across a growing set of products.
Execution latency, AI spend, memory, CPU), then fixing the bottlenecks they expose at both the application and infrastructure level.
Handling dozens of file formats, very large spreadsheets and hundreds of parallel uploads, under a strong reliability requirement.
Have designed, built and operated distributed…