Senior Software Engineer | Kubernetes Automation | EU

Creandum

Deutschland

Remote

EUR 73.000 - 100.000

Vollzeit

Vor 6 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, fertig zum Versenden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Equity options
Remote-first global environment
Learning budget

Zusammenfassung

Cast AI is seeking a senior Go/DevOps engineer to design and operate autonomous Kubernetes infrastructure. You will build real‑time cloud resource management services, move workloads between nodes without disruption, and optimize costs across clouds. The role demands deep systems knowledge, production intensity, and a startup mindset.

Join a global, remote-first team delivering cutting-edge automation that reduces manual tuning and tickets while scaling AI and cloud-native infrastructure.

Qualifikationen

  • Strong software engineering fundamentals with clean abstractions.
  • Go production experience or strong systems programming as alternative.
  • Understand Kubernetes autoscaling and networking.
  • Led a complex project end-to-end in production.
  • Experience with AI coding agents and agentic workflows.
  • Proficient in observability tooling in production environments.
  • CI/CD and DevOps practices experience.
  • Strong English communication skills.
  • Startup mindset with adaptability and initiative.

Aufgaben

  • Design and build distributed systems automating Kubernetes infrastructure at scale.
  • Develop production Go services interacting with cloud provider APIs.
  • Own features end-to-end from design to production rollout (1–4 weeks).
  • Debug complex production issues across clouds and clusters.
  • Collaborate with product and engineering teams to solve non‑textbook problems.
  • Handle time-series data and Kubernetes control plane internals.
  • Contribute to product direction and backlog prioritization.
  • Participate in on-call rotation as needed.

Kenntnisse

Go programming
Kubernetes internals
End-to-end project lead
Observability in prod
CI/CD/DevOps
English proficiency
Startup mindset
AI coding agents

Tools

AWS
GCP
Azure
Karpenter
Terraform

Jobbeschreibung

Why Cast AI?

Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale. By embedding autonomous decision-making directly into Kubernetes and cloud environments, Cast AI continuously optimizes performance, reliability, and efficiency in production.

The old way doesn’t work. As Kubernetes and AI environments grow, manual decisions don’t. Cast AI replaces tickets, alerts, and manual tuning with continuous automation that adapts infrastructure as conditions change. Efficiency and cost savings follow naturally from that automation.

Over 2,100 companies already rely on Cast AI, including Akamai, BMW, Cisco, FICO, HuggingFace, NielsenIQ, Swisscom, and TGS.

Global team, diverse perspectives
We’re headquartered in Miami, but our impact is international. We take a global and intentional approach to diversity. Today, Cast AI operates across 34 countries spanning Europe, North America, Latin America, and APAC, bringing a wide range of perspectives into how we build and lead.

Unicorn momentum
In January 2026, we achieved unicorn status with a strategic investment from Pacific Alliance Ventures, the corporate venture arm of Shinsegae Group (a $50+ billion Korean conglomerate). Our valuation now exceeds $1 billion, and we’re just getting started.

Join us as we build the future of autonomous infrastructure.

About the role
What you’ll be solving

A pod is mid-request. The node it’s running on is about to disappear – Spot is pulling the instance, or Autoscaler has flagged the node as underutilized and marked it for removal. The safe move is to wait. The profitable move is to act now. Every team in this pipeline is built to close that gap: make the aggressive call and still guarantee nothing breaks.

That guarantee is the hard part. Deciding which nodes to kill, which pods to move, and which storage volume to resize is a genuinely hard optimization problem – several of them are P=NP-hard in the general case – and it has to run continuously, in production, faster than a human would ever attempt it by hand. Get it right and clusters run at half the cost. Get it wrong once and you’ve taken down something live.

This posting covers multiple teams working on different pieces of the same automation problem node placement algorithms and provisioning, deep integrations with tools like Karpenter, workload rightsizing (vertical and horizontal), moving running workloads between nodes without interrupting them, storage that scales itself, and GPU capacity optimization across clouds and regions. All of it runs in real time, in production, with no human tuning YAML in the loop.


This is a location-specific opportunity. We are currently accepting applications from candidates residing in the following European countries: Bulgaria, Croatia, Estonia, Greece, Hungary, Latvia, Lithuania, Poland, Romania, Slovakia, Slovenia, and Ukraine.

Requirements:
  • Strong software engineering fundamentals: you design clean abstractions, write maintainable code, and can justify your architecture choices – not just make something work.

  • Production experience with Go is strongly preferred; candidates without Go should demonstrate strong systems programming skills in a comparable language.

  • Understanding of Kubernetes internals – autoscaling and networking.

  • You’ve personally driven a complex project end-to-end.

  • Experience with AI coding agents and agentic development workflows.

  • You’ve used observability tooling in production.

  • CI/CD and DevOps practices experience.

  • Strong English skills, both verbal and written.

  • Startup mindset: adaptable, proactive, and comfortable with ambiguity.

Nice-to-haves:
  • Deep hands-on experience with cloud platforms (AWS, GCP, or Azure) – including real understanding of how compute, networking, and storage work under the hood.

  • Deep knowledge of EKS, GKE, or AKS internals.

  • Container runtime experience CRI-O, runc, containerd.

  • Cloud storage depth – block storage, CSI, volume management, filesystem internals.

  • Low-level Linux experience: process internals, packet-level networking (NAT, iptables, conntrack, eBPF, SDN), filesystems, storage.

  • Low-level systems programming experience in C or C++.

  • Kubernetes or cloud-native OSS contributions.

  • Experience with AI coding agents and agentic development workflows.

Responsibilities:
  • Design and build distributed systems that operate Kubernetes infrastructure autonomously at scale.

  • Write production Go services that interact with AWS, GCP, and Azure APIs for real-time cloud resource management.

  • Own features end-to-end: from design through implementation, testing, and production rollout (most projects ship in 1-4 weeks).

  • Debug complex production issues across cloud providers, Kubernetes clusters, and distributed services.

  • Collaborate with product and other engineering teams to solve problems that don’t have textbook solutions.

  • Work with time-series data, cloud provider APIs and Kubernetes control plane internals.

  • Contribute to product direction and backlog – shape what gets built next, not just how it’s built.

  • Participate in on-call rotation for teams that require it.

Tools we use daily:
  • Programming Languages: Go

  • Cloud & Orchestration: Kubernetes, AWS, GCP, Azure

  • Infrastructure as Code: Terraform

  • Databases & Storage: PostgreSQL, Cloud Object Storage, ClickHouse

  • Messaging & APIs: GCP Pub/Sub, gRPC for internal communication, REST for public APIs

  • Observability: Prometheus, Grafana, Loki, Tempo

  • CI/CD & GitOps: GitLab/Github CI with ArgoCD and Kargo

What’s in it for you?
  • Competitive salary (€6,500 – €9,000 gross, depending on the level of experience)
  • Enjoy a flexible, remote-first global environment.
  • Collaborate with a global team of cloud experts and innovators, passionate about pushing the boundaries of Kubernetes technology.
  • Equity options.
  • Get quick feedback with a fast-paced workflow. Most feature projects are completed in 1 to 4 weeks.
  • Spend 10% of your work time on personal projects or self-improvement.
  • Learning budget for professional and personal development – including access to international conferences and courses that elevate your skills.
  • Annual hackathon to spark new ideas and strengthen team bonds.
  • Team-building budget and company events to connect your colleagues.
  • Equipment budget to ensure you have everything you need.
  • Extra days off to help maintain a healthy work-life balance.
Hiring process
  • Screening call with Recruiter
  • Hiring Manager interview
  • Technical interview (system design)
  • Live coding
  • Culture Check interview with an executive

*As part of our standard hiring process, we would like to inform you that a background check may be conducted at the final stage of recruitment through our third-party provider, Checkr.
*Please note that Cast AI does not provide any form of visa sponsorship/work permit.

#LI-Remote

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Software Engineer | Kubernetes Automation
Senior Software Engineer | Kubernetes Automation

Creandum • Deutschland

Remote
EUR 73.000 - 100.000
Equity options
Remote-friendly environment
Learning budget
+4
Senior Software Engineer | Platform Engineering
Senior Software Engineer | Platform Engineering

Creandum • Deutschland

Remote
EUR 78.000 - 108.000
Remote-first environment
Equity options
Learning budget
+4
Senior Talent Acquisition Partner
Senior Talent Acquisition Partner

Creandum • Deutschland

Remote
EUR 45.000 - 56.000
Remote-friendly
Equity options
Learning budget
+3
Senior Software Engineer | Kimchi (Agent Platform Team)
Senior Software Engineer | Kimchi (Agent Platform Team)

Creandum • Deutschland

Remote
EUR 73.000 - 100.000
Remote-first
Equity options
Learning budget
+2
Senior Software Engineer | Kimchi (Harness Team)
Senior Software Engineer | Kimchi (Harness Team)

Creandum • Deutschland

Remote
EUR 73.000 - 100.000
Equity options
Remote-first global environment
Learning budget for conferences and发展
+4
Senior Software Engineer | Kimchi Studio
Senior Software Engineer | Kimchi Studio

Creandum • Deutschland

Remote
EUR 73.000 - 100.000
Competitive salary
Remote-friendly environment
Equity options
+2
Senior/Staff Platform Engineer (m/f/x)
Senior/Staff Platform Engineer (m/f/x)

Cortea • Berlin

Hybrid
EUR 90.000 - 140.000
Equity
Competitive salary
Autonomy
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Vcluster • Deutschland

Remote
EUR 90.000 - 130.000
Competitive Salary
Premium Insurance
Flexible Working Schedule
+1
Platform Engineer Turn complex infrastructure into simple, dependable platform primitives. Remote (CET ±1h) · Full-time →
Platform Engineer Turn complex infrastructure into simple, dependable platform primitives. Remote (CET ±1h) · Full-time →

Interloom Ltd • Deutschland

Remote
EUR 90.000 - 130.000
Competitive salary
VSOP stock options
Choose workstation and hardware
+1
Engineering Team Lead, Managed K8s
Engineering Team Lead, Managed K8s

Mistral • München

Vor Ort
EUR 140.000 - 210.000
Healthcare coverage
Relocation support
Meal and transportation allowances