Senior Software Engineer | Kubernetes Automation | EU

Creandum

United States

Remote

USD 7,300 - 10,000

Full time

47 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity options
Remote-first environment
Learning budget
Annual hackathon
Team events
Equipment budget
Extra days off

Job summary

Cast AI is seeking a senior Go engineer to design and build distributed systems that autonomously manage Kubernetes infrastructure at scale. You will write production-grade Go services interacting with AWS, GCP, and Azure APIs for real-time cloud resource management.

You will own features end-to-end, collaborate across teams, and tackle on-call rotations while advancing cutting-edge cloud-native automation in a remote-first, globally distributed environment.

Qualifications

  • Strong software engineering fundamentals; design clean abstractions and maintainable code.
  • Production Go experience preferred; systems programming in another language acceptable.
  • Understanding of Kubernetes internals - autoscaling and networking.
  • Experience delivering complex projects end-to-end.
  • Experience with AI coding agents and agentic development workflows.
  • Observability tooling in production experience.
  • CI/CD and DevOps practices experience.
  • Strong English, both verbal and written.
  • Startup mindset: adaptable and proactive.

Responsibilities

  • Design and build distributed systems that operate Kubernetes infrastructure autonomously at scale.
  • Write production Go services that interact with AWS, GCP, and Azure APIs for real-time cloud resource management.
  • Own features end-to-end from design through implementation, testing, and production rollout (1–4 weeks).
  • Debug complex production issues across cloud providers, Kubernetes clusters, and distributed services.
  • Collaborate with product and other engineering teams to solve non-textbook problems.
  • Work with time-series data, cloud provider APIs, and Kubernetes control plane internals.
  • Contribute to product direction and backlog – shape what gets built next.
  • Participate in on-call rotation for teams that require it.

Skills

Go
Kubernetes
CI/CD
DevOps
English

Tools

Terraform
PostgreSQL
Cloud Platforms

Job description

What you'll be solving

A pod is mid-request. The node it's running on is about to disappear - Spot is pulling the instance, or Autoscaler has flagged the node as underutilized and marked it for removal. The safe move is to wait. The profitable move is to act now. Every team in this pipeline is built to close that gap: make the aggressive call and still guarantee nothing breaks.

That guarantee is the hard part. Deciding which nodes to kill, which pods to move, and which storage volume to resize is a genuinely hard optimization problem - several of them are P=NP-hard in the general case - and it has to run continuously, in production, faster than a human would ever attempt it by hand. Get it right and clusters run at half the cost. Get it wrong once and you've taken down something live.

This posting covers multiple teams working on different pieces of the same automation problem node placement algorithms and provisioning, deep integrations with tools like Karpenter, workload rightsizing (vertical and horizontal), moving running workloads between nodes without interrupting them, storage that scales itself, and GPU capacity optimization across clouds and regions. All of it runs in real time, in production, with no human tuning YAML in the loop.


This is a location-specific opportunity. We are currently accepting applications from candidates residing in the following European countries: Bulgaria, Croatia, Estonia, Greece, Hungary, Latvia, Lithuania, Poland, Romania, Slovakia, Slovenia, and Ukraine.

Requirements:
  • Strong software engineering fundamentals: you design clean abstractions, write maintainable code, and can justify your architecture choices - not just make something work.

  • Production experience with Go is strongly preferred; candidates without Go should demonstrate strong systems programming skills in a comparable language.

  • Understanding of Kubernetes internals - autoscaling and networking.

  • You've personally driven a complex project end-to-end.

  • Experience with AI coding agents and agentic development workflows.

  • You've used observability tooling in production.

  • CI/CD and DevOps practices experience.

  • Strong English skills, both verbal and written.

  • Startup mindset: adaptable, proactive, and comfortable with ambiguity.

Nice-to-haves:
  • Deep hands-on experience with cloud platforms (AWS, GCP, or Azure) - including real understanding of how compute, networking, and storage work under the hood.

  • Deep knowledge of EKS, GKE, or AKS internals.

  • Container runtime experience CRI-O, runc, containerd.

  • Cloud storage depth - block storage, CSI, volume management, filesystem internals.

  • Low-level Linux experience: process internals, packet-level networking (NAT, iptables, conntrack, eBPF, SDN), filesystems, storage.

  • Low-level systems programming experience in C or C++.

  • Kubernetes or cloud-native OSS contributions.

  • Experience with AI coding agents and agentic development workflows.

Responsibilities:
  • Design and build distributed systems that operate Kubernetes infrastructure autonomously at scale.

  • Write production Go services that interact with AWS, GCP, and Azure APIs for real-time cloud resource management.

  • Own features end-to-end: from design through implementation, testing, and production rollout (most projects ship in 1-4 weeks).

  • Debug complex production issues across cloud providers, Kubernetes clusters, and distributed services.

  • Collaborate with product and other engineering teams to solve problems that don't have textbook solutions.

  • Work with time-series data, cloud provider APIs, and Kubernetes control plane internals.

  • Contribute to product direction and backlog - shape what gets built next, not just how it's built.

  • Participate in on-call rotation for teams that require it.

Tools we use daily:
  • Programming Languages: Go

  • Cloud & Orchestration: Kubernetes, AWS, GCP, Azure

  • Infrastructure as Code: Terraform

  • Databases & Storage: PostgreSQL, Cloud Object Storage, ClickHouse

  • Messaging & APIs: GCP Pub/Sub, gRPC for internal communication, REST for public APIs

  • Observability: Prometheus, Grafana, Loki, Tempo

  • CI/CD & GitOps: GitLab/Github CI with ArgoCD and Kargo

What’s in it for you?
  • Competitive salary (€6,500 - €9,000 gross, depending on the level of experience)
  • Enjoy a flexible, remote-first global environment.
  • Collaborate with a global team of cloud experts and innovators, passionate about pushing the boundaries of Kubernetes technology.
  • Equity options.
  • Get quick feedback with a fast-paced workflow. Most feature projects are completed in 1 to 4 weeks.
  • Spend 10% of your work time on personal projects or self-improvement.
  • Learning budget for professional and personal development - including access to international conferences and courses that elevate your skills.
  • Annual hackathon to spark new ideas and strengthen team bonds.
  • Team-building budget and company events to connect with your colleagues.
  • Equipment budget to ensure you have everything you need.
  • Extra days off to help maintain a healthy work-life balance.
Hiring process
  • Screening call with Recruiter
  • Hiring Manager interview
  • Technical interview (system design)
  • Live coding
  • Culture Check interview with an executive

*As part of our standard hiring process, we would like to inform you that a background check may be conducted at the final stage of recruitment through our third-party provider, Checkr.*
*Please note that Cast AI does not provide any form of visa sponsorship/work permit.

#LI-Remote

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Account Executive | France
Account Executive | France

Cast AI • United States

Remote
USD 200,000 - 250,000
Equity options
Remote-friendly
Learning budget
+3
Site Reliability Engineer
Site Reliability Engineer

sportygroup • United States

Remote
USD 150,000 - 210,000
Remote-first company
Competitive salary
Kubernetes Engineer
Kubernetes Engineer

Satori Analytics • Kentucky

On-site
USD 90,000 - 150,000
Hybrid work model
Training budget for certifications
Private insurance
+1
Senior Software Engineer | Kimchi Studio
Senior Software Engineer | Kimchi Studio

Creandum • Northern (KY)

Hybrid
USD 82,000 - 114,000
Flexible remote-first
Equity options
Learning budget
+2
Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

sportygroup • United States

Remote
USD 140,000 - 170,000
Remote-first company
Competitive salary with quarterlyBonu
28 days paid annual leave
+4
Kimchi Sales Engineer | EMEA (Remote)
Kimchi Sales Engineer | EMEA (Remote)

Cast AI • United States

Remote
USD 120,000 - 190,000
Equity options
Learning budget
Annual hackathon
+1
Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

Sporty Group • United States

On-site
USD 120,000 - 180,000
Remote first
Bonuses (quarterly)
28 days leave
+4
Partner Account Manager | UK
Partner Account Manager | UK

Cast AI • United States

Remote
USD 119,000 - 172,000
Competitive salary
Remote-first environment
Equity options
+5
Commercial Counsel
Commercial Counsel

Creandum • Northern (KY)

Hybrid
USD 140,000 - 170,000
401(k) company match
Health, dental, vision
Equity
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Camunda • Atlanta (GA)

Remote
USD 150,000 - 242,000
Remote work
Annual company events
Health & wellbeing
+2