Development Engineer, Facility Software Automation

Fluidstack

Austin (TX)

On-site

USD 250,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options

Job summary

Fluidstack is seeking an experienced infrastructure engineer to own the Facility Software Automation platform’s core. You will design cluster architecture, implement deployment configurations, and ensure reliability across the stack.

You will also lead containerization pipelines and the observability layer to detect issues before they impact customers. You will build and enforce infrastructure-as-code standards, enable fast yet safe deployments, and collaborate with engineers to take telemetry

Qualifications

  • Owned container-based infrastructure in production.
  • Telemetry pipeline health and data integrity are non-negotiable.
  • Hands-on experience with containerization technologies (Docker, Kubernetes, or equivalent).
  • Built CI/CD pipelines engineers trust for safe production deployments.
  • Infrastructure-as-code is version-controlled, reviewed, and auditable.
  • Built observability stacks (Prometheus, Grafana, OpenTelemetry).
  • Experience in reliability‑critical environments.
  • Bonus: GitOps tooling such as ArgoCD or Flux; operations at scale across sites.

Responsibilities

  • Own the infrastructure underpinning the Facility Software Automation platform: cluster architecture, deployment configuration, and the operational reliability of every service in the stack.
  • Own and manage containerization workflows and CI/CD pipelines for telemetry services, making deployments from commit to live in production fast, safe, and repeatable without manual intervention.
  • Own the observability layer, Prometheus, alerting, and dashboards, that gives the platform team and Fluidstack operations the signal to catch problems in telemetry pipelines before they hit downstream teams.
  • Drive infrastructure-as-code standards across the platform so every environment change is version‑controlled, reviewed, and auditable at the pace of a team shipping continuously.
  • Work directly with engineers to take new telemetry services from first commit to production‑ready deployment, setting the bar for what production‑ready means on this platform.

Skills

Containerization
CI/CD pipelines
Observability
Infrastructure as Code
Prometheus
Grafana
OpenTelemetry

Tools

Docker
Kubernetes
ArgoCD

Job description

Facility Software Automation Team

Examples of key problems the team is working on

  • We do not just watch infrastructure run. We use the signals we collect to build automations that act on it -- compressing the time between a facility coming online and compute being in customers' hands. Detect, decide, act. If we fail to do any one of those three, we lose the trust of the customers who depend on us to keep the frontier moving.
  • Bad telemetry is a trust failure at gigawatt scale. Every facility Fluidstack operates runs on the data we collect. A missed signal is a missed alert. A missed alert is a customer running frontier AI workloads on infrastructure that cannot see itself. At the scale we are building, that is not a monitoring gap – it is an outage.
  • The platform has to scale as fast as the build does. Fluidstack delivers gigawatts of compute in six months -- a fraction of the industry's 18 to 24 month timeline -- and every new site adds more devices, more signals, and more automation decisions that depend on clean data. The telemetry platform either grows ahead of the business or it becomes the constraint that slows it down.
Role Scope
  • Own the infrastructure underpinning the Facility Software Automation platform: cluster architecture, deployment configuration, and the operational reliability of every service in the stack.
  • Own and manage containerization workflows and CI/CD pipelines for telemetry services, making deployments from commit to live in production fast, safe, and repeatable without manual intervention.
  • Own the observability layer, Prometheus, alerting, and dashboards, that gives the platform team and Fluidstack operations the signal to catch problems in telemetry pipelines before they hit downstream teams.
  • Drive infrastructure-as-code standards across the platform so every environment change is version‑controlled, reviewed, and auditable at the pace of a team shipping continuously.
  • Work directly with engineers to take new telemetry services from first commit to production‑ready deployment, setting the bar for what production‑ready means on this platform.
What We’re Looking For
  • You have owned container‑based infrastructure in production: designed the architecture, debugged the failure modes that only appear under load, and carried the pager for it.
  • You treat telemetry pipeline health and data integrity as non‑negotiable, since the automation that delivers compute to customers depends entirely on the reliability of the data flowing underneath it, and you build and operate accordingly.
  • You have hands‑on experience with containerization technologies (Docker, Kubernetes, or equivalent) and have used them to build deployment workflows that engineering teams depend on without thinking about it.
  • You have built CI/CD pipelines that engineers actually trust, where a merged PR reaches production safely without anyone watching over it, and you understand modern development workflows well enough to build them from scratch.
  • You drive infrastructure‑as‑code as a functioning discipline: every environment change is version‑controlled, reviewed, and auditable, and you hold the platform to that standard.
  • You have built observability stacks from scratch (Prometheus, Grafana, OpenTelemetry) and you know the difference between a dashboard that looks good and one that actually catches problems before they become incidents.
  • You have worked in environments where reliability matters because real operational decisions depend on the data flowing correctly, not just internal tooling.
  • Bonus: GitOps deployment tooling (ArgoCD, Flux, or equivalent). NATS and ClickHouse operational experience, tuning, capacity planning, failure recovery. Compute or critical infrastructure telemetry environments. Experience deploying and operating services at scale across distributed sites.

Compensation: $250,000 – $300,000 per year, depending on experience, skills, qualifications, and location. Offers equity in the form of stock options.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Facility Software Automation
Software Engineer, Facility Software Automation

Fluidstack • Austin (TX)

On-site
USD 250,000 - 300,000
Equity in the form of stock options
Technical Lead, Facilities Telemetry Platform
Technical Lead, Facilities Telemetry Platform

Fluidstack • New York (NY)

On-site
USD 150,000 - 210,000
Software Engineer, Cloud Infrastructure
Software Engineer, Cloud Infrastructure

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Equity
Health insurance
Retirement plan
+1
Facilities Telemetry Platform, Technical Lead
Facilities Telemetry Platform, Technical Lead

Fluidstack • Austin (TX)

On-site
USD 150,000 - 190,000
Distributed Systems Engineer
Distributed Systems Engineer

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Equity
Retirement plan
Health, dental, and vision insurance
+1
Distributed Systems Engineer
Distributed Systems Engineer

Fluidstack • Seattle (WA)

On-site
USD 175,000 - 300,000
Stock options
Health/dental/vision insurance
Generous PTO
Software Engineer, Ontology
Software Engineer, Ontology

Fluidstack • New York (NY)

On-site
USD 200,000 - 250,000
Health, dental, and vision insurance
Generous PTO
Retirement plan
Facilities Production Technical Lead
Facilities Production Technical Lead

Fluidstack • Seattle (WA)

On-site
USD 208,000 - 269,000
Pay equity and transparency
Software Engineer, Infrastructure Platform
Software Engineer, Infrastructure Platform

Fluidstack • Austin (TX)

On-site
USD 200,000 - 250,000
Health, dental, and vision insurance
Generous PTO policy
Equity and competitive compensation
Software Engineer, Infrastructure Platform
Software Engineer, Infrastructure Platform

Fluidstack • Seattle (WA)

On-site
USD 200,000 - 250,000
Health, dental, vision insurance
Generous PTO