Staff Software Engineer, Cloud Monitoring Service (Distributed Systems)

Crusoe

San Francisco (CA)

On-site

USD 215,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity packages
RSUs
PTO & holidays
Health insurance
HSA contributions
Parental leave
Life insurance
Professional development
Mental health support
Commuter benefits
Cell phone stipend
401(k) match
Volunteer time off
Travel insurance
Meals allowance
Location perks

Job summary

Crusoe is building the observability backbone for Crusoe Cloud and is seeking a Staff Software Engineer to lead distributed systems design for our Cloud Monitoring Service. You will own high-throughput telemetry pipelines from edge collection to storage and query, with a focus on correctness, availability, and cost-efficiency at scale.

We value cross-functional collaboration, on-call reliability, and mentorship of senior engineers.

Qualifications

  • 8+ years of software development experience with production distributed systems.
  • 2+ years leading large-scale telemetry or observability platforms.
  • Strong Go background and modern language fundamentals.
  • Experience with Kubernetes and CI/CD in production environments.
  • Ability to scope ambiguous problems and tradeoffs for customers and business outcomes.
  • Track record of cross-functional technical influence across teams.
  • Mentorship through design reviews and code guidance.
  • 8+ years of production distributed systems experience.

Responsibilities

  • Own the architecture and evolution of large-scale telemetry pipelines (ingestion, processing, storage, query).
  • Design scalable, multi-tenant data systems with high availability and cost-efficiency.
  • Improve pipeline reliability and reduce on-call burden through robust design.
  • Set technical direction and lead design reviews for distributed systems.
  • Collaborate with product, compute, networking, and platform teams to align observability decisions.
  • Mentor engineers through design work, code reviews, and incident response.

Skills

Distributed systems
Observability data
Go programming
Kubernetes
CI/CD
Design judgment
On-call / incident response
Cross-functional influence
Mentorship
Production experience

Tools

Prometheus
VictoriaMetrics
Loki
OpenTelemetry
Kafka
Vector

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role:

We are seeking a Staff Software Engineer to lead distributed systems design and development within Crusoe Cloud's Cloud Monitoring Service. This team builds the observability backbone of Crusoe Cloud: the telemetry systems that collect, process, store, and serve the metrics and logs our customers depend on to run AI workloads at scale. You will own the design and evolution of high-throughput, multi-tenant data systems, from collection at the edge through ingestion, storage, and query. This is a full-time position.

What You'll Be Working On:
  • Distributed Systems Ownership: Own the architecture and evolution of large-scale telemetry pipelines, including high-volume ingestion, stream processing, time-series and log storage, and low-latency query paths. Design systems that stay correct, available, and cost-efficient as data volume grows 10x.

  • Scalable, Multi-Tenant Design: Design services that are highly scalable, durable, and fair across tenants. Solve the hard problems in this space: hot shards, high-cardinality data, noisy neighbors, backpressure, retention and compaction at scale, and graceful degradation under load.

  • Reliability and Operational Excellence: Build for operability from day one. Improve pipeline reliability and data freshness, reduce on-call burden through better system design rather than more process, and participate in a customer-facing on-call rotation, leading by example.

  • Technical Leadership: Set the technical direction for the team's distributed systems work. Drive design reviews, identify one-way door decisions early, and raise the bar on how the team scopes, builds, and operates systems.

  • Cross-Team Collaboration: Work with product, compute, networking, and platform teams to make sure observability decisions are made with full context. Represent the team's technical position in cross-org conversations.

  • Mentorship: Coach senior and mid-level engineers through design work, code review, and incident response. Build patterns and frameworks that make the team better without requiring your direct involvement.

What You'll Bring to the Team:
  • Distributed Systems Depth: Deep, hands-on experience designing and operating distributed systems at scale. You have solved real problems in sharding, replication, consistency, load balancing, and concurrency, not just studied them.

  • Observability Data Experience: Experience building or operating large-scale data infrastructure such as time-series databases, log aggregation, streaming pipelines, or distributed tracing backends. Familiarity with technologies like Prometheus, VictoriaMetrics, Loki, OpenTelemetry, Kafka, Vector, or similar.

  • Technical Proficiency: Strong programming fundamentals in Go or another modern compiled language (Go strongly preferred). Comfort with Kubernetes, microservices, and CI/CD as the operating environment for everything you build.

  • Design Judgment at Staff Level: You proactively scope ambiguous problems, surface non-functional requirements without being prompted, and reason about tradeoffs in terms of customer and business outcomes, not just technical elegance.

  • Operational Mindset: On-call experience on a customer-facing service. You treat incidents as design feedback and systematically eliminate the conditions that cause them.

  • Cross-Functional Influence: A track record of driving technical outcomes that span team boundaries, and the communication skills to bring product, support, and adjacent engineering teams along.

  • Mentorship: You make the engineers around you better through design guidance, thoughtful review, and honest, constructive feedback.

  • Professional Experience: 8+ years of software development experience, with sustained ownership of production distributed systems.

Benefits:
  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $215,000 -$260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Cloud Monitoring Service (Distributed Systems)
Staff Software Engineer, Cloud Monitoring Service (Distributed Systems)

Crusoe Energy Systems LLC • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation
Equity packages
Paid time off
+3
Senior Software Engineer, Cloud Monitoring Service
Senior Software Engineer, Cloud Monitoring Service

Crusoe Energy Systems LLC • San Francisco (CA)

On-site
USD 175,000 - 210,000
Competitive compensation
Equity
Paid time off
+2
Senior Software Engineer, Cloud Monitoring Service
Senior Software Engineer, Cloud Monitoring Service

Crusoe • San Francisco (CA)

On-site
USD 175,000 - 210,000
Competitive compensation and equity
Restricted Stock Units
Paid time off and holidays
+13
Engineering Manager, Telemetry Agent and Edge
Engineering Manager, Telemetry Agent and Edge

Crusoe Energy Systems LLC • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off and holidays
+13
Engineering Manager, Managed Platform Services
Engineering Manager, Managed Platform Services

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off and holidays
+12
Engineering Manager, Telemetry Agent and Edge
Engineering Manager, Telemetry Agent and Edge

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+12
Senior Manager, Engineering Operations
Senior Manager, Engineering Operations

Crusoe • San Francisco (CA)

On-site
USD 230,000 - 280,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+13
Senior Manager, Engineering Operations
Senior Manager, Engineering Operations

Crusoe • San Francisco (CA)

On-site
USD 230,000 - 280,000
Equity
Paid time off
Health insurance
+4
Software Engineer, Control Plane
Software Engineer, Control Plane

crusoe • San Francisco (CA)

On-site
USD 136,000 - 161,000
Equity
Paid time off
Health insurance
+12
Senior Software Engineer, Datacenter Operations Platform Engineering
Senior Software Engineer, Datacenter Operations Platform Engineering

Linuxconfig • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 205,000
Health insurance
Restricted Stock Units
401(k) with 100% match up to 4%
+2