Software Engineer - Datacenter

Worky

Southaven (MS)

On-site

USD 120,000 - 160,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SpaceXAI is seeking a Datacenter Engineer to build and operate software stacks that enable scalable, auditable site operations across its data centers. You will partner with operations, research, and infrastructure teams to deliver high-leverage tools that translate raw data into actionable insights.

The role emphasizes ownership of system correctness, production-grade APIs, data pipelines, and end-to-end reliability.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related fields.
  • 3+ years building and operating production software (backend and/or full-stack).
  • Strong fundamentals in data structures, algorithms, OS, and networking.
  • Proficiency in programming languages such as Rust, Python, JavaScript, Java, C++.
  • Experience designing RESTful APIs for mission-critical applications.
  • Experience with databases like Postgres, MongoDB, MySQL, DynamoDB.
  • Experience with cloud platforms (GCP, Azure, AWS, OCI or similar).
  • Experience using observability tools and dashboards (New Relic, Splunk, Grafana).
  • Experience writing unit tests and integration tests.

Responsibilities

  • Build and operate the software stacks for scalable, auditable site operations.
  • Design and run multi-service production systems (UI, APIs, data pipelines, auth, on-call).
  • Own correctness of operational state, queues, and audit trails.
  • Develop integrations with ticketing, inventory, telemetry stores, and vendor portals.
  • Ensure reliability: uptime, data integrity, reconciliation, access control, safe Deploys.
  • Collaborate with SiteOps and NOC to measure workflow adoption and impact.

Skills

Programming languages (Rust)
Python
JavaScript
Java
C++
Data structures & algorithms
RESTful APIs
Unit tests
Observability

Education

Bachelor's degree in Computer Science/Engineering
MS in Computer Science or related field

Tools

Postgres
MongoDB
MySQL
DynamoDB
GCP
AWS
Azure
New Relic
Splunk
Grafana
CI/CD (Azure DevOps)
ArgoCD
Jenkins
GitHub Actions

Job description

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

The Datacenter Engineering team builds the internal systems and platforms that keep our datacenters running at the scale and reliability required for frontier AI training and inference. We partner closely with datacenter operations, research, and infrastructure teams to deliver high-leverage tools that turn raw operational data into clear insight and action.

RESPONSIBILITIES:
  • Build and operate the software stacks that make site operations scalable, auditable, and fast — including repair trackers, vendor turnback workflows, operational dashboards, and their integrations. Your users are the technicians, managers, NOC operators, and leadership who run the fleet.
  • Design, build, and operate multi-service production systems (UI, APIs, data pipelines, auth, and on-call) for systems such as SRT-style repair/maintenance trackers and vendor turnover/turnback state machines.
  • Own correctness of operational state: node state accuracy, queue ownership, and audit trails — the data that decides what work happens on the floor.
  • Build and maintain integrations with ticketing, inventory/rack systems, telemetry stores, and vendor portals.
  • Keep the tools themselves reliable: uptime, data integrity, reconciliation, access control, and safe deploys.
  • Embed with SiteOps and NOC users; measure workflow adoption, not just feature delivery.
BASIC QUALIFICATIONS:
  • Bachelor’s degree in Computer Science, Engineering, or related fields.
  • 3+ years building and operating production software (backend and/or full-stack).
  • Strong fundamental knowledge of computer science - data structures, algorithms, operating systems and networking.
  • Strong proficiency in at least one programming language e.g. Rust, Python, JavaScript, Java, C++, etc.
  • Experience designing and developing RESTful APIs for mission-critical applications.
  • Experience working with at least one database like Postgres, MongoDB, MySQL, DynamoDB, etc.
  • Experience collaborating with cross-functional teams.
  • Experience working in cloud platforms like GCP, Microsoft Azure, AWS, OCI or similar.
  • Experience using observability tools and dashboards like New Relic, Splunk, Grafana, etc.
  • Experience writing unit tests and integration tests.
PREFERRED SKILLS AND EXPERIENCE:
  • Full-stack or backend + data experience, especially with workflow and state-machine systems.
  • Experience automating deployments using CI/CD tools like Azure Devops, ArgoCD, Jenkins, GitHub Actions, etc.
  • Experience building high-correctness operational UIs where a wrong value can dispatch a human to the wrong rack.
  • On-call discipline and a track record of treating internal platforms with production rigor.
  • Experience designing and shipping multi-service systems (APIs, data stores, and at least one of: UI, pipelines, or auth).
  • Experience delivering high-quality internal tools in rapidly changing environments (as a tech lead, former founder, etc.).
  • Proven ownership of correctness-sensitive systems — state machines, workflows, or operational data where inaccurate state has real-world impact.
  • Experience integrating with external systems via APIs (e.g. ticketing, inventory, telemetry, or vendor portals).
  • MS in Computer Science or related field.
  • Experience collaborating closely with operations, NOC, and infrastructure teams.
  • Experience working with performance load testing tools like BlazeMeter, k6, etc.
ADDITIONAL REQUIREMENTS:
  • The role is fully onsite in Memphis, TN or Southhaven, MS. Candidates are expected to be located near the area or open to relocation.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Datacenter
Software Engineer - Datacenter

AI Chopping Block • Memphis (TN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Software Engineer - Data Center
Software Engineer - Data Center

Pantera Capital • Southaven (MS)

On-site
USD 90,000 - 130,000
Software Engineer - Data Center
Software Engineer - Data Center

Socket.dev • Memphis (TN)

On-site
USD 100,000 - 150,000
Software Engineer - Datacenter
Software Engineer - Datacenter

SpaceXAI • Memphis (TN)

On-site
USD 120,000 - 160,000
Software Engineer - Data Center
Software Engineer - Data Center

SpaceXAI • Memphis (TN)

On-site
USD 95,000 - 130,000
Supervisor, Data Center Operations
Supervisor, Data Center Operations

Pantera Capital • Southaven (MS)

On-site
USD 90,000 - 130,000
Manager, Data Center Operations
Manager, Data Center Operations

Xai • Memphis (TN)

On-site
USD 110,000 - 150,000
Manager, Site Operations
Manager, Site Operations

Pantera Capital • Memphis (TN)

On-site
USD 90,000 - 130,000
Site Reliability Engineer - Datacenter
Site Reliability Engineer - Datacenter

AI Need That • Memphis (TN), Northern (KY)

Hybrid
USD 90,000 - 130,000
Operations Engineer (Facility Operations)
Operations Engineer (Facility Operations)

Socket.dev • Memphis (TN)

On-site
USD 85,000 - 110,000