SITE RELIABILITY ENGINEER

Onyxodds

New York (NY)

On-site

USD 145,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity ownership
Fully covered health benefits

Job summary

Onyx is hiring a Site Reliability Engineer for our NoHo, Manhattan office to build and scale reliable production systems on AWS. You’ll own infrastructure, deployment pipelines, and observability, collaborating with engineers across product and security to keep latency low and uptime high.

You’ll manage Terraform-based infrastructure, strengthen monitoring with Datadog, and lead incident response, post-incident reviews, and runbooks while balancing long-term improvements with immediate

Qualifications

  • Experience operating production systems in AWS.
  • Proficient with Terraform and infrastructure as code.
  • Hands-on observability, incident response, and production troubleshooting.
  • Understanding of distributed systems, networking, databases, and cloud architecture.
  • Built automation using Python, TypeScript, Go, Bash, or similar languages.
  • Familiar with deployment pipelines, release strategies, and rollback procedures.
  • Ability to balance long-term infra improvements with immediate production needs.
  • Calm under pressure, ownership from alert to fix.

Responsibilities

  • Build, operate, and improve scalable production infrastructure in AWS.
  • Manage cloud infrastructure through Terraform and IaC.
  • Strengthen monitoring, logging, tracing, alerting, and observability using Datadog.
  • Improve availability, performance, resilience, and security of services.
  • Build tooling and automation to make infrastructure safer and easier to operate.
  • Improve deployment systems, release processes, environment management, and rollback capabilities.
  • Partner with engineers to design reliable, observable services.
  • Monitor production systems and lead incident investigation and resolution.
  • Develop runbooks, escalation processes, and post-incident reviews.
  • Improve database reliability, backups, DR, and capacity planning.
  • Participate in on-call rotation and support during high-volume events.

Skills

AWS production systems
Terraform
Observability
Incident response
Distributed systems
Networking
Databases
Deployment pipelines
Python/Go/TypeScript
On-call ownership

Tools

Datadog
Kubernetes
Docker
ECS/EKS

Job description

Onyx is building the market for people who enjoy being right. Our consumer platform lets people take positions on what happens next across sports, prediction markets, crypto, financial markets, politics, culture, and real-world events.

More than one million users have already found Onyx, driving over $2 billion in monthly notional volume across our social sports platform. Now, we're bringing that same energy into regulated financial markets through Onyx Predictions, a CFTC-registered introducing broker and NFA member.

In June 2026, we raised a $20 million Series A led by Payward, the parent company of Kraken, valuing Onyx at $220 million less than one year after emerging from beta. Our team comes from Jane Street, DraftKings, Robinhood, xAI, SIG, HRT, and Harvard.

We're building Onyx in New York with a small team, serious momentum, and no interest in moving slowly. The product is live. The market is growing. The category is still ours to define.

The Role.

We're hiring a Site Reliability Engineer to build the infrastructure, systems, and operational practices that keep Onyx fast and reliable as we scale.

Markets move in real time. Traffic spikes without warning. Customers expect orders, balances, positions, and settlement to work every time. You'll help ensure our platform is ready for all of it.

Our infrastructure runs primarily on AWS and is managed using Terraform, with Datadog supporting observability and incident response. You'll work directly with application engineers, product, trading, security, and operations to improve reliability, automate infrastructure, strengthen production systems, and respond when something goes wrong.

This is a high-ownership role. You won't just monitor infrastructure or maintain someone else’s playbook. You'll help define how reliability engineering works at Onyx.

What You'll Own:
  • Build, operate, and improve scalable production infrastructure in AWS.
  • Manage cloud infrastructure through Terraform and infrastructure as code.
  • Strengthen monitoring, logging, tracing, alerting, and observability using Datadog.
  • Improve the availability, performance, resilience, and security of critical services.
  • Build tooling and automation that make infrastructure safer and easier to operate.
  • Improve deployment systems, release processes, environment management, and rollback capabilities.
  • Partner with engineers to design reliable, observable, and scalable services.
  • Monitor production systems and lead incident investigation and resolution.
  • Build runbooks, escalation processes, and post-incident review practices.
  • Improve database reliability, backups, disaster recovery, and capacity planning.
  • Participate in an on-call rotation and provide support during high-volume market events.
You'll Do Well Here If You:
  • Have experience operating production systems in AWS.
  • Are proficient with Terraform and infrastructure as code.
  • Have hands-on experience with observability, incident response, and production troubleshooting.
  • Understand distributed systems, networking, databases, and cloud architecture.
  • Have built automation using Python, TypeScript, Go, Bash, or a similar language.
  • Understand deployment pipelines, release strategies, and rollback procedures.
  • Can balance long-term infrastructure improvements with immediate production needs.
  • Remain calm during incidents and take ownership from the first alert through the permanent fix.
Even Better If You Have:
  • Operated infrastructure for trading, crypto, payments, gaming, or consumer fintech.
  • Supported real-time, transactional, or highly concurrent systems.
  • Worked with PostgreSQL in a high-availability production environment.
  • Experience with Docker, ECS, EKS, Kubernetes, or similar container systems.
  • Built CI/CD pipelines and automated deployment processes.
  • Worked with event-driven systems, WebSockets, streaming platforms, or live data feeds.
  • Experience with disaster recovery, IAM, secrets management, or cloud security.
  • Established SRE practices or scaled infrastructure at an early-stage company.
Location and Schedule:
  • This role is based in our NoHo, Manhattan office, with core in-office days Tuesday through Thursday.
  • This position includes participation in an on-call rotation and may require support during major sporting events, market-moving news, or production incidents.
  • Candidates must be authorized to work in the United States. Onyx is unable to provide employment sponsorship for this position.
Compensation and Benefits:
  • The anticipated base salary range for this position is $145,000 - $190,000 depending on experience, impact, and scope, plus target bonus and equity.
Benefits include:
  • Fully covered medical, dental, and vision insurance.
  • Meaningful equity ownership.
Why Onyx?

Onyx is for people with conviction.

We hire talented people, give them meaningful ownership, and expect them to use it. Our teams stay small, the work stays visible, and the distance between an idea and production stays short.

Onyx moves quickly because our markets move quickly. Titles matter less than judgment, initiative, and the quality of what you ship. Good ideas can come from anywhere, but the person willing to make the call is expected to follow it through.

We value people who are ambitious without ego, intellectually honest, comfortable with uncertainty, and energized by hard problems. We debate openly, decide quickly, and take responsibility for the result.

Come build what's next with us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SOFTWARE ENGINEER, FULL-STACK
SOFTWARE ENGINEER, FULL-STACK

Onyxodds • New York (NY)

On-site
USD 150,000 - 200,000
Medical, dental, and vision insurance
Meaningful equity ownership
CUSTOMER EXPERIENCE
CUSTOMER EXPERIENCE

Onyxodds • New York (NY)

On-site
USD 100,000 - 150,000
Medical insurance
Dental insurance
Vision insurance
+2
TRADE OPERATIONS
TRADE OPERATIONS

Onyxodds • New York (NY)

On-site
USD 80,000 - 120,000
Fully covered medical, dental, and vis
Meaningful equity ownership
SOFTWARE ENGINEER, FRONTEND
SOFTWARE ENGINEER, FRONTEND

Onyxodds • New York (NY)

On-site
USD 150,000 - 200,000
Equity ownership
Bonus potential
Health, dental, vision insurance
Customer Support Manager
Customer Support Manager

Onyx Odds • New York (NY)

On-site
USD 120,000 - 160,000
Medical, dental, and vision insurance
Equity ownership
CUSTOMER SUPPORT MANAGER
CUSTOMER SUPPORT MANAGER

Onyxodds • New York (NY)

On-site
USD 120,000 - 160,000
Health insurance
Equity ownership
Paid time off
PRODUCT MANAGER, PREDICTIONS
PRODUCT MANAGER, PREDICTIONS

Onyxodds • New York (NY)

On-site
USD 150,000 - 200,000
Medical, dental, and vision insurance
Equity ownership
Flexible workdays
CRM ASSOCIATE
CRM ASSOCIATE

Onyxodds • New York (NY)

On-site
USD 65,000 - 100,000
Fully covered medical, dental, and vis
Meaningful equity ownership
Product Mangager, Predictions
Product Mangager, Predictions

Onyxodds • New York (NY)

On-site
USD 150,000 - 200,000
Fully covered medical, dental, and vis
Meaningful equity ownership
GROWTH ANALYST
GROWTH ANALYST

Onyxodds • New York (NY)

On-site
USD 65,000 - 100,000
Medical, dental, vision insurance
Meaningful equity ownership