Site Reliability Engineer

FLUIX

Palo Alto (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth

Job summary

A technology company in California seeks a skilled Site Reliability Engineer to ensure reliability and performance of their hybrid-based platform. The ideal candidate will collaborate with engineering and operations teams to enhance system reliability, develop automation tools, and respond to incidents. Candidates should have a strong background in managing cloud infrastructure, proven experience in SaaS environments, and proficiency in programming languages such as Python.

Qualifications

  • Proven experience as a Site Reliability Engineer or similar role in a SaaS environment.
  • Experience with ML and AI technologies, and familiarity with data center operations integrations.
  • Excellent problem-solving skills and strong communication abilities.

Responsibilities

  • Design, implement, and maintain scalable systems while optimizing performance.
  • Develop and maintain automation tools to enhance system reliability.
  • Respond to and resolve incidents effectively and timely.

Skills

Site Reliability Engineering
Cloud infrastructure management
Machine Learning and AI technologies
Proficiency in Python
Kubernetes

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

AWS
GCP
Azure
CI/CD pipelines

Job description

FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge Machine Learning (ML) and Artificial Intelligence (AI) technologies. Our mission is to double America’s compute capacity without building new data centers.

We are seeking a skilled Site Reliability Engineer to join our growing team. The ideal candidate will help ensure the reliability, scalability, and performance of our hybrid-based (Cloud & On-Prem) platform while supporting our AI/ML infrastructure. You will work closely with our engineering, AI, and operations teams to build and maintain robust systems that support our cutting-edge solutions. Your expertise in ML/AI and experience with data center sites will be crucial in driving the success of our platform.

Who you’ll work closely with

Founder & CEO

Chase Overcash

CTO

What you’ll do

Design, implement, and maintain scalable systems while optimizing performance, ensuring high availability and disaster recovery, and assisting with codebase refactoring for modular deployment.

Develop and maintain automation tools to streamline operations, improve efficiency, and automate repetitive tasks to enhance system reliability.

Collaborate with engineering and data science teams to integrate ML and AI models into production environments, while ensuring seamless integration and high performance of cutting-edge models within our technology stack.

Identify areas for improvement and drive initiatives to enhance system reliability and performance, while staying updated on industry trends and advancements in SRE practices, ML, and AI technologies.

Respond to and resolve incidents to minimize impact and ensure timely resolution, while conducting post-incident reviews and implementing improvements to prevent recurrence.

Create and manage multiple cloud instances (dev, staging, test), optimize cloud infrastructure and data center operations, and ensure the security and compliance of both infrastructure and applications.

Your background

Bachelorʼs degree in Computer Science, Engineering, or a related field (or equivalent experience).

Proven experience as a Site Reliability Engineer or similar role in a SaaS environment, with a strong background in managing and optimizing cloud infrastructure (AWS preferred, or GCP, Azure), experience with ML and AI technologies, and familiarity with data center operations integrations.

Proficiency in programming and scripting languages (e.g., Python), experience with containerization and orchestration tools (Kubernetes), a strong understanding of networking, security, and performance optimization, and knowledge of CI/CD pipelines and DevOps practices.

Excellent problem-solving skills with attention to detail, strong communication and collaboration abilities, and the capacity to thrive in a fast-paced, dynamic startup environment.

Culture Fit

We are looking for obsessed individuals who want to give it their all.

We are not afraid to get our hands dirty with physical and software systems.

We are eager to visit and work with clients and understand the importance and gravitas of their mission-critical work.

We are eager to come into the office and on-site, as our work directly affects physical environments.

Due to our mission-critical work, we understand and our eager to help our teammates and co-workers during holidays, weekends, and emergencies.

We are cordial and over-communicate with teammates, co-workers, and management.

Attractive compensation package, including equity options.

Comprehensive health, dental, and vision insurance, along with other standard benefits.

A dynamic and collaborative San Francisco Bay Area work environment.

Opportunities for professional growth and development, with the chance to shape the future of technology in the industry.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Engineer
AI/ML Engineer

FLUIX • San Francisco (CA)

On-site
USD 180,000 - 250,000
Attractive compensation package, including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth and development
+1
Software Engineering Internship
Software Engineering Internship

FLUIX • Santa Clara (CA)

Hybrid
Paid internship
Direct mentorship from senior engineers
Dynamic work environment
Site Reliability Engineer
Site Reliability Engineer

Inclusion Services S.A • Chicago (IL)

On-site
USD 90,000 - 130,000
100% company-covered health insurance
401k plan with 4% match
15 days paid time off
+3
Site Reliability Engineer - NYC
Site Reliability Engineer - NYC

Mistral • New York (NY)

Hybrid
USD 140,000 - 190,000
Competitive salary and equity
Healthcare: Medical/Dental/Vision for你
401K with match
+6
Site Reliability Engineer Austin, TX
Site Reliability Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 120,000 - 160,000
Platform DevOps Engineer Austin, TX
Platform DevOps Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 130,000
Competitive salary
Flexible work environment
High-performance culture
Site Reliability Engineer
Site Reliability Engineer

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000