Position Summary:
We’re looking for a DevOps / Site Reliability Engineer to join our team. You will play a key role in
ensuring that our systems are efficient, reliable, and scalable, while helping us improve
developer productivity and application performance. You’ll collaborate closely with developers,
QA, product, and our cloud security engineer to streamline builds and deployments, maintain
application infrastructure and proactively solve issues before they impact our users.
We are a product development team full of fun, committed, and hardworking engineers, designers and product managers distributed across the United States. We are scaling our network and building innovative tools to empower student athletes, college coaches, and event operators. Our tools are built on top of technologies that span mobile and web applications, computer vision, and LLMs.
We emphasize performance, security, and maintainability—and we love solving problems that have real-world impact on student-athletes, coaches, and partners.
Position Responsibilities:
CI/CD & Deployments:
- Configure, manage, and improve Bitbucket pipelines for deploying our
applications to staging and production.
- Improve CI pipeline speed, reliability, and security in collaboration with our Cloud Security Engineer.
- Assist developers and QA teams with deployments.
- Work with Docker and AWS ECR for container builds and deployment workflows.
Monitoring & Incident Response:
- Review and investigate system issues flagged by Sentry, NewRelic, and CloudWatch.
- Monitor application performance, identify bottlenecks, and propose solutions.
- Respond to production and staging issues, including database latency, unresponsive resources, or failed jobs.
Environment & Infrastructure Management:
- Maintain and support non-production environments used by developers and QA.
- Maintain and improve AWS infrastructure and Terraform resources.
- Perform updates and upgrades to AWS services as needed to ensure reliability and ability to scale.
Collaboration & Continuous Improvement:
- Partner with engineers to design systems that are scalable, observable, and resilient.
- Work closely with our cloud security engineer to ensure secure configurations in CI/CD, AWS, and containerized workloads.
- Contribute ideas and improvements to workflows, automation, and monitoring strategies.
- Leverage AI to automate monitoring and diagnosis.
Knowledge, Skills and Abilities:
- 3+ years of experience in DevOps, SRE, or related engineering roles.
- Strong experience configuring CI/CD pipelines (Bitbucket Pipelines, GitHub Actions, or similar).
- Experience configuring, debugging and deploying PHP applications
- Hands-on experience with Docker and AWS ECR for container builds and deployments.
- Strong experience with AWS services (EC2, RDS, ECS, Lambda, etc.) and Terraform for infrastructure as code.
- Familiarity with monitoring and observability tools such as New Relic, Sentry, CloudWatch, or similar.
- Strong troubleshooting skills for debugging performance issues in databases, applications, and distributed systems.
- Experience with modern software development workflows (agile teams, code reviews, branching strategies).
- Strong scripting and automation skills (Bash, Python, or similar).
- Excellent communication skills and a collaborative mindset.
- Interest in leveraging AI agents to automate monitoring and diagnosis workflows
#LI-TR1