Role: Site Reliability Engineer
Location: Southlake / Austin, TX - Onsite 4 days weekly
Duration: 12 Months
Job Summary
We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure, and production operations. The ideal candidate has a strong automation-first mindset and enjoys solving complex operational challenges. This role focuses on improving reliability, observability, operational efficiency, and automation across on-premises and cloud platforms.
Key Responsibilities
Automation & Platform Engineering
- Develop Python-based automation solutions to reduce manual operational effort.
- Automate infrastructure management across Linux, Windows, Kubernetes, GCP, and cloud-native environments.
- Integrate tools and platforms through APIs and client libraries.
- Assist in implementing infrastructure automation using Ansible, Terraform, or similar technologies.
- Support CI/CD automation and deployment reliability initiatives.
Reliability & Operations
- Monitor and maintain production systems to meet reliability and availability objectives.
- Participate in incident response, troubleshooting, and root cause analysis activities.
- Develop automation and operational improvements to prevent recurring issues.
- Support disaster recovery, failover testing, and operational readiness activities.
- Perform performance analysis and system health reviews.
Observability & Monitoring
- Build and maintain dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, GCP Operations Suite, or similar tools.
- Improve visibility into application and infrastructure health through metrics, logs, and traces.
- Investigate alerts and identify opportunities to reduce noise and improve detection.
AIOps & Intelligence
- Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics.
- Assist in developing automation solutions that leverage AI to improve operational efficiency.
- Participate in evaluating emerging AIOps capabilities and observability technologies.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
- 3 to 5 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Platform Engineering.
- Strong programming skills in Python for automation and tooling development.
- Experience supporting Kubernetes and cloud platforms (GCP, AWS, or Azure).
- Familiarity with infrastructure automation and configuration management tools.
- Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
- Understanding of Linux systems, networking, and distributed applications.
- Strong analytical, troubleshooting, and problem-solving skills.
- Ability to work effectively in fast-paced, mission-critical environments.
Preferred Qualifications
- Experience with Terraform, Ansible, or Infrastructure as Code solutions.
- Exposure to OpenTelemetry and modern observability practices.
- Experience with CI/CD pipelines and deployment automation.
- Knowledge of AI/ML, AIOps, or intelligent operational tooling.
- Experience supporting highly available production systems in regulated or enterprise environments.
Must Have:
- 3-5 years of hands‑on Site Reliability Engineering or Production Engineering experience
- Must have supported production systems at scale.
- Operations ownership, incident response, and reliability engineering experience required.
- Pure DevOps, build/release, or CI/CD-only backgrounds are not a fit.
- Strong Python and operation automation Development Skills
- Demonstrated experience building automation tools, scripts, frameworks, or operational solutions.
- Candidate should be able to provide examples of automation they personally developed.
- Python must be a primary skill, not just basic scripting.
- Production Operations & Incident Management Experience
- Experience troubleshooting critical production incidents.
- Root Cause Analysis (RCA) participation and problem remediation.
- Experience reducing operational toil through automation.