An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Get past ATS filters
Job summary
A technology solutions company is seeking a Software Engineer (Site Reliability Engineer) to work remotely. This role involves creating and analyzing production incident scenarios, evaluating AI model performance, and improving incident response processes. The ideal candidate has strong experience in SRE, DevOps, and observability tools, along with excellent debugging skills. This hourly contract offers compensation between $100 and $160 per hour.
Qualifications
Strong experience in SRE, DevOps, or production engineering.
Experience with on-call operations and managing production incidents.
Experience with root cause analysis and post-mortem processes.
Experience with observability tools such as Prometheus, Grafana, Datadog, or PagerDuty.
Knowledge of Linux systems, networking, and container orchestration.
Experience with infrastructure-as-code and CI/CD pipelines.
Strong debugging skills across application and system levels.
Ability to evaluate complex system reliability scenarios.
Responsibilities
Create and review realistic production incident scenarios for AI evaluation.
Analyze system failures including root cause analysis, monitoring, and remediation.
Evaluate AI model performance on infrastructure and reliability problem-solving.
Develop scenarios involving observability, alerting, and capacity planning.
Assess best practices in incident response and post-mortem processes.
Contribute to improving AI reasoning in production system environments.
Skills
SRE
DevOps
Production engineering
Root cause analysis
Incident management
Observability tools
Linux systems
Networking
Container orchestration
Infrastructure-as-code
CI/CD pipelines
Debugging skills
Tools
Prometheus
Grafana
Datadog
PagerDuty
Job description
A technology solutions company is seeking a Software Engineer (Site Reliability Engineer) to work remotely. This role involves creating and analyzing production incident scenarios, evaluating AI model performance, and improving incident response processes. The ideal candidate has strong experience in SRE, DevOps, and observability tools, along with excellent debugging skills. This hourly contract offers compensation between $100 and $160 per hour.