AI Ops Engineer

TechDigital Group

Bryan (TX)

On-site

USD 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading tech solutions provider in Texas seeks an experienced IT Operations Specialist. The role involves proactive monitoring, maintaining system health, and managing incidents effectively. Candidates should have a degree in computer science or a related field, with at least 5 years of experience in IT operations or L1 support roles. The ideal candidate has hands-on experience with IT monitoring tools and scripting skills, coupled with strong attention to detail and communication skills. Opportunities for professional growth and a focus on automation and reliability are emphasized.

Qualifications

  • 5+ years in IT operations or L1 support roles.
  • Hands-on experience with IT monitoring tools.
  • Understanding of scripting for basic automation tasks.

Responsibilities

  • Proactive monitoring of alerts and detect anomalies.
  • Perform daily health checks until full automation is implemented.
  • Update incident status every 4 hours for P1/P2 tickets.

Skills

Splunk
PowerShell
Python
Logs Monitoring
Confluence
SharePoint

Education

Bachelor's or master's degree in computer science, Engineering, or a related field

Tools

Nagios
Zabbix
Prometheus
ServiceNow
Jira
Remedy

Job description

Experience Requirements
  • 5+ years in IT operations or L1 support roles.
  • Exposure to AIOps environments or automated monitoring solutions is a plus.
Qualifications
  • Bachelor's or master's degree in computer science, Engineering, or a related field.
Key Skills

Splunk, PowerShell, or Python, Logs Monitoring, Confluence and SharePoint

Skill Requirements
  • Hands‑on experience with IT monitoring tools (e.g., Nagios, Zabbix, Prometheus, Splunk, or similar).
  • Understanding of scripting (PowerShell, Python, or Shell) for basic automation tasks.
  • Understanding of AIOps concepts and automation frameworks.
  • Proficiency in Confluence and SharePoint for status updates and documentation.
  • Ability to interpret logs and detect anomalies proactively.
  • Familiarity with ITIL processes for incident, problem, and change management.
  • Experience using ticketing systems (e.g., ServiceNow, Jira, Remedy).
  • Skilled in creating and updating runbooks and SOPs.
  • Ability to follow documented procedures accurately.
  • Strong attention to detail for maintaining health check reports and incident updates.
  • Analytical thinking for quick problem identification and escalation.
  • Excellent communication and documentation skills.
  • Proactive mindset with a passion for reliability and automation.
  • Strong problem‑solving and debugging skills.
Preferred
  • ITIL Foundation Certification.
  • Experience with anomaly detection, time‑series forecasting, and log analysis.
  • Basic certifications in monitoring tools or cloud platforms (AWS, Azure).
Key Responsibilities
  • Proactive Monitoring of alerts and detect anomalies from logs.
  • Perform daily health checks until full automation and application monitoring are implemented.
  • Follow status checks as per existing runbooks.
  • Create and update runbooks as needed to reflect current processes.
  • Update system health status every 2 hours during the shift in Confluence or SharePoint.
  • Acknowledge incidents promptly and route them to the correct team.
  • Update incident status every 4 hours for P1/P2 tickets.
  • Communicate with users and provide timely updates on their requests.
  • Ensure timely acknowledgment, follow‑up, and closure of incidents within SLA.
  • Complete service tasks on time as per SLA to release queues quickly.
  • Work strictly as per SOPs documented by the team.
  • Familiarity with incident management processes and ITIL principles.
  • Ability to follow documented procedures and create/update runbooks.
  • Strong communication and coordination skills.
  • Understanding of Confluence, SharePoint, and ticketing systems.
  • Implement best practices in ML operations and productionization.
  • Ensure compliance with enterprise data security, governance, and regulatory requirements.
  • Collaborate with data engineers, analysts, DevOps/SRE teams and business teams to ensure reliability and security.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIOps Engineer
AIOps Engineer

Career Listings • Fort Belvoir (VA)

On-site
USD 110,000 - 150,000
401(k)
401(k) matching
Health insurance
+2
AI OPS Engineer
AI OPS Engineer

Nowges • Fort Belvoir (VA)

On-site
USD 160,000 - 175,000
AIOps Engineer
AIOps Engineer

Compunnel, Inc. • Reston (VA)

On-site
USD 120,000 - 150,000
AIOps Engineer
AIOps Engineer

Sarela Technology Solutions • Fort Belvoir (VA)

On-site
USD 140,000 - 210,000
401(k)
401(k) matching
Dental insurance
+6
Senior AI OPS Engineer
Senior AI OPS Engineer

JCS Solutions LLC • Arlington (VA)

On-site
USD 150,000 - 174,000
Health, dental, and vision insurance
401k retirement plan with employer match
Paid time off (PTO)
AIOps Engineer Location: VA-Ft Belvoir-22060 Full / Part Time
AIOps Engineer Location: VA-Ft Belvoir-22060 Full / Part Time

Sarela Technology Solutions LLC • Mission (KS)

On-site
USD 110,000 - 140,000
401(k)
401(k) matching
Dental insurance
+6
IT Operations Analyst
IT Operations Analyst

Compunnel, Inc. • Aliso Viejo (CA)

On-site
USD 55,000 - 75,000
Operations Analyst
Operations Analyst

Encore Technologies • Cincinnati (OH)

On-site
USD 75,000 - 105,000
Manager of DevOps Operational Support
Manager of DevOps Operational Support

Request Technology, LLC • Chicago (IL)

On-site
USD 140,000 - 190,000
Bonus eligible
Operations Analyst
Operations Analyst

Encore Talent • Cincinnati (IA)

Hybrid
USD 60,000 - 80,000