Job Details
Technical Support Specialist
Job ID#: 64554
Job Category: Technical Support Specialist
Position Type: Full Time
Remaining Positions: 3
We are seeking a skilled and proactive Apollo L1 Support Engineer to join our AI Platform Operations team. In this role, you will serve as the first line of technical support for our AI/LLM platform, providing expert-level troubleshooting and incident resolution for complex technical issues involving APIs, networking, and Python-based applications.
This is a highly technical support role requiring strong developer skills and a deep understanding of IT infrastructure. You will work closely with development teams, platform engineers, and end-users to ensure the reliability and performance of our AI platform.
Key Responsibilities
Technical Support & Incident Management
- In cident Handling: Receive, triage, and resolve Level 1 technical incidents related to the Apollo AI platform, ensuring timely resolution and minimal service disruption.
- API Troubleshooting: Diagnose and resolve API connectivity, authentication, and performance issues using tools like Postman, curl, and logging platforms.
- Networking Support: Troubleshoot network-related issues including connectivity, latency, DNS, and firewall configurations affecting platform access.
- Python Application Support: Debug and resolve issues with Python-based automation scripts, data pipelines, and integration workflows.
System Monitoring & Operations
- Proactive Monitoring: Monitor platform health and performance using observability tools, identifying potential issues before they impact users.
- Alert Response: Respond to system alerts, perform initial diagnostics, and elevate complex issues to Level 2/3 engineers as needed.
- Runbook Execution: Follow documented runbooks and standard operating procedures for incident resolution and system maintenance tasks.
Documentation & Knowledge Management
- Knowledge Base: Create and maintain detailed documentation, knowledge articles, and troubleshooting guides to support end-users and internal teams.
- Incident Reports: Document incident root causes, resolution steps, and preventive measures to build a comprehensive knowledge repository.
- Continuous Improvement: Contribute to the improvement of support processes, runbooks, and automation scripts.
Collaboration & Communication
- C ross-functional Collaboration: Work closely with development, platform engineering, and product teams to resolve complex technical issues and communicate platform updates.
- Stakeholder Communication: Provide clear and professional updates to stakeholders on incident status, resolution timelines, and root cause analysis.
- Knowledge Sharing: Actively participate in knowledge transfer sessions, team stand-ups, and post-incident reviews.
Escalation & Incident Management
- Escalation: Escalate complex or unresolved issues to Level 2/3 engineers with clear documentation and diagnostic information.
- Triage & Prioritization: Prioritize incidents based on business impact and service level agreements (SLAs), ensuring critical issues are addressed immediately.
- Incident Documentation: Ensure accurate and detailed logging of all incidents and service requests in the ticketing system.
Job Requirements Details:
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- 3+ years of experience in a technical support, software development, or systems engineering role.
- Proven experience troubleshooting complex technical issues in a production environment.
Technical Skills (Must-Haves)
- Python Strong proficiency in Python – ability to read, debug, and write automation scripts
- APIs Deep understanding of RESTful APIs, authentication (OAuth, API keys), and troubleshooting tools (Postman, curl)
- Networking Solid understanding of networking fundamentals (TCP/IP, DNS, firewalls, load balancers, VPNs)
- Linux/Unix Familiarity with Linux command line for log analysis and system troubleshooting
- CI/CD Understanding of CI/CD pipelines and deployment processes
- Monitoring Experience with monitoring and observability tools
- Ticketing Systems Experience with ServiceNow or similar ITSM platforms
Nice-to-Have Skills
- AI/LLM Concepts Knowledge of Large Language Models, prompting, model inference, and AI platform operations
- Docker/Kubernetes Experience with containerization and orchestration platforms
- Cloud Platforms Familiarity with AWS, Azure, or GCP
- ITIL Certification Understanding of ITIL processes (incident, problem, change management)
- Security Concepts Familiarity with authentication, authorization, and security best practices
- Bash/Shell Scripting Proficiency in Bash or Shell scripting for automation
Soft Skills
- Excellent Communication: Clear and professional verbal and written communication in English.
- Structured Mindset: Highly organized with strong attention to detail and ability to prioritize effectively.
- Problem-Solving: Strong analytical and troubleshooting skills with a proactive approach to issue resolution.
- Knowledge Sharing: Willingness to share knowledge and contribute to team development.
- Customer Focus: Strong customer service orientation and commitment to user satisfaction.
- Flexibility: Willingness to work across morning, mid, and night shifts as required.
Shift Coverage Requirement
- Morning Shift: 6:00 AM – 2:00 PM
- Mid Shift: 2:00 PM – 10:00 PM
- Night Shift: 10:00 PM – 6:00 AM
- Flexibility: The role requires availability to cover all three shifts on a rotational basis, with shift schedules determined by business needs.
- Global Support: This 24/7 coverage model ensures seamless support for international teams and users across different time zones.
#LI-GA1, #LI-ONSITE
Pay Range: Based on Experience
Already have an account? Log in here