You’ll join a collaborative, dynamic, and innovative team that provides 24x7x365 coverage for Client’s technology environment. You’ll work closely with teams across Engineering, including application developers, systems engineers, Incident Management, and Site Reliability Engineers.
You’ll play an important role in incident response and in the design, implementation, automation, and support of operational procedures across the platforms that make up Bloomberg’s computing environment. You’ll also have opportunities to contribute to projects that improve reliability, reduce operational toil, modernize tooling, and evolve the processes used by Triage Operations.
Responsibilities:
- Improve operational efficiency by identifying opportunities for automation, standardization, and process improvement
- Use strong technical, troubleshooting, multitasking, and interpersonal skills to help drive incidents through investigation and resolution
- Proactively investigate issues using logs, metrics, dashboards, dependency information, and other available telemetry
- Work effectively both independently and collaboratively in a fast-paced environment
- Develop, maintain, and improve runbooks, procedures, and other technical documentation
- Identify recurring operational problems and work with Engineering and SRE teams to drive long‑term improvements
- Stay current with evolving technologies and identify tools and strategies that can improve automation, troubleshooting, and operational efficiency
- Learn and adopt new workflows and technologies as the Triage Operations Engineer role continues to evolve
- Communicate clearly during incidents, providing useful technical findings and status to partner teams
Requirements:
- 2+ years of relevant experience in computer operations, production support, systems administration, SRE, or a related technical field
- Experience working with Unix/Linux environments, including RHEL, Ubuntu or similar distributions
- An understanding of datacenter infrastructure, distributed systems, and data communications
- A degree in Computer Science, Engineering, Mathematics, or a related field, or equivalent practical experience
- Experience using AI productivity and agentic tools such as Claude, ChatGPT, or Copilot
- Experience with orchestration and configuration‑management technologies such as Salt or Ansible, with Salt experience preferred
- Strong written and verbal communication, interpersonal, and customer‑service skills
- Working knowledge of Jira and software development lifecycle (SDLC) concepts
- Experience performing and automating day‑to‑day systems administration tasks
- Scripting or programming experience using Shell, Python, or similar languages
- Experience using observability, log‑analysis, metrics, and visualization platforms such as Grafana or Humio (CrowdStrike Falcon LogScale)
- Familiarity with development, automation, and management tools such as GitHub, Artifactory, Chef, Ruby, or similar technologies
- An interest in using automation and AI to reduce operational toil and improve incident response