UK remote
Network Operations Engineer
About this role
Qube Research & Technologies (QRT) is a global quantitative and systematic investment manager, operating in all liquid asset classes across the world. We are a technology and data driven group implementing a scientific approach to investing. Combining data, research, technology and trading expertise has shaped QRT’s collaborative mindset which enables us to solve the most complex challenges. QRT’s culture of innovation continuously drives our ambition to deliver high quality returns for our investors.
Your future role within QRT: The successful candidate will join the global infrastructure team at QRT and you will be responsible for managing and maintaining the QRT Network Infrastructure and Automations for the firm. We implement a strict IAC (Infrastructure as Code) mindsight at QRT, and therefore it is crucial that any candidates must be able to demonstrate operation of complex infrastructure at scale, which have been delivered through a software and automation driven mindset.
The team will be focused on maintaining the service availability, reliability, and security of the platform. Developing observability tooling, self-healing/event driven automations, and performing troubleshooting activities. Providing on-going support and operational improvements for existing network infrastructure, whilst building our capability to operate a high-performance compute datacentre. Your present skillsets: Experience monitoring and resolving incidents across Low Latency LAN, Datacentre LAN, WAN transit, and Internet /cloud connectivity.
Expert troubleshooting skills of network infrastructure and collaborating with vendor support teams when required to perform deep investigation. Oversight and development of monitoring dashboards, responding proactively to alerts with appropriate level of priority. Taking responsibility for continuous improvement of alerting and the overall observability stack. Running post incident reviews finding opportunities to improve the availability and reliability of our infrastructure offerings.
Delivery of BAU changes through automation wherever possible, working with research and other infrastructure engineering teams in a highly collaborative manner. Perform trend analysis using various data sources aiming to seek out potential issues, improve correlation and finding capacity concerns. Defining SLOs to ensure the high availability of network services and infrastructure. Ability to provide considerable support of network automation and associated tooling such as CI/CD pipelines, orchestration, Ansible, Python, GitOps practices, amongst others.