EU remote
Enterprise Architect for AI Cloud Infrastructure (m/f/d)
About this role
NVIDIA and Deutsche Telekom are jointly developing the world’s first industrial AI cloud for European manufacturers. This AI factory in Germany will host 10,000 GPUs across NVIDIA DGX B200 systems and RTX Pro Servers. Deutsche Telekom provides secure, sovereign and fast infrastructure, including data centers, operations, security, and AI solutions. Role Overview: We are seeking an Enterprise Architect for Network Infrastructure at Industrial AI Cloud to design, build, automate network platform for automation and operation related network components such as Switches, Firewalls, Routers, Border Gateways as part of core environment of the Industrial AI Cloud.
In this role you will design, provision and manage above mentioned stack, implement and fine-tune monitoring, and deploy additional components if necessary. You’ll be working and coordinating between multiple teams (such as Infrastructure, Platform) to deliver and continuously improve infrastructure services following ITIL processes. Enterprise Architect Considers and defines design to enable automated configuration management, release management, build, test and deployment activities.
This is a customer facing role/ tailor made solutions and implementations for the customer including consultancy. Proprietary technologies used for managing above scope: InfiniBand, Cumullus OS, RoCE, UFM, FortiGate friewalls, Cisco Border gateways. WHAT WILL YOU DO? Coordinate Operations together with Data Center, IaaS & PaaS layer: Coordinate and support network lifecycle activities (installs, upgrades, changes, firmware updates) and manage /network interconnections and related documentation Switch & Firewall Management: Provision and maintain InfiniBand switches according to ITIL Standards Automation: Develop and maintain automation scripts to orchestrate overall scope.
Fine tuning, configuration changes through whole project lifetime OS & Firmware Management: Maintain network-based environments, apply patches, and manage firmware upgrades at scale. Monitoring & Observability: ITIL Processes: Follow and improve incident, problem, and change management workflows; document runbooks and standard operating procedures. Adhere to ZERO Outage guidelines. Cross-Team Collaboration: Work closely with Platform Engineers and AI solution teams to ensure smooth deployments and operations.