Leads customer implementations of LLM and generative-AI systems from architecture through production, building full-stack solutions and improving model quality.
About this role
This role sits within the AI & HPC Infrastructure division of a technology organization building large-scale AI data centers and computing platforms. The team covers GPU servers, high-speed storage, network technologies, liquid-cooling facilities, Kubernetes-based GPU cluster management, and supporting software. The position focuses on the network architecture that connects these components and supports an external GPU cloud service.
The work spans architecture design, validation, construction, and operational improvement for large-scale data-center networks. It includes selecting switches, NICs, optical transceivers, and cable types; designing InfiniBand or RoCEv2 GPU interconnects; monitoring GPU-node communication; building EVPN/VXLAN overlays for multi-tenant separation; and designing physical network architectures based on GPU vendor references. The role also covers service specifications, SLA definition, and provisioning flows from application through tenant handover, including automation with Ansible, Terraform, or Python where applicable.
Applicants need experience designing, building, or operating large-scale Layer 2/Layer 3 data-center or carrier/ISP networks, knowledge of routing protocols such as BGP and OSPF, and experience coordinating specifications with vendors and internal teams. Depending on the role, approximately two or five years of relevant network design and construction experience is required.
This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.