Network Solutions Consultant, AI/GPU Infrastructure
1 week ago
Houston, Texas, United States
AMSYS Innovative Solutions
Remote
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
Work Arrangement: U.
S. Remote | Travel Requirement: 50% (East Coast client sites) We're looking for a senior network architect to design and validate the fabrics behind large-scale GPU clusters for a global technology leader's AI and Hybrid Cloud engineering team. You'll build reference architectures, deployment runbooks, and performance baselines for AI training, inference, and HPC environments using NVIDIA Spectrum-X Ethernet, Quantum InfiniBand, and BlueField-3 DPUs. Your designs become the playbook that pre-sales, professional services, and delivery teams take into the field. This is a hands-on, customer-facing role, working closely with product engineering on high-density, liquid-cooled rack deployments. What you bring
• 10+ years designing and deploying network infrastructure for AI, HPC, or large GPU clusters
• Deep hands-on InfiniBand
experience:
rail-optimized topologies, Adaptive Routing, SHARP, and congestion control
• RoCEv2/RDMA tuning at scale, plus strong BGP, EVPN/VXLAN, and multi-tenant design
• Experience with BlueField DPUs and DOCA
• A track record of writing architecture docs and runbooks that engineers actually use Nice to have
• NVIDIA networking or InfiniBand certification, or CCIE/CCNP Data Center
• Automation with Ansible, Terraform, or Python
• NeoCloud or managed service provider background
• Exposure to FedRAMP or SOC 2 environments
S. Remote | Travel Requirement: 50% (East Coast client sites) We're looking for a senior network architect to design and validate the fabrics behind large-scale GPU clusters for a global technology leader's AI and Hybrid Cloud engineering team. You'll build reference architectures, deployment runbooks, and performance baselines for AI training, inference, and HPC environments using NVIDIA Spectrum-X Ethernet, Quantum InfiniBand, and BlueField-3 DPUs. Your designs become the playbook that pre-sales, professional services, and delivery teams take into the field. This is a hands-on, customer-facing role, working closely with product engineering on high-density, liquid-cooled rack deployments. What you bring
• 10+ years designing and deploying network infrastructure for AI, HPC, or large GPU clusters
• Deep hands-on InfiniBand
experience:
rail-optimized topologies, Adaptive Routing, SHARP, and congestion control
• RoCEv2/RDMA tuning at scale, plus strong BGP, EVPN/VXLAN, and multi-tenant design
• Experience with BlueField DPUs and DOCA
• A track record of writing architecture docs and runbooks that engineers actually use Nice to have
• NVIDIA networking or InfiniBand certification, or CCIE/CCNP Data Center
• Automation with Ansible, Terraform, or Python
• NeoCloud or managed service provider background
• Exposure to FedRAMP or SOC 2 environments