Technical Solutions Architect
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
Technical Enablement Architect – GPU Platform
Location: United States (SF Bay Area Preferred, Remote OK)
The Role
We are seeking a Technical Enablement Architect – Rafay GPU Platform to join our product team. You will work closely with product managers, engineering, technical publications, and field solution architects to define, validate, and document how customers deploy and operate Rafay’s GPU Platform.
This is a very hands-on and technical role focused on turning product capabilities into real-world architectures, deployment patterns, use cases, and troubleshooting guidance. You will build reference environments, validate end-to-end deployment scenarios, document architecture and operational best practices, and help customers understand how to successfully deploy GPU infrastructure and AI workloads using Rafay.
The ideal candidate combines strong infrastructure and cloud-native architecture skills with the ability to communicate complex technical concepts clearly. You should be comfortable working across GPU infrastructure, networking, storage, virtualization, Kubernetes and AI/ML workloads.
You will have the opportunity to work on cutting-edge infrastructure spanning GPUs, AI/ML, Generative AI, Kubernetes, virtualization, and modern data center technologies.
Key Responsibilities
Solution Architecture & Reference Architectures
- Design and document GPU PaaS reference architectures for cloud providers, enterprises, service providers, and AI infrastructure operators.
- Develop architecture patterns covering GPU compute, Kubernetes, virtual machines, bare metal, networking, storage, security, observability, and AI workloads.
- Create detailed architecture diagrams illustrating platform components, infrastructure dependencies, traffic flows, integrations, and deployment models.
- Build and maintain reproducible reference environments that demonstrate recommended deployment patterns.
- Define architecture guidance for production considerations including scalability, high availability, multi-tenancy, security, and operational resilience.
Deployment Guides & Implementation Patterns
- Develop comprehensive, hands-on deployment guides that help customers move from infrastructure prerequisites to production-ready environments.
- Document infrastructure prerequisites, installation workflows, configuration options, integrations, validation steps, and operational best practices.
- Create deployment patterns for common GPU PaaS environments, including Kubernetes clusters, GPU VMs, bare-metal GPU servers, AI development environments, and inference services.
- Develop configuration examples using technologies such as Kubernetes YAML, Helm, Terraform, APIs, CLI tools, and automation frameworks.
- Validate documented procedures against real environments to ensure they are accurate and reproducible.
GPU PaaS Use Cases
- Identify and document common customer use cases and translate them into repeatable solution architectures and implementation patterns.
- Explain architectural tradeoffs and provide guidance on selecting the appropriate deployment model for different workloads.
Troubleshooting & Operational Guidance
- Develop troubleshooting guides, runbooks, and diagnostic workflows for common deployment and operational issues.
- Reproduce customer and field issues in reference environments to understand root causes and document resolution procedures.
- Create troubleshooting decision trees covering infrastructure, Kubernetes, GPU drivers/operators, networking, storage, scheduling, and AI workloads.
- Work with engineering and support teams to identify recurring issues and convert lessons learned into reusable operational guidance.
- Document health checks, validation procedures, logs, metrics, and diagnostic commands customers can use to operate environments effectively.&l