{"schemaVersion":"jobsearcher.job.v1","id":"4b9232422e8667bfde53ea0e","url":"https://jobsearcher.com/jobs/4b9232422e8667bfde53ea0e","canonicalUrl":"https://jobsearcher.com/jobs/4b9232422e8667bfde53ea0e","title":"Technical Solutions Architect – GPU Platform","description":"The Role\n\nWe are seeking a Technical Solutions Architect – Rafay GPU Platform to join our product team. You will work closely with product managers, engineering, technical publications, and field solution architects to define, validate, and document how customers deploy and operate Rafay’s GPU Platform.\n\nThis is a very hands-on and technical role focused on turning product capabilities into real-world architectures, deployment patterns, use cases, and troubleshooting guidance. You will build reference environments, validate end-to-end deployment scenarios, document architecture and operational best practices, and help customers understand how to successfully deploy GPU infrastructure and AI workloads using Rafay.\n\nThe ideal candidate combines strong infrastructure and cloud-native architecture skills with the ability to communicate complex technical concepts clearly. You should be comfortable working across GPU infrastructure, networking, storage, virtualization, Kubernetes and AI/ML workloads.\n\nYou will have the opportunity to work on cutting-edge infrastructure spanning GPUs, AI/ML, Generative AI, Kubernetes, virtualization, and modern data center technologies.\n\nKey Responsibilities\nSolution Architecture & Reference Architectures\nDesign and document GPU PaaS reference architectures for cloud providers, enterprises, service providers, and AI infrastructure operators.\nDevelop architecture patterns covering GPU compute, Kubernetes, virtual machines, bare metal, networking, storage, security, observability, and AI workloads.\nCreate detailed architecture diagrams illustrating platform components, infrastructure dependencies, traffic flows, integrations, and deployment models.\nBuild and maintain reproducible reference environments that demonstrate recommended deployment patterns.\nDefine architecture guidance for production considerations including scalability, high availability, multi-tenancy, security, and operational resilience.\n\n‍\n\nDeployment Guides & Implementation Patterns\nDevelop comprehensive, hands-on deployment guides that help customers move from infrastructure prerequisites to production-ready environments.\nDocument infrastructure prerequisites, installation workflows, configuration options, integrations, validation steps, and operational best practices.\nCreate deployment patterns for common GPU PaaS environments, including Kubernetes clusters, GPU VMs, bare-metal GPU servers, AI development environments, and inference services.\nDevelop configuration examples using technologies such as Kubernetes YAML, Helm, Terraform, APIs, CLI tools, and automation frameworks.\nValidate documented procedures against real environments to ensure they are accurate and reproducible.\nGPU PaaS Use Cases\nIdentify and document common customer use cases and translate them into repeatable solution architectures and implementation patterns.\nExplain architectural tradeoffs and provide guidance on selecting the appropriate deployment model for different workloads.\nTroubleshooting & Operational Guidance\nDevelop troubleshooting guides, runbooks, and diagnostic workflows for common deployment and operational issues.\nReproduce customer and field issues in reference environments to understand root causes and document resolution procedures.\nCreate troubleshooting decision trees covering infrastructure, Kubernetes, GPU drivers/operators, networking, storage, scheduling, and AI workloads.\nWork with engineering and support teams to identify recurring issues and convert lessons learned into reusable operational guidance.\nDocument health checks, validation procedures, logs, metrics, and diagnostic commands customers can use to operate environments effectively.\nCustomer & Field Intelligence\nGather feedback from customer deployments, POCs, support cases, and field engagements.\nIdentify recurring architecture patterns, deployment challenges, and operational issues.\nTranslate field learnings into improved reference architectures, deployment guides, troubleshooting content, and product requirements.\nPartner with product management to identify gaps that can be addressed through product enhancements or improved operational workflows.\nTechnical Demonstrations & Enablement\nBuild and maintain demo environments showcasing GPU PaaS architectures and use cases.\nCreate technical demonstrations and walkthroughs that explain deployment workflows and architecture concepts.\nDevelop presentations and enablement materials for solution architects, partners, customers, and internal teams.\nRecord technical demos and deployment walkthroughs for training and product launches.\nContribute hands-on technical blogs and architecture content highlighting real-world implementation patterns.\nTechnical Qualifications\n3–6+ years of experience in solutions architecture, solutions engineering, infrastructure engineering, DevOps, platform engineering, or a related technical role.\nStrong understanding of Kubernetes and cloud-native infrastructure.\nExperience designing or operating infrastructure across public cloud, private cloud, or data center environments.\nUnderstanding of GPU infrastructure and AI/ML workloads, including GPU scheduling, drivers, operators, resource allocation, and workload lifecycle.\nFamiliarity with virtualization, bare-metal infrastructure, networking, storage, and infrastructure automation.\nExperience with technologies such as:\n\nKubernetes, Helm, and container runtimes\nTerraform and infrastructure-as-code\nPython, Bash, YAML, and REST APIs\nGPU infrastructure and NVIDIA software stacks\nNetworking and load balancing\nPersistent storage and CSI-based storage platforms\nObservability, metrics, logging, and troubleshooting tools\nAbility to build environments, deploy workloads, troubleshoot failures, and document the complete process.\n\nExperience with technologies such as Slurm, KubeVirt, NVIDIA GPU Operator, NVIDIA DRA, MIG, vLLM, Ray, Kubeflow, or AI inference platforms is a plus.\n\nCommunication & Technical Writing\nExceptional written and verbal communication skills for highly technical audiences.\nAbility to translate complex infrastructure concepts into clear architectures, deployment procedures, and operational guidance.\nStrong ability to create architecture diagrams and visually communicate system designs.\nExperience creating technical presentations, demonstrations, deployment guides, or reference architectures.\nComfortable explaining both how a solution works and why a particular architecture should be used.\nCollaboration & Execution\nComfortable working cross-functionally with product management, engineering, support, technical publications, solution architects, partners, and customers.\nStrong hands-on problem-solving and troubleshooting skills.\nAbility to independently build and validate technical environments.\nProven ability to manage multiple technical deliverables in a fast-moving environment.\nSelf-starter who enjoys exploring new infrastructure technologies and turning that knowledge into practical guidance for customers.\nWhy Join Us\nWork at the intersection of GPU infrastructure, Kubernetes, cloud platforms, and AI.\nDefine how customers architect and deploy production GPU PaaS environments.\nBuild reference architectures and deployment patterns that directly influence customer success.\nWork hands-on with emerging GPU, AI, networking, storage, and cloud-native technologies.\nPartner closely with engineering and product teams and help shape the evolution of the Rafay platform.","company":"Rafay Systems","rawCompany":"rafay systems","city":"Alameda","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-10T11:20:27.400Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Technical Solutions Architect – GPU Platform","description":"The Role\n\nWe are seeking a Technical Solutions Architect – Rafay GPU Platform to join our product team. You will work closely with product managers, engineering, technical publications, and field solution architects to define, validate, and document how customers deploy and operate Rafay’s GPU Platform.\n\nThis is a very hands-on and technical role focused on turning product capabilities into real-world architectures, deployment patterns, use cases, and troubleshooting guidance. You will build reference environments, validate end-to-end deployment scenarios, document architecture and operational best practices, and help customers understand how to successfully deploy GPU infrastructure and AI workloads using Rafay.\n\nThe ideal candidate combines strong infrastructure and cloud-native architecture skills with the ability to communicate complex technical concepts clearly. You should be comfortable working across GPU infrastructure, networking, storage, virtualization, Kubernetes and AI/ML workloads.\n\nYou will have the opportunity to work on cutting-edge infrastructure spanning GPUs, AI/ML, Generative AI, Kubernetes, virtualization, and modern data center technologies.\n\nKey Responsibilities\nSolution Architecture & Reference Architectures\nDesign and document GPU PaaS reference architectures for cloud providers, enterprises, service providers, and AI infrastructure operators.\nDevelop architecture patterns covering GPU compute, Kubernetes, virtual machines, bare metal, networking, storage, security, observability, and AI workloads.\nCreate detailed architecture diagrams illustrating platform components, infrastructure dependencies, traffic flows, integrations, and deployment models.\nBuild and maintain reproducible reference environments that demonstrate recommended deployment patterns.\nDefine architecture guidance for production considerations including scalability, high availability, multi-tenancy, security, and operational resilience.\n\n‍\n\nDeployment Guides & Implementation Patterns\nDevelop comprehensive, hands-on deployment guides that help customers move from infrastructure prerequisites to production-ready environments.\nDocument infrastructure prerequisites, installation workflows, configuration options, integrations, validation steps, and operational best practices.\nCreate deployment patterns for common GPU PaaS environments, including Kubernetes clusters, GPU VMs, bare-metal GPU servers, AI development environments, and inference services.\nDevelop configuration examples using technologies such as Kubernetes YAML, Helm, Terraform, APIs, CLI tools, and automation frameworks.\nValidate documented procedures against real environments to ensure they are accurate and reproducible.\nGPU PaaS Use Cases\nIdentify and document common customer use cases and translate them into repeatable solution architectures and implementation patterns.\nExplain architectural tradeoffs and provide guidance on selecting the appropriate deployment model for different workloads.\nTroubleshooting & Operational Guidance\nDevelop troubleshooting guides, runbooks, and diagnostic workflows for common deployment and operational issues.\nReproduce customer and field issues in reference environments to understand root causes and document resolution procedures.\nCreate troubleshooting decision trees covering infrastructure, Kubernetes, GPU drivers/operators, networking, storage, scheduling, and AI workloads.\nWork with engineering and support teams to identify recurring issues and convert lessons learned into reusable operational guidance.\nDocument health checks, validation procedures, logs, metrics, and diagnostic commands customers can use to operate environments effectively.\nCustomer & Field Intelligence\nGather feedback from customer deployments, POCs, support cases, and field engagements.\nIdentify recurring architecture patterns, deployment challenges, and operational issues.\nTranslate field learnings into improved reference architectures, deployment guides, troubleshooting content, and product requirements.\nPartner with product management to identify gaps that can be addressed through product enhancements or improved operational workflows.\nTechnical Demonstrations & Enablement\nBuild and maintain demo environments showcasing GPU PaaS architectures and use cases.\nCreate technical demonstrations and walkthroughs that explain deployment workflows and architecture concepts.\nDevelop presentations and enablement materials for solution architects, partners, customers, and internal teams.\nRecord technical demos and deployment walkthroughs for training and product launches.\nContribute hands-on technical blogs and architecture content highlighting real-world implementation patterns.\nTechnical Qualifications\n3–6+ years of experience in solutions architecture, solutions engineering, infrastructure engineering, DevOps, platform engineering, or a related technical role.\nStrong understanding of Kubernetes and cloud-native infrastructure.\nExperience designing or operating infrastructure across public cloud, private cloud, or data center environments.\nUnderstanding of GPU infrastructure and AI/ML workloads, including GPU scheduling, drivers, operators, resource allocation, and workload lifecycle.\nFamiliarity with virtualization, bare-metal infrastructure, networking, storage, and infrastructure automation.\nExperience with technologies such as:\n\nKubernetes, Helm, and container runtimes\nTerraform and infrastructure-as-code\nPython, Bash, YAML, and REST APIs\nGPU infrastructure and NVIDIA software stacks\nNetworking and load balancing\nPersistent storage and CSI-based storage platforms\nObservability, metrics, logging, and troubleshooting tools\nAbility to build environments, deploy workloads, troubleshoot failures, and document the complete process.\n\nExperience with technologies such as Slurm, KubeVirt, NVIDIA GPU Operator, NVIDIA DRA, MIG, vLLM, Ray, Kubeflow, or AI inference platforms is a plus.\n\nCommunication & Technical Writing\nExceptional written and verbal communication skills for highly technical audiences.\nAbility to translate complex infrastructure concepts into clear architectures, deployment procedures, and operational guidance.\nStrong ability to create architecture diagrams and visually communicate system designs.\nExperience creating technical presentations, demonstrations, deployment guides, or reference architectures.\nComfortable explaining both how a solution works and why a particular architecture should be used.\nCollaboration & Execution\nComfortable working cross-functionally with product management, engineering, support, technical publications, solution architects, partners, and customers.\nStrong hands-on problem-solving and troubleshooting skills.\nAbility to independently build and validate technical environments.\nProven ability to manage multiple technical deliverables in a fast-moving environment.\nSelf-starter who enjoys exploring new infrastructure technologies and turning that knowledge into practical guidance for customers.\nWhy Join Us\nWork at the intersection of GPU infrastructure, Kubernetes, cloud platforms, and AI.\nDefine how customers architect and deploy production GPU PaaS environments.\nBuild reference architectures and deployment patterns that directly influence customer success.\nWork hands-on with emerging GPU, AI, networking, storage, and cloud-native technologies.\nPartner closely with engineering and product teams and help shape the evolution of the Rafay platform.","datePosted":"2026-09-10T11:20:27.400Z","dateModified":"2026-09-10T11:20:27.400Z","hiringOrganization":{"@type":"Organization","name":"Rafay Systems","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Alameda","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"4b9232422e8667bfde53ea0e"},"url":"https://jobsearcher.com/jobs/4b9232422e8667bfde53ea0e"}}