{"schemaVersion":"jobsearcher.job.v1","id":"a0e3de7c3988e87f6d2edf5d","url":"https://jobsearcher.com/jobs/a0e3de7c3988e87f6d2edf5d","canonicalUrl":"https://jobsearcher.com/jobs/a0e3de7c3988e87f6d2edf5d","title":"Principal Solutions Architect","description":"Overview\nWe are seeking an elite Solutions Architect to lead the end-to-end design, sizing, and deployment of NVIDIA AI Factory-aligned infrastructure. In this highly technical, customer-facing role you will translate complex AI and machine learning workload requirements into fully engineered infrastructure solutions spanning colocation facilities, GPU compute, high-performance networking, parallel storage, and the complete NVIDIA AI software stack.\nYou will serve as a trusted technical advisor to enterprise and hyperscale customers, partnering with sales, product, and engineering teams to win and deliver transformational AI infrastructure programs. Your expertise will directly shape how organizations build and operate production AI Factories capable of training frontier models, running large-scale inference fleets, and accelerating data science pipelines at scale.\nYour Impact\nSolution Design & Architecture\nLead discovery workshops to capture AI/ML workload requirements, including model training scale, inference SLAs, data pipeline throughput, and multi-tenancy needs.\nArchitect full-stack AI Factory solutions aligned to NVIDIA reference architectures, integrating colocation, GPU compute, networking, storage, and software layers.\nDevelop detailed Bills of Materials (BOMs), rack elevation diagrams, network topology drawings, and power/cooling budgets for customer proposals.\nDefine GPU cluster architectures using NVIDIA DGX, HGX, and MGX systems with B200, B300, and GB300 Blackwell SXM and NVLink-Switch configurations.\nDesign RTX PRO 6000 Blackwell Server Edition deployments for inference-optimized and enterprise AI workloads.\nConduct workload sizing and TCO/ROI modeling to validate infrastructure dimensioning for training, finetuning, and inference at scale.\nColocation & Facility Planning\nSpecify colocation requirements including critical power load (MW-scale), UPS and generator configurations, and PUE targets.\nDesign high-density GPU deployments utilizing air-cooled, direct liquid cooling (DLC), and rear-door heat exchanger configurations.\nDefine meet-me room (MMR) and cross-connect requirements; specify carrier-neutral telecom diversity strategies.\nEngage colocation providers and data center operators to validate capacity availability and negotiate technical SLAs.\nCoordinate with facilities and MEP engineers to validate power infrastructure from utility feed through PDU to rack level.\nGPU Compute Infrastructure\nArchitect multi-node GPU clusters optimized for large language model (LLM) pre-training, fine-tuning, and reinforcement learning from human feedback (RLHF).\nSize and configure DGX SuperPOD, HGX H/B-series, and MGX modular systems based on model parameter count, dataset size, and iteration timelines.\nDefine server firmware, BIOS, BMC, and DGXOS baselines for production GPU infrastructure.\nEstablish GPU health monitoring, RAS (Reliability, Availability, Serviceability) policies, and lifecycle management procedures.\nHigh-Performance Networking\nDesign backend GPU fabric networks using NVIDIA Quantum InfiniBand (NDR 400Gb/s and HDR 200Gb/s) for distributed training traffic.\nArchitect Spectrum-X Ethernet-based AI networking solutions for inference clusters requiring highbandwidth, low-latency connectivity.\nSpecify ConnectX-8/7 HCA deployments and configure RDMA over Converged Ethernet (RoCEv2) or InfiniBand transport for NCCL collective operations.\nIntegrate BlueField-3 DPUs for GPU-accelerated network functions, storage offload, zero-trust security isolation, and bare-metal provisioning.\nDesign leaf-spine and fat-tree topologies for non-blocking bisectional bandwidth in GPU training clusters.\nDefine Quality of Service (QoS) policies separating storage, compute fabric, and management plane traffic.\nParallel Storage Architecture\nDesign high-performance parallel file system solutions using VAST Data, Hammerspace, and Pure Storage FlashBlade//E for AI training and checkpoint storage.\nSize storage capacity, IOPS, and throughput based on dataset characteristics, checkpoint frequency, and concurrent reader/writer counts.\nArchitect multi-tier storage hierarchies: hot NVMe flash (VAST/FlashBlade) for active datasets, warm object storage for model archives, and cold tape/cloud for long-term retention.\nConfigure VAST Data Universal Storage for disaggregated storage with NFS, S3, and POSIX access; tune for large sequential read performance.\nDeploy Hammerspace Global Data Environment for distributed data management and NFS-over-RDMA acceleration across geographically dispersed GPU clusters.\nDefine data pipeline architectures ingesting from cloud object stores (S3, GCS, ABS) to local flash for GPUlocal data loading without I/O bottlenecks.\nAI Software Stack & Orchestration\nDeploy and configure NVIDIA AI Enterprise (NVAIE) software stack including NVIDIA GPU Operator, NIM microservices, and RAPIDS accelerated data science libraries.\nArchitect inference serving infrastructure using NVIDIA NIM (NVIDIA Inference Microservices) for optimized LLM and vision model deployment with autoscaling.\nImplement NVIDIA Dynamo for distributed inference and disaggregated serving of large-scale generative AI models.\nConfigure and optimize CUDA toolkit, cuDNN, NCCL communication libraries, and custom kernel environments for training workloads.\nDeploy Base Command Manager and DGXOS for cluster lifecycle management, node provisioning, health dashboards, and job scheduling integration.\nIntegrate NVIDIA Mission Control for AI Factory operations, observability, and multi-cluster fleet management.\nDesign and deploy Kubernetes-based AI platforms using NVIDIA GPU Operator, integrating with Run:ai for dynamic GPU resource scheduling and multi-tenant workload isolation.\nConfigure SLURM workload manager for traditional HPC-style job scheduling on bare-metal GPU clusters, including preemption policies, fair-share scheduling, and burst-to-cloud integration.\nEstablish MLOps toolchain integrations with popular frameworks (PyTorch, JAX, TensorFlow) and experiment tracking platforms (MLflow, Weights & Biases).\n\nCustomer Engagement & Delivery\nServe as primary technical point of contact throughout the pre-sales and delivery lifecycle, from initial discovery through post-deployment optimization.\nProduce and present architecture design documents, technical proposals, and executive-level briefings to CTO/CIO and VP-level stakeholders.\nLead proof-of-concept (POC) and pilot deployments, including benchmark design, execution, and results analysis.\nCollaborate with procurement, logistics, and deployment teams to ensure on-time delivery of complex infrastructure programs.\nProvide post-deployment hypercare support, performance tuning, and capacity planning advisory services.\nContribute to internal knowledge bases, solution playbooks, and reference architectures for repeatable AI Factory deployments.\nTechnology Stack\nCandidates must demonstrate deep, hands-on expertise across the following technology domains:\nGPU Compute\n\nDGX B200 / B300, DGX H100 / H200, HGX B200 / B300, HGX H100 / H200,\nMGX platforms, GB300 NVL72 / GB200 NVL72, RTX PRO 6000 Blackwell\nServer Edition, NVLink Switch System, NVLink-C2C\n\nNetworking\n\nNVIDIA Quantum InfiniBand (NDR 400G, HDR 200G), Spectrum-X Ethernet, ConnectX-8 / ConnectX-7 HCAs, BlueField-3 DPU, SHARP in-network computing, UFM Fabric Manager, RDMA / RoCEv2 / InfiniBand\n\nStorage\n\nVAST Data Universal Storage (NFS/S3/POSIX), Hammerspace Global Data Environment, Pure Storage FlashBlade//E (Evergreen//One), NFS-over-RDMA, parallel file systems (Lustre, GPFS/WEKA), S3-compatible object storage\n\nAI Software\nNVIDIA AI Enterprise (NVAIE), NIM Microservices, RAPIDS (cuDF, cuML, cuGraph), NVIDIA Dynamo, CUDA Toolkit, cuDNN, NCCL, TensorRT, Triton Inference Server\n\nCluster Mgmt\n\nBase Command Manager, DGXOS, NVIDIA Mission Control, DGX Cloud, UFM, IPMI / Redfish BMC management\n\nOrchestration\n\nKubernetes (K8s), NVIDIA GPU Operator, Run:ai GPU scheduling, SLURM, OpenMPI, Helm, Argo Workflows, Kubeflow, KServe\n\nColocation\n\nCritical power design (kW – MW), UPS / generator, CRAC / CRAH / DLC / immersion cooling, hot-aisle containment, PUE optimization, carrier-neutral telecom, cross-connects, MMR design\n\nFrameworks\n\nPyTorch, JAX, TensorFlow, Hugging Face Transformers, DeepSpeed, Megatron-LM, vLLM, LMDeploy\n\nQualifications\nBachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred.\n8+ years of solutions architecture, systems engineering, or technical pre-sales experience, with at least 4 years focused on GPU infrastructure or HPC environments.\nProven track record designing and deploying NVIDIA DGX or HGX-based GPU clusters in production AI/ML environments.\nDeep understanding of distributed deep learning concepts: tensor parallelism, pipeline parallelism, data parallelism, gradient checkpointing, and mixed-precision training.\nHands-on experience with InfiniBand or high-speed Ethernet fabric design, RDMA configuration, and collective communication tuning (NCCL, MPI).\nDirect experience sizing and deploying parallel storage systems (VAST, Hammerspace, or Lustre/WEKA/GPFS) for AI training workloads.\nStrong working knowledge of Kubernetes, GPU Operator, and at least one GPU workload scheduler (Run:ai or SLURM).\nExperience with Linux system administration, CUDA development environment configuration, and GPU driver/firmware management.\nDemonstrated ability to create compelling technical proposals, architecture diagrams (Visio/Lucidchart/draw.io), and BOM-level documentation.\nExceptional communication skills with proven ability to present to both deep technical audiences and Clevel executives.\nPreferred Qualifications:\nNVIDIA-certified professional credentials (DCA-Core, NCP-DS, or equivalent).\nExperience with NVIDIA Base Command Platform or Mission Control for multi-cluster AI Factory operations.\nFamiliarity with sovereign AI, government cloud, or regulated industry AI infrastructure requirements.\nExperience integrating AI Factory infrastructure with public cloud (AWS, Azure, GCP) for hybrid and burstto-cloud architectures.\nBackground in MLOps, LLMOps, or platform engineering for production AI model lifecycle management.\nPrior experience with colocation data center procurement, RFP development, and SLA negotiation.\nContributions to open-source AI infrastructure projects or published technical content (blogs, whitepapers, conference presentations).\nActive participation in the NVIDIA Partner Network (NPN) ecosystem or prior experience at an NVIDIA Elite Solution Provider.\nCore Competencies\nTechnical Depth\nEnd-to-end AI infrastructure expertise from silicon to software; ability to go deep on any layer of the stack.\n\nSystems Thinking\nAbility to reason holistically about performance, reliability, power, cost, and operability trade-offs across complex integrated systems.\n\nCustomer Obsession\nRelentless focus on understanding customer AI objectives and delivering solutions that accelerate time-to-value.\n\nExecutive Presence\nConfidence and clarity when presenting complex technical architectures to senior business and technology leaders.\n\nAnalytical Rigor\nData-driven approach to workload sizing, performance modeling, and TCO analysis with attention to detail.\n\nCollaborative Leadership\nAbility to lead cross-functional pursuit teams, align internal stakeholders, and orchestrate complex delivery programs.\n\nPosition Specifics\nThe initial base salary range for this position is expected to be between $170,000 and $190,000 annually. The final base salary offered will be determined by multiple factors, including, but not limited to, job-related knowledge, depth of experience, skills, certifications, and geographic location. In addition to the base salary, our compensation structure may include other components such as commissions and discretionary bonuses.\nePlus offers a full range of medical, financial, and/or other benefits (including 401(k) eligibility, employee stock purchase program and various paid time off benefits, such as vacation, sick time, and personal leave), dependent on the position offered. Details of participation in these benefit plans will be provided if an offer of employment is extended.\nIf hired, employee will be in an \"at-will position\" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.\n#LI-DY1\n#IND1\n\nWho We Are\nAt ePlus, we believe technology is a people business. Our team is passionate, skilled, and driven to deliver solutions that make a real difference. Join us and be part of a culture that values collaboration, innovation, and extraordinary results.\nCorporate Values\nRespectful communication and cooperation: We prioritize respectful communication, fostering an environment where everyone is treated with dignity and respect.\nTeamwork and employee participation: Collaboration and teamwork thrive through diverse perspectives, both within our teams and in our interactions with our customers.\nWork/life balance that supports our employees' varying needs: We value the well-being of our employees, recognizing that a healthy work-life balance is pivotal to our collective success.\nEmbracing communities: We embrace and support the communities that nurture us. Our employees' dedication to fostering positive change is a source of immense pride for us.\nCommitment to Diversity, Inclusion and Belonging\nWe are an equal opportunity employer that does not discriminate or allow discrimination based on race, color, religion, sex, sexual orientation, gender identity, age, national origin, citizenship, disability, veteran status, or any other classification protected by federal, state, or local law.\nePlus is dedicated to fostering, cultivating, and preserving a culture that represents diversity, enables inclusion, and makes our employees feel comfortable bringing their full, unique selves to work.\nPhysical Requirements\nWhile performing this role, you will engage in both seated and occasional standing or walking activities. We provide reasonable accommodations, in accordance with relevant laws, to support success in this position.\nBy embracing our values, you will contribute to our collective mission of making a positive impact within our organization and the broader community. We understand that this job description serves as a guide and is not an employment contract.\nePlus maintains a California Consumer Privacy Act (CCPA) Privacy Notice on our Trust Center, available here: CCPA Privacy Notice.\nNotice to Recruiting Agencies: ePlus only accepts unsolicited resumes when presented directly by a candidate. Unsolicited resumes submitted to ePlus from any other source will be considered ePlus property and will not qualify for any placement or referral fees. ePlus will only pay such fees in connection with a valid written agreement between ePlus and the referring agency, and then only after providing advance written approval to the referring agency to submit resumes in connection with a particular opportunity.","company":"Eplus","rawCompany":"eplus","city":"Irvine","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-03T19:34:42.593Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1241.00","title":"Computer Network Architects","slug":"computer-network-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541513","title":"Computer Facilities Management Services","slug":"computer-facilities-management-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Principal Solutions Architect","description":"Overview\nWe are seeking an elite Solutions Architect to lead the end-to-end design, sizing, and deployment of NVIDIA AI Factory-aligned infrastructure. In this highly technical, customer-facing role you will translate complex AI and machine learning workload requirements into fully engineered infrastructure solutions spanning colocation facilities, GPU compute, high-performance networking, parallel storage, and the complete NVIDIA AI software stack.\nYou will serve as a trusted technical advisor to enterprise and hyperscale customers, partnering with sales, product, and engineering teams to win and deliver transformational AI infrastructure programs. Your expertise will directly shape how organizations build and operate production AI Factories capable of training frontier models, running large-scale inference fleets, and accelerating data science pipelines at scale.\nYour Impact\nSolution Design & Architecture\nLead discovery workshops to capture AI/ML workload requirements, including model training scale, inference SLAs, data pipeline throughput, and multi-tenancy needs.\nArchitect full-stack AI Factory solutions aligned to NVIDIA reference architectures, integrating colocation, GPU compute, networking, storage, and software layers.\nDevelop detailed Bills of Materials (BOMs), rack elevation diagrams, network topology drawings, and power/cooling budgets for customer proposals.\nDefine GPU cluster architectures using NVIDIA DGX, HGX, and MGX systems with B200, B300, and GB300 Blackwell SXM and NVLink-Switch configurations.\nDesign RTX PRO 6000 Blackwell Server Edition deployments for inference-optimized and enterprise AI workloads.\nConduct workload sizing and TCO/ROI modeling to validate infrastructure dimensioning for training, finetuning, and inference at scale.\nColocation & Facility Planning\nSpecify colocation requirements including critical power load (MW-scale), UPS and generator configurations, and PUE targets.\nDesign high-density GPU deployments utilizing air-cooled, direct liquid cooling (DLC), and rear-door heat exchanger configurations.\nDefine meet-me room (MMR) and cross-connect requirements; specify carrier-neutral telecom diversity strategies.\nEngage colocation providers and data center operators to validate capacity availability and negotiate technical SLAs.\nCoordinate with facilities and MEP engineers to validate power infrastructure from utility feed through PDU to rack level.\nGPU Compute Infrastructure\nArchitect multi-node GPU clusters optimized for large language model (LLM) pre-training, fine-tuning, and reinforcement learning from human feedback (RLHF).\nSize and configure DGX SuperPOD, HGX H/B-series, and MGX modular systems based on model parameter count, dataset size, and iteration timelines.\nDefine server firmware, BIOS, BMC, and DGXOS baselines for production GPU infrastructure.\nEstablish GPU health monitoring, RAS (Reliability, Availability, Serviceability) policies, and lifecycle management procedures.\nHigh-Performance Networking\nDesign backend GPU fabric networks using NVIDIA Quantum InfiniBand (NDR 400Gb/s and HDR 200Gb/s) for distributed training traffic.\nArchitect Spectrum-X Ethernet-based AI networking solutions for inference clusters requiring highbandwidth, low-latency connectivity.\nSpecify ConnectX-8/7 HCA deployments and configure RDMA over Converged Ethernet (RoCEv2) or InfiniBand transport for NCCL collective operations.\nIntegrate BlueField-3 DPUs for GPU-accelerated network functions, storage offload, zero-trust security isolation, and bare-metal provisioning.\nDesign leaf-spine and fat-tree topologies for non-blocking bisectional bandwidth in GPU training clusters.\nDefine Quality of Service (QoS) policies separating storage, compute fabric, and management plane traffic.\nParallel Storage Architecture\nDesign high-performance parallel file system solutions using VAST Data, Hammerspace, and Pure Storage FlashBlade//E for AI training and checkpoint storage.\nSize storage capacity, IOPS, and throughput based on dataset characteristics, checkpoint frequency, and concurrent reader/writer counts.\nArchitect multi-tier storage hierarchies: hot NVMe flash (VAST/FlashBlade) for active datasets, warm object storage for model archives, and cold tape/cloud for long-term retention.\nConfigure VAST Data Universal Storage for disaggregated storage with NFS, S3, and POSIX access; tune for large sequential read performance.\nDeploy Hammerspace Global Data Environment for distributed data management and NFS-over-RDMA acceleration across geographically dispersed GPU clusters.\nDefine data pipeline architectures ingesting from cloud object stores (S3, GCS, ABS) to local flash for GPUlocal data loading without I/O bottlenecks.\nAI Software Stack & Orchestration\nDeploy and configure NVIDIA AI Enterprise (NVAIE) software stack including NVIDIA GPU Operator, NIM microservices, and RAPIDS accelerated data science libraries.\nArchitect inference serving infrastructure using NVIDIA NIM (NVIDIA Inference Microservices) for optimized LLM and vision model deployment with autoscaling.\nImplement NVIDIA Dynamo for distributed inference and disaggregated serving of large-scale generative AI models.\nConfigure and optimize CUDA toolkit, cuDNN, NCCL communication libraries, and custom kernel environments for training workloads.\nDeploy Base Command Manager and DGXOS for cluster lifecycle management, node provisioning, health dashboards, and job scheduling integration.\nIntegrate NVIDIA Mission Control for AI Factory operations, observability, and multi-cluster fleet management.\nDesign and deploy Kubernetes-based AI platforms using NVIDIA GPU Operator, integrating with Run:ai for dynamic GPU resource scheduling and multi-tenant workload isolation.\nConfigure SLURM workload manager for traditional HPC-style job scheduling on bare-metal GPU clusters, including preemption policies, fair-share scheduling, and burst-to-cloud integration.\nEstablish MLOps toolchain integrations with popular frameworks (PyTorch, JAX, TensorFlow) and experiment tracking platforms (MLflow, Weights & Biases).\n\nCustomer Engagement & Delivery\nServe as primary technical point of contact throughout the pre-sales and delivery lifecycle, from initial discovery through post-deployment optimization.\nProduce and present architecture design documents, technical proposals, and executive-level briefings to CTO/CIO and VP-level stakeholders.\nLead proof-of-concept (POC) and pilot deployments, including benchmark design, execution, and results analysis.\nCollaborate with procurement, logistics, and deployment teams to ensure on-time delivery of complex infrastructure programs.\nProvide post-deployment hypercare support, performance tuning, and capacity planning advisory services.\nContribute to internal knowledge bases, solution playbooks, and reference architectures for repeatable AI Factory deployments.\nTechnology Stack\nCandidates must demonstrate deep, hands-on expertise across the following technology domains:\nGPU Compute\n\nDGX B200 / B300, DGX H100 / H200, HGX B200 / B300, HGX H100 / H200,\nMGX platforms, GB300 NVL72 / GB200 NVL72, RTX PRO 6000 Blackwell\nServer Edition, NVLink Switch System, NVLink-C2C\n\nNetworking\n\nNVIDIA Quantum InfiniBand (NDR 400G, HDR 200G), Spectrum-X Ethernet, ConnectX-8 / ConnectX-7 HCAs, BlueField-3 DPU, SHARP in-network computing, UFM Fabric Manager, RDMA / RoCEv2 / InfiniBand\n\nStorage\n\nVAST Data Universal Storage (NFS/S3/POSIX), Hammerspace Global Data Environment, Pure Storage FlashBlade//E (Evergreen//One), NFS-over-RDMA, parallel file systems (Lustre, GPFS/WEKA), S3-compatible object storage\n\nAI Software\nNVIDIA AI Enterprise (NVAIE), NIM Microservices, RAPIDS (cuDF, cuML, cuGraph), NVIDIA Dynamo, CUDA Toolkit, cuDNN, NCCL, TensorRT, Triton Inference Server\n\nCluster Mgmt\n\nBase Command Manager, DGXOS, NVIDIA Mission Control, DGX Cloud, UFM, IPMI / Redfish BMC management\n\nOrchestration\n\nKubernetes (K8s), NVIDIA GPU Operator, Run:ai GPU scheduling, SLURM, OpenMPI, Helm, Argo Workflows, Kubeflow, KServe\n\nColocation\n\nCritical power design (kW – MW), UPS / generator, CRAC / CRAH / DLC / immersion cooling, hot-aisle containment, PUE optimization, carrier-neutral telecom, cross-connects, MMR design\n\nFrameworks\n\nPyTorch, JAX, TensorFlow, Hugging Face Transformers, DeepSpeed, Megatron-LM, vLLM, LMDeploy\n\nQualifications\nBachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred.\n8+ years of solutions architecture, systems engineering, or technical pre-sales experience, with at least 4 years focused on GPU infrastructure or HPC environments.\nProven track record designing and deploying NVIDIA DGX or HGX-based GPU clusters in production AI/ML environments.\nDeep understanding of distributed deep learning concepts: tensor parallelism, pipeline parallelism, data parallelism, gradient checkpointing, and mixed-precision training.\nHands-on experience with InfiniBand or high-speed Ethernet fabric design, RDMA configuration, and collective communication tuning (NCCL, MPI).\nDirect experience sizing and deploying parallel storage systems (VAST, Hammerspace, or Lustre/WEKA/GPFS) for AI training workloads.\nStrong working knowledge of Kubernetes, GPU Operator, and at least one GPU workload scheduler (Run:ai or SLURM).\nExperience with Linux system administration, CUDA development environment configuration, and GPU driver/firmware management.\nDemonstrated ability to create compelling technical proposals, architecture diagrams (Visio/Lucidchart/draw.io), and BOM-level documentation.\nExceptional communication skills with proven ability to present to both deep technical audiences and Clevel executives.\nPreferred Qualifications:\nNVIDIA-certified professional credentials (DCA-Core, NCP-DS, or equivalent).\nExperience with NVIDIA Base Command Platform or Mission Control for multi-cluster AI Factory operations.\nFamiliarity with sovereign AI, government cloud, or regulated industry AI infrastructure requirements.\nExperience integrating AI Factory infrastructure with public cloud (AWS, Azure, GCP) for hybrid and burstto-cloud architectures.\nBackground in MLOps, LLMOps, or platform engineering for production AI model lifecycle management.\nPrior experience with colocation data center procurement, RFP development, and SLA negotiation.\nContributions to open-source AI infrastructure projects or published technical content (blogs, whitepapers, conference presentations).\nActive participation in the NVIDIA Partner Network (NPN) ecosystem or prior experience at an NVIDIA Elite Solution Provider.\nCore Competencies\nTechnical Depth\nEnd-to-end AI infrastructure expertise from silicon to software; ability to go deep on any layer of the stack.\n\nSystems Thinking\nAbility to reason holistically about performance, reliability, power, cost, and operability trade-offs across complex integrated systems.\n\nCustomer Obsession\nRelentless focus on understanding customer AI objectives and delivering solutions that accelerate time-to-value.\n\nExecutive Presence\nConfidence and clarity when presenting complex technical architectures to senior business and technology leaders.\n\nAnalytical Rigor\nData-driven approach to workload sizing, performance modeling, and TCO analysis with attention to detail.\n\nCollaborative Leadership\nAbility to lead cross-functional pursuit teams, align internal stakeholders, and orchestrate complex delivery programs.\n\nPosition Specifics\nThe initial base salary range for this position is expected to be between $170,000 and $190,000 annually. The final base salary offered will be determined by multiple factors, including, but not limited to, job-related knowledge, depth of experience, skills, certifications, and geographic location. In addition to the base salary, our compensation structure may include other components such as commissions and discretionary bonuses.\nePlus offers a full range of medical, financial, and/or other benefits (including 401(k) eligibility, employee stock purchase program and various paid time off benefits, such as vacation, sick time, and personal leave), dependent on the position offered. Details of participation in these benefit plans will be provided if an offer of employment is extended.\nIf hired, employee will be in an \"at-will position\" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.\n#LI-DY1\n#IND1\n\nWho We Are\nAt ePlus, we believe technology is a people business. Our team is passionate, skilled, and driven to deliver solutions that make a real difference. Join us and be part of a culture that values collaboration, innovation, and extraordinary results.\nCorporate Values\nRespectful communication and cooperation: We prioritize respectful communication, fostering an environment where everyone is treated with dignity and respect.\nTeamwork and employee participation: Collaboration and teamwork thrive through diverse perspectives, both within our teams and in our interactions with our customers.\nWork/life balance that supports our employees' varying needs: We value the well-being of our employees, recognizing that a healthy work-life balance is pivotal to our collective success.\nEmbracing communities: We embrace and support the communities that nurture us. Our employees' dedication to fostering positive change is a source of immense pride for us.\nCommitment to Diversity, Inclusion and Belonging\nWe are an equal opportunity employer that does not discriminate or allow discrimination based on race, color, religion, sex, sexual orientation, gender identity, age, national origin, citizenship, disability, veteran status, or any other classification protected by federal, state, or local law.\nePlus is dedicated to fostering, cultivating, and preserving a culture that represents diversity, enables inclusion, and makes our employees feel comfortable bringing their full, unique selves to work.\nPhysical Requirements\nWhile performing this role, you will engage in both seated and occasional standing or walking activities. We provide reasonable accommodations, in accordance with relevant laws, to support success in this position.\nBy embracing our values, you will contribute to our collective mission of making a positive impact within our organization and the broader community. We understand that this job description serves as a guide and is not an employment contract.\nePlus maintains a California Consumer Privacy Act (CCPA) Privacy Notice on our Trust Center, available here: CCPA Privacy Notice.\nNotice to Recruiting Agencies: ePlus only accepts unsolicited resumes when presented directly by a candidate. Unsolicited resumes submitted to ePlus from any other source will be considered ePlus property and will not qualify for any placement or referral fees. ePlus will only pay such fees in connection with a valid written agreement between ePlus and the referring agency, and then only after providing advance written approval to the referring agency to submit resumes in connection with a particular opportunity.","datePosted":"2026-08-03T19:34:42.593Z","dateModified":"2026-08-03T19:34:42.593Z","hiringOrganization":{"@type":"Organization","name":"Eplus","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Irvine","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a0e3de7c3988e87f6d2edf5d"},"url":"https://jobsearcher.com/jobs/a0e3de7c3988e87f6d2edf5d"}}