JOBSEARCHER

GPU Solution Architect

Why Zensar?We’re a bunch of hardworking, fun-loving, people-oriented technology enthusiasts. We love what we do, and we’re passionate about helping our clients thrive in an increasingly complex digital world. Zensar is an organization focused on building relationships with our clients and with each other—and happiness is at the core of everything we do. In fact, we’re so into happiness that we’ve created a Global Happiness Council, and we send out a Happiness Survey to our employees each year. We’ve learned that employee happiness requires more than a competitive paycheck, and our employee value proposition—grow, own, achieve, learn (GOAL)—lays out the core opportunities we seek to foster for every employee. Teamwork and collaboration are critical to Zensar’s mission and success, and our teams work on a diverse and challenging mix of technologies across a broad industry spectrum. These industries include banking and financial services, high-tech and manufacturing, healthcare, insurance, retail, and consumer services. Our employees enjoy flexible work arrangements and a competitive benefits package, including medical, dental, vision, 401(k), among other benefits. If you are looking for a place to have an immediate impact, to grow and contribute, where we work hard, play hard, and support each other, consider joining team Zensar!Zensar is looking for a GPU Solution Architect in the United States (Hybrid). This position is open for Full Time with excellent benefits and professional growth opportunities.About the RoleSolve hard Day 2 operations problems at scale. Work alongside partner engineers to find the cause, prototype an approach, validate it under representative load, and leave behind a practice their team can operate.Make new technology Day 2 ready. Help partners prepare the operating model for new platforms, capacity, services, and use cases before customers depend on them,and help drive adoption in live environments without degrading service.Improve reliability, performance, and economics together. Use measures such as incident frequency, recovery time, utilization, and cost per token to show where the cloud is losing performance or margin - and whether the fix worked.Raise each partner's Day 2 maturity. Identify and help close the gaps that matter across people, process, tooling, telemetry, security, and incident response.Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows that other client partners can integrate into their standard operating model.Create the feedback loop. Spot patterns across partners early and bring clear field evidence to account teams, support, product, and engineering so repeated problems are fixed at the right level.Experience:8+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.Experience building, operating, or improving distributed infrastructure under real production load - not only designing or deploying it.Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure. Relevant technologies may include DCGM, BMC/Redfish, and firmware and driver lifecycle; InfiniBand or high-speed Ethernet, NCCL, and UFM; or high-performance storage such as Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.A detailed evidence-led approach to troubleshooting across system boundaries, paired with the judgment to make difficult technical findings clear.The ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority or taking ownership away from the operator.Strong communication, prioritization, and time-management skills across multiple partner engagements.Education:Bachelor’s / Master’s Degree – Information TechnologyZensar believes that diversity of backgrounds, thought, experience, and expertise fosters the robust exchange of ideas that enables the highest quality collaboration and work product. Zensar is an equal opportunity employer. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by federal, state, or local law. Zensar is committed to providing veteran employment opportunities to our service men and women. Zensar is committed to providing equal employment opportunities for people with disabilities or religious observances, including reasonable accommodation when needed. Accommodation made to facilitate the recruiting process are not a guarantee of future or continued accommodation once hired.All applicants must be legally authorized to work with Zensar.Zensar values your privacy. We’ll use your data in accordance with our privacy statement located at: https://zensar.com/privacy-notice