{"schemaVersion":"jobsearcher.job.v1","id":"a1dfedfcedbebc4d0c5164a5","url":"https://jobsearcher.com/jobs/a1dfedfcedbebc4d0c5164a5","canonicalUrl":"https://jobsearcher.com/jobs/a1dfedfcedbebc4d0c5164a5","title":"GPU Systems Engineer","description":"Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.\nTower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.\nEngineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.\nAt Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.\nAt Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.\nSummary: Trading and research at the firm run around the clock and across the globe, and they run on infrastructure this team designs, builds, and operates.\nAs part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes.\nThe role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.\nResponsibilities Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.\nTrack down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.\nPartner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.\nBuild the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.\nOwn infrastructure projects end to end, from scope and design through implementation and long-term support.\nQualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.\nQualifications 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.\nDeep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.\nHands‑on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.\nWorking experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.\nSolid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.\nFamiliarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.\nComfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.\nClear communication. You will work daily with researchers, engineers, and vendors.\nNice to Have Experience with the rest of the NVIDIA stack, such as NCCL and NVLink.\nAnticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.\nTower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.\nAt Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive – without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.\nBenefits Generous paid time off policies\nSavings plans and other financial wellness tools available in each region\nHybrid working opportunities\nFree breakfast, lunch, and snacks daily\nIn-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)\nCompany-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)\nVolunteer opportunities and charitable giving\nSocial events, happy hours, treats, and celebrations throughout the year\nWorkshops and continuous learning opportunities\nAt Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work – together.\nTower Research Capital is an equal opportunity employer.\n\n#J-18808-Ljbffr","company":"Socket","rawCompany":"socket","city":"New York","state":"NY","isRemote":false,"isActive":true,"createdAt":"2026-08-18T03:35:00.816Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"GPU Systems Engineer","description":"Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.\nTower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.\nEngineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.\nAt Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.\nAt Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.\nSummary: Trading and research at the firm run around the clock and across the globe, and they run on infrastructure this team designs, builds, and operates.\nAs part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes.\nThe role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.\nResponsibilities Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.\nTrack down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.\nPartner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.\nBuild the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.\nOwn infrastructure projects end to end, from scope and design through implementation and long-term support.\nQualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.\nQualifications 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.\nDeep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.\nHands‑on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.\nWorking experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.\nSolid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.\nFamiliarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.\nComfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.\nClear communication. You will work daily with researchers, engineers, and vendors.\nNice to Have Experience with the rest of the NVIDIA stack, such as NCCL and NVLink.\nAnticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.\nTower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.\nAt Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive – without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.\nBenefits Generous paid time off policies\nSavings plans and other financial wellness tools available in each region\nHybrid working opportunities\nFree breakfast, lunch, and snacks daily\nIn-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)\nCompany-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)\nVolunteer opportunities and charitable giving\nSocial events, happy hours, treats, and celebrations throughout the year\nWorkshops and continuous learning opportunities\nAt Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work – together.\nTower Research Capital is an equal opportunity employer.\n\n#J-18808-Ljbffr","datePosted":"2026-08-18T03:35:00.816Z","dateModified":"2026-08-18T03:35:00.816Z","hiringOrganization":{"@type":"Organization","name":"Socket","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"New York","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a1dfedfcedbebc4d0c5164a5"},"url":"https://jobsearcher.com/jobs/a1dfedfcedbebc4d0c5164a5"}}