{"schemaVersion":"jobsearcher.job.v1","id":"161d26253acb1920e8c8524c","url":"https://jobsearcher.com/jobs/161d26253acb1920e8c8524c","canonicalUrl":"https://jobsearcher.com/jobs/161d26253acb1920e8c8524c","title":"Solutions Engineer","description":"About the Role We’re looking for a Forward Deployment Engineer (FDE) to work directly with customers and partners to design, deploy, and validate Inference dedicated endpoint & Model-as-a-Service products on GMI’s global infrastructure.\nThis is a high-impact, hybrid engineering role that sits at the intersection of platform engineering, applied ML, and customer success. You’ll be embedded with customers during early-stage deployments—turning research ideas, datasets, and business requirements into working, performant systems on real GPU clusters.\nIf you enjoy being close to users, debugging real systems, and shipping results fast (not just writing docs), this role is for you.\nWhat You’ll Do Own customer POCs end-to-end\nDeploy and optimize LLM and multi-modal inference workflows on GMI clusters\nTranslate customer requirements into concrete system designs and experiments\nForward-deploy with customers\nWork hands-on with research teams, startups, and enterprise customers\nDebug performance, stability, and correctness issues in real environments\nInference deployment\nStand up and tune inference stacks (e.g. vLLM / SGLang / Ray Serve–style architectures)\nOptimize latency, throughput, GPU utilization, and cost efficiency\nModel-as-a-Service enablement\nHelp customers test, evaluate, and adopt the most frontier LLM and multi-modal models through GMI's unified API\nGuide model selection, API integration, and migration across providers; shorten the \"idea → production\" cycle\nValidate correctness, compatibility, and performance across the MaaS model catalog\nPerformance & reliability\nDiagnose GPU, networking, and distributed system bottlenecks\nRun benchmarks, profiling, and stress tests on multi-GPU / multi-node setups\nFeedback loop to product\nFeed real-world customer learnings back into GMI's platform, SDKs, and APIs\nHelp shape reference architectures, cookbooks, and best practices\nWhat We’re Looking For Core Requirements\nProficiency in at least one programming language (Python and Golang preferred)\nSolid understanding of software systems and distributed systems\nHands-on experience with ML inference or serving systems\nComfort working directly with customers and ambiguous requirements\nAbility to debug end-to-end systems (code, infra, networking, performance)\nNice to Have\nExperience with:\nLLM inference frameworks (vLLM, SGLang, Ray Serve, Triton, etc.)\nGlobal, distributed systems\nHands-on experience developing and maintaining production services on Kubernetes\nGPU performance profiling, optimization, and inference benchmarking\nPrior experience as:\nForward Deployed Engineer\nSolutions Engineer\nML Platform Engineer\nApplied Research Engineer\nWhat Makes This Role Special You’re close to real users and real GPUs — not abstract roadmaps\nYou’ll work on cutting-edge inference and frontier models , not toy demos\nYou’ll influence product direction through direct customer feedback\nFast iteration, high ownership, and visible impact\nWho Thrives Here - Engineers who like shipping over theorizing\n- People who enjoy being the “last mile” problem solver\n- Builders who want exposure to buth deep systems and applied ML\n- Those excited by early-stage POCs that turn into real production systems\n\n#J-18808-Ljbffr","company":"Socket","rawCompany":"socket","city":"Mountain View","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-06T03:36:58.849Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"41-9031.00","title":"Sales Engineers","slug":"sales-engineers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Solutions Engineer","description":"About the Role We’re looking for a Forward Deployment Engineer (FDE) to work directly with customers and partners to design, deploy, and validate Inference dedicated endpoint & Model-as-a-Service products on GMI’s global infrastructure.\nThis is a high-impact, hybrid engineering role that sits at the intersection of platform engineering, applied ML, and customer success. You’ll be embedded with customers during early-stage deployments—turning research ideas, datasets, and business requirements into working, performant systems on real GPU clusters.\nIf you enjoy being close to users, debugging real systems, and shipping results fast (not just writing docs), this role is for you.\nWhat You’ll Do Own customer POCs end-to-end\nDeploy and optimize LLM and multi-modal inference workflows on GMI clusters\nTranslate customer requirements into concrete system designs and experiments\nForward-deploy with customers\nWork hands-on with research teams, startups, and enterprise customers\nDebug performance, stability, and correctness issues in real environments\nInference deployment\nStand up and tune inference stacks (e.g. vLLM / SGLang / Ray Serve–style architectures)\nOptimize latency, throughput, GPU utilization, and cost efficiency\nModel-as-a-Service enablement\nHelp customers test, evaluate, and adopt the most frontier LLM and multi-modal models through GMI's unified API\nGuide model selection, API integration, and migration across providers; shorten the \"idea → production\" cycle\nValidate correctness, compatibility, and performance across the MaaS model catalog\nPerformance & reliability\nDiagnose GPU, networking, and distributed system bottlenecks\nRun benchmarks, profiling, and stress tests on multi-GPU / multi-node setups\nFeedback loop to product\nFeed real-world customer learnings back into GMI's platform, SDKs, and APIs\nHelp shape reference architectures, cookbooks, and best practices\nWhat We’re Looking For Core Requirements\nProficiency in at least one programming language (Python and Golang preferred)\nSolid understanding of software systems and distributed systems\nHands-on experience with ML inference or serving systems\nComfort working directly with customers and ambiguous requirements\nAbility to debug end-to-end systems (code, infra, networking, performance)\nNice to Have\nExperience with:\nLLM inference frameworks (vLLM, SGLang, Ray Serve, Triton, etc.)\nGlobal, distributed systems\nHands-on experience developing and maintaining production services on Kubernetes\nGPU performance profiling, optimization, and inference benchmarking\nPrior experience as:\nForward Deployed Engineer\nSolutions Engineer\nML Platform Engineer\nApplied Research Engineer\nWhat Makes This Role Special You’re close to real users and real GPUs — not abstract roadmaps\nYou’ll work on cutting-edge inference and frontier models , not toy demos\nYou’ll influence product direction through direct customer feedback\nFast iteration, high ownership, and visible impact\nWho Thrives Here - Engineers who like shipping over theorizing\n- People who enjoy being the “last mile” problem solver\n- Builders who want exposure to buth deep systems and applied ML\n- Those excited by early-stage POCs that turn into real production systems\n\n#J-18808-Ljbffr","datePosted":"2026-08-06T03:36:58.849Z","dateModified":"2026-08-06T03:36:58.849Z","hiringOrganization":{"@type":"Organization","name":"Socket","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Mountain View","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"161d26253acb1920e8c8524c"},"url":"https://jobsearcher.com/jobs/161d26253acb1920e8c8524c"}}