{"schemaVersion":"jobsearcher.job.v1","id":"6238e7a0ecfe58a148be241c","url":"https://jobsearcher.com/jobs/6238e7a0ecfe58a148be241c","canonicalUrl":"https://jobsearcher.com/jobs/6238e7a0ecfe58a148be241c","title":"Founding Platform Engineer","description":"About us\n\nGeneral Compute is the neocloud for alternative chips.\n\nInference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.\n\nWe closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.\n\nAbout the role\n\nYou will build the inference cloud itself — the control plane, API, and serving layer that turn racks into a sellable product. There's no existing platform team to inherit or manage, no legacy system to work around, and no established playbook to follow — just the platform itself to build, with reliability treated as core infrastructure from day one rather than something bolted on after the first outage.The technical problem is also genuinely unsolved elsewhere. The fleet is heterogeneous by design — GPUs for prefill, multiple ASIC vendors for decode — so there's no single-vendor playbook to lean on; you'll be defining how a mixed-hardware inference cloud gets scheduled, routed, and served reliably, in close partnership with the teams standing up the physical fleet.\n\nWhat you'll do:\n\nBuild/own the control plane — routing, model placement, scheduling across a mixed ASIC/GPU pool\n\nBuild the API and serving layer exposing rack capacity as a sellable product\n\nBuild in reliability and observability from day one\n\nScale the platform ahead of the demand curve\n\nPartner closely with data center deployment and model bring-up teams\n\nBe a founding technical voice on platform architecture\n\nWhat we need from you:\n\nStrong systems engineering background on cloud control planes/serving infra at scale\n\nComfort being a high-impact IC rather than a manager\n\nTrack record building reliability from scratch\n\nComfort with hardware heterogeneity/ambiguity\n\nGenuine interest in being an early hire at a ~6–7 person company.\n\nNice-to-haves:\n\nLLM-serving infra experience (vLLM, TGI, Ray Serve, etc.)\n\nExperience running non-NVIDIA accelerators (TPUs/ASICs) in production.","company":"General Compute","rawCompany":"general compute","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-15T09:53:52.913Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Founding Platform Engineer","description":"About us\n\nGeneral Compute is the neocloud for alternative chips.\n\nInference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.\n\nWe closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.\n\nAbout the role\n\nYou will build the inference cloud itself — the control plane, API, and serving layer that turn racks into a sellable product. There's no existing platform team to inherit or manage, no legacy system to work around, and no established playbook to follow — just the platform itself to build, with reliability treated as core infrastructure from day one rather than something bolted on after the first outage.The technical problem is also genuinely unsolved elsewhere. The fleet is heterogeneous by design — GPUs for prefill, multiple ASIC vendors for decode — so there's no single-vendor playbook to lean on; you'll be defining how a mixed-hardware inference cloud gets scheduled, routed, and served reliably, in close partnership with the teams standing up the physical fleet.\n\nWhat you'll do:\n\nBuild/own the control plane — routing, model placement, scheduling across a mixed ASIC/GPU pool\n\nBuild the API and serving layer exposing rack capacity as a sellable product\n\nBuild in reliability and observability from day one\n\nScale the platform ahead of the demand curve\n\nPartner closely with data center deployment and model bring-up teams\n\nBe a founding technical voice on platform architecture\n\nWhat we need from you:\n\nStrong systems engineering background on cloud control planes/serving infra at scale\n\nComfort being a high-impact IC rather than a manager\n\nTrack record building reliability from scratch\n\nComfort with hardware heterogeneity/ambiguity\n\nGenuine interest in being an early hire at a ~6–7 person company.\n\nNice-to-haves:\n\nLLM-serving infra experience (vLLM, TGI, Ray Serve, etc.)\n\nExperience running non-NVIDIA accelerators (TPUs/ASICs) in production.","datePosted":"2026-09-15T09:53:52.913Z","dateModified":"2026-09-15T09:53:52.913Z","hiringOrganization":{"@type":"Organization","name":"General Compute","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6238e7a0ecfe58a148be241c"},"url":"https://jobsearcher.com/jobs/6238e7a0ecfe58a148be241c"}}