{"schemaVersion":"jobsearcher.job.v1","id":"ce5413aaefeca7fdf543956f","url":"https://jobsearcher.com/jobs/ce5413aaefeca7fdf543956f","canonicalUrl":"https://jobsearcher.com/jobs/ce5413aaefeca7fdf543956f","title":"Founding Inference Engineer","description":"About us\n\nGeneral Compute is the neocloud for alternative chips.\n\nInference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.\n\nWe closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.\n\nAbout the Role\n\nGetting a model correct and fast on our silicon is only half the problem — the other half is serving it. You'll build and own the inference layer that sits between a bought-up model and a live customer request: request scheduling, batching, KV-cache management, autoscaling across our ASIC fleet, and the failure modes that only show up at real traffic and real scale.\n\nThis is a founding role on a small team, which means the scope is wide and the ownership is real: there's no separate SRE org to hand reliability to and no platform team to hand infra to. You'll design the serving architecture, then be the person paged when it breaks. The bet is that a serving stack built specifically for our hardware — not adapted from a GPU-first framework — is a durable edge, and you're the person who proves that out in production.\n\nWhat You'll Do:\n\nOwn the inference serving stack end-to-end. Design and build the system that takes a bring-up-verified model and serves it in production: request routing, batching, scheduling, and autoscaling across our ASIC fleet.\n\nPush cost-per-token down. Continuously tune batching strategy, KV-cache handling, and hardware utilization to widen the throughput advantage over GPU-based serving.\n\nBuild for reliability from day one. Put in place the monitoring, alerting, and failover that make a fast-moving inference stack trustworthy under real customer load — and be the one who responds when it isn't.\n\nWork at the boundary with the compiler and bring-up team. Define the interface between \"a model is correct and compiled\" and \"a model is live and fast,\" and push issues back to the right side of that line.\n\nShape the roadmap, not just the backlog. As a founding engineer, you'll help decide what we build next in serving — multi-tenant isolation, speculative decoding, new scheduling strategies — not just execute a spec someone else wrote.\n\nSet the technical bar for the team you're helping build. Early architecture and code-quality decisions you make here will shape how the serving team operates as it grows.\n\nWhat We Need From You:\n\n5+ years building and operating production systems at the infrastructure layer, ideally including a high-throughput or low-latency serving system.\n\nDirect experience with LLM inference serving — request batching, KV-cache management, continuous batching, or similar — in a production environment, not just research code.\n\nComfortable owning reliability: you've been on call for a system that mattered, and you design for failure rather than reacting to it after the fact.\n\nStrong systems fundamentals — concurrency, networking, scheduling — deep enough to reason about performance at the hardware level, not just the application level.\n\nSelf-directed and comfortable with ambiguity. This is a founding role: there's no existing playbook to follow, and you'll help write it.\n\nNice-to-Haves:\n\nExperience serving models on non-NVIDIA accelerators (TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar).\n\nFamiliarity with serving frameworks such as vLLM, TGI, TensorRT-LLM, or SGLang, and an opinion on where they fall short.\n\nExperience running infrastructure at a small company or in a founding/early-engineer capacity before.\n\nExposure to capacity planning or fleet management for specialized hardware.","company":"General Compute","rawCompany":"general compute","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-15T09:54:08.645Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Founding Inference Engineer","description":"About us\n\nGeneral Compute is the neocloud for alternative chips.\n\nInference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.\n\nWe closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.\n\nAbout the Role\n\nGetting a model correct and fast on our silicon is only half the problem — the other half is serving it. You'll build and own the inference layer that sits between a bought-up model and a live customer request: request scheduling, batching, KV-cache management, autoscaling across our ASIC fleet, and the failure modes that only show up at real traffic and real scale.\n\nThis is a founding role on a small team, which means the scope is wide and the ownership is real: there's no separate SRE org to hand reliability to and no platform team to hand infra to. You'll design the serving architecture, then be the person paged when it breaks. The bet is that a serving stack built specifically for our hardware — not adapted from a GPU-first framework — is a durable edge, and you're the person who proves that out in production.\n\nWhat You'll Do:\n\nOwn the inference serving stack end-to-end. Design and build the system that takes a bring-up-verified model and serves it in production: request routing, batching, scheduling, and autoscaling across our ASIC fleet.\n\nPush cost-per-token down. Continuously tune batching strategy, KV-cache handling, and hardware utilization to widen the throughput advantage over GPU-based serving.\n\nBuild for reliability from day one. Put in place the monitoring, alerting, and failover that make a fast-moving inference stack trustworthy under real customer load — and be the one who responds when it isn't.\n\nWork at the boundary with the compiler and bring-up team. Define the interface between \"a model is correct and compiled\" and \"a model is live and fast,\" and push issues back to the right side of that line.\n\nShape the roadmap, not just the backlog. As a founding engineer, you'll help decide what we build next in serving — multi-tenant isolation, speculative decoding, new scheduling strategies — not just execute a spec someone else wrote.\n\nSet the technical bar for the team you're helping build. Early architecture and code-quality decisions you make here will shape how the serving team operates as it grows.\n\nWhat We Need From You:\n\n5+ years building and operating production systems at the infrastructure layer, ideally including a high-throughput or low-latency serving system.\n\nDirect experience with LLM inference serving — request batching, KV-cache management, continuous batching, or similar — in a production environment, not just research code.\n\nComfortable owning reliability: you've been on call for a system that mattered, and you design for failure rather than reacting to it after the fact.\n\nStrong systems fundamentals — concurrency, networking, scheduling — deep enough to reason about performance at the hardware level, not just the application level.\n\nSelf-directed and comfortable with ambiguity. This is a founding role: there's no existing playbook to follow, and you'll help write it.\n\nNice-to-Haves:\n\nExperience serving models on non-NVIDIA accelerators (TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar).\n\nFamiliarity with serving frameworks such as vLLM, TGI, TensorRT-LLM, or SGLang, and an opinion on where they fall short.\n\nExperience running infrastructure at a small company or in a founding/early-engineer capacity before.\n\nExposure to capacity planning or fleet management for specialized hardware.","datePosted":"2026-09-15T09:54:08.645Z","dateModified":"2026-09-15T09:54:08.645Z","hiringOrganization":{"@type":"Organization","name":"General Compute","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"ce5413aaefeca7fdf543956f"},"url":"https://jobsearcher.com/jobs/ce5413aaefeca7fdf543956f"}}