ML systems
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
At Decart I shipped the first realtime diffusion model in the world to run on Trainium which was presented on the main stage of AWS re:Invent, a stage reserved for the president and other leaders of Amazon. This was even the first such live demo in re:Invent history! For me, this is another level of fulfillment, far beyond just improving kernels for existing open-source models.I’m looking for an exceptional AI Kernel Engineer in San Francisco to own low-level performance for realtime diffusion models: custom CUDA/Triton/NKI kernels, compiler/runtime integration, and end-to-end optimizations. If you’ve shipped real speedups on modern accelerators, DM me a link to your work, or email me at heba@decart.ai Learn more about Decart https://decart.ai/company Re:Invent Demo https://lnkd.in/ggBrfVcW