{"schemaVersion":"jobsearcher.job.v1","id":"0cac3aa121b014b7af702fcb","url":"https://jobsearcher.com/jobs/0cac3aa121b014b7af702fcb","canonicalUrl":"https://jobsearcher.com/jobs/0cac3aa121b014b7af702fcb","title":"Lead RTL Design Engineer","description":"Efficient is developing the world's most energy-efficient general-purpose computer processor. Efficient's patented technology uses 100x less energy than state of the art commercially available ultra-low-power processors and is programmable using standard high-level programming languages and AI/ML frameworks. This level of efficiency makes perpetual, pervasive intelligence possible: run AI/ML continuously on a AA battery for 5-10 years. Our platform's unprecedented level of efficiency enables IoT devices to intelligently capture and curate first-party data to drive the next major computing revolution\nWe are looking for a Lead RTL Design Engineer to own microarchitecture definition and RTL implementation across the dataflow execution fabric, memory subsystem, on-chip interconnect/NoC, low-power logic, and standard peripheral IP (RiscV, NVM, I2S, I2C) integration. You will work from architecture spec through synthesis-ready RTL, collaborating with architects, microarchitects, DV leads, physical design, and firmware teams to tape out an industry leading power-efficient SoC.\nThis is a unique opportunity to be a part of a newly formed HW engineering org and have an influence on our products and processes as we move from the initial stages of product development to market release and scaled volume production. Join our team and help us shape the future of computing at the edge and beyond!\nKey Responsibilities\nMicroarchitecture definition (core focus): Own the design and definition of processor and compute-unit microarchitecture, including dataflow pipelines, execution units, and interfaces. Set performance, power, and area targets, and guide the team toward achieving them.\nOn-chip interconnects and system integration: Define and drive the design of on-chip networks and data movement across the fabric, balancing performance, scalability, and implementation constraints in collaboration with physical design.\nMemory subsystem & system architecture: Define the interface to the memory subsystem, including data movement, ordering, and synchronization behavior, ensuring a clean and scalable model for software and future system expansion.\nReconfiguration and execution model: Lead the architecture of configuration, scheduling, and execution of workloads on the fabric, including multi-kernel support and interaction with host systems.\nPower management and low-power design: Drive power architecture across the design, including clocking, reset, power domains, and low-power strategies to meet aggressive energy and efficiency goals.\nHW/SW co-design: Collaborate closely with compiler and software teams to define the hardware execution model, ensuring efficient mapping of workloads onto the architecture.\nSpecifications, Documentations and Reviews: Author and own uArch specification documents for assigned blocks; drive design reviews with architecture, compiler, DV, and physical design stakeholders.\nMentoring and process improvement: Mentor senior and junior RTL engineers; review RTL, flag microarchitecture risks, and enforce coding style and lint-clean standards across the team.\nDriving PPA Metrics: Participate in PPA analysis loops: synthesize blocks regularly, review area/timing/power reports, and make data-driven tradeoffs against performance and feature requirements.\nDV Collaboration: Collaborate with DV leads to define/review verification plans; provide directed test scenarios for graph execution corner cases, back-pressure conditions, and power state transitions.\nSilicon Bring-up: Support silicon bring-up: contribute scan/ATPG guidelines, review DFT insertion, and provide RTL-level debug assistance during lab validation.\nRequired Qualifications & Experience\n8+ years of RTL design experience with tape-out ownership of dataflow based design, on chip networks, memory subsystems or peripheral integration on a processor or accelerator SoC.\nDeep proficiency in SystemVerilog for RTL — synthesis-clean, lint-clean, timing-aware; able to design complex state machines, arbiters, token flow controllers, and datapath logic from scratch.\nSolid understanding of parallel execution models: dataflow, SIMD, or systolic array architectures; familiarity with the hardware challenges of token-based firing-rule evaluation and producer-consumer synchronization.\nHands-on experience with on-chip memory design: SRAM wrappers, scratchpad/TCM, banking, and memory-mapped register interfaces.\nExperience with low-power RTL techniques: UPF-driven flows, clock gating, power domains, retention registers, and AON wakeup logic.\nFamiliarity with at least one standard on-chip bus protocol (AXI, AHB, APB, TileLink, or NoC equivalent) at the RTL implementation level.\nExperience taking RTL through synthesis and timing closure; ability to read and act on SDC constraints, STA reports, and synthesis QoR summaries.\nStrong written communication skills; able to produce uArch specs and design review material independently.\nExperience with memory compiler toolchains\nDesired Qualifications & Experience Requirements\nPrior RTL ownership of a dataflow engine, neural processing unit (NPU), or streaming DSP architecture with explicit producer-consumer token management.\nExperience collaborating with compiler or graph-optimization teams to co-design hardware execution models and graph IR representations.\nFamiliarity with NVM controller RTL (MRAM, RRAM) including ECC, program/erase sequencing, and model weight storage use cases.\nExperience with IoT-class power budgets (sub-10 mW active, sub-100 µW standby) and the RTL design choices they necessitate.\nFamiliarity with functional safety standards (ISO 26262, IEC 61508) as applied to execution fabric error detection and power domain isolation.\nExposure to AI framework graph formats (ONNX, TFLite) and understanding of how graph compilation maps to hardware execution primitives.\nTape-out credits on an edge-AI, IoT, or wearable SoC at 12nm or below.\nExperience with formal verification of flow-control logic, deadlock freedom, or bus protocol compliance.\nWe offer a competitive salary for this role, generally ranging from $160,000 to $250,000, along with meaningful equity and comprehensive benefits. The final compensation package will be based on your experience and location, with some flexibility to ensure we align with the right candidate.\nWhy Join Efficient?\nEfficient offers a competitive compensation and benefits package, including 401K match, company-paid benefits, equity program, paid parental leave, and flexibility. We are committed to personal and professional development and strive to grow together as people and as a company.","company":"Efficient Computer","rawCompany":"efficient computer","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-23T14:45:19.353Z","occupations":[{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"17-2072.00","title":"Electronics Engineers, Except Computer","slug":"electronics-engineers-except-computer"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"334413","title":"Semiconductor and Related Device Manufacturing","slug":"semiconductor-and-related-device-manufacturing"},{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Lead RTL Design Engineer","description":"Efficient is developing the world's most energy-efficient general-purpose computer processor. Efficient's patented technology uses 100x less energy than state of the art commercially available ultra-low-power processors and is programmable using standard high-level programming languages and AI/ML frameworks. This level of efficiency makes perpetual, pervasive intelligence possible: run AI/ML continuously on a AA battery for 5-10 years. Our platform's unprecedented level of efficiency enables IoT devices to intelligently capture and curate first-party data to drive the next major computing revolution\nWe are looking for a Lead RTL Design Engineer to own microarchitecture definition and RTL implementation across the dataflow execution fabric, memory subsystem, on-chip interconnect/NoC, low-power logic, and standard peripheral IP (RiscV, NVM, I2S, I2C) integration. You will work from architecture spec through synthesis-ready RTL, collaborating with architects, microarchitects, DV leads, physical design, and firmware teams to tape out an industry leading power-efficient SoC.\nThis is a unique opportunity to be a part of a newly formed HW engineering org and have an influence on our products and processes as we move from the initial stages of product development to market release and scaled volume production. Join our team and help us shape the future of computing at the edge and beyond!\nKey Responsibilities\nMicroarchitecture definition (core focus): Own the design and definition of processor and compute-unit microarchitecture, including dataflow pipelines, execution units, and interfaces. Set performance, power, and area targets, and guide the team toward achieving them.\nOn-chip interconnects and system integration: Define and drive the design of on-chip networks and data movement across the fabric, balancing performance, scalability, and implementation constraints in collaboration with physical design.\nMemory subsystem & system architecture: Define the interface to the memory subsystem, including data movement, ordering, and synchronization behavior, ensuring a clean and scalable model for software and future system expansion.\nReconfiguration and execution model: Lead the architecture of configuration, scheduling, and execution of workloads on the fabric, including multi-kernel support and interaction with host systems.\nPower management and low-power design: Drive power architecture across the design, including clocking, reset, power domains, and low-power strategies to meet aggressive energy and efficiency goals.\nHW/SW co-design: Collaborate closely with compiler and software teams to define the hardware execution model, ensuring efficient mapping of workloads onto the architecture.\nSpecifications, Documentations and Reviews: Author and own uArch specification documents for assigned blocks; drive design reviews with architecture, compiler, DV, and physical design stakeholders.\nMentoring and process improvement: Mentor senior and junior RTL engineers; review RTL, flag microarchitecture risks, and enforce coding style and lint-clean standards across the team.\nDriving PPA Metrics: Participate in PPA analysis loops: synthesize blocks regularly, review area/timing/power reports, and make data-driven tradeoffs against performance and feature requirements.\nDV Collaboration: Collaborate with DV leads to define/review verification plans; provide directed test scenarios for graph execution corner cases, back-pressure conditions, and power state transitions.\nSilicon Bring-up: Support silicon bring-up: contribute scan/ATPG guidelines, review DFT insertion, and provide RTL-level debug assistance during lab validation.\nRequired Qualifications & Experience\n8+ years of RTL design experience with tape-out ownership of dataflow based design, on chip networks, memory subsystems or peripheral integration on a processor or accelerator SoC.\nDeep proficiency in SystemVerilog for RTL — synthesis-clean, lint-clean, timing-aware; able to design complex state machines, arbiters, token flow controllers, and datapath logic from scratch.\nSolid understanding of parallel execution models: dataflow, SIMD, or systolic array architectures; familiarity with the hardware challenges of token-based firing-rule evaluation and producer-consumer synchronization.\nHands-on experience with on-chip memory design: SRAM wrappers, scratchpad/TCM, banking, and memory-mapped register interfaces.\nExperience with low-power RTL techniques: UPF-driven flows, clock gating, power domains, retention registers, and AON wakeup logic.\nFamiliarity with at least one standard on-chip bus protocol (AXI, AHB, APB, TileLink, or NoC equivalent) at the RTL implementation level.\nExperience taking RTL through synthesis and timing closure; ability to read and act on SDC constraints, STA reports, and synthesis QoR summaries.\nStrong written communication skills; able to produce uArch specs and design review material independently.\nExperience with memory compiler toolchains\nDesired Qualifications & Experience Requirements\nPrior RTL ownership of a dataflow engine, neural processing unit (NPU), or streaming DSP architecture with explicit producer-consumer token management.\nExperience collaborating with compiler or graph-optimization teams to co-design hardware execution models and graph IR representations.\nFamiliarity with NVM controller RTL (MRAM, RRAM) including ECC, program/erase sequencing, and model weight storage use cases.\nExperience with IoT-class power budgets (sub-10 mW active, sub-100 µW standby) and the RTL design choices they necessitate.\nFamiliarity with functional safety standards (ISO 26262, IEC 61508) as applied to execution fabric error detection and power domain isolation.\nExposure to AI framework graph formats (ONNX, TFLite) and understanding of how graph compilation maps to hardware execution primitives.\nTape-out credits on an edge-AI, IoT, or wearable SoC at 12nm or below.\nExperience with formal verification of flow-control logic, deadlock freedom, or bus protocol compliance.\nWe offer a competitive salary for this role, generally ranging from $160,000 to $250,000, along with meaningful equity and comprehensive benefits. The final compensation package will be based on your experience and location, with some flexibility to ensure we align with the right candidate.\nWhy Join Efficient?\nEfficient offers a competitive compensation and benefits package, including 401K match, company-paid benefits, equity program, paid parental leave, and flexibility. We are committed to personal and professional development and strive to grow together as people and as a company.","datePosted":"2026-07-23T14:45:19.353Z","dateModified":"2026-07-23T14:45:19.353Z","hiringOrganization":{"@type":"Organization","name":"Efficient Computer","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0cac3aa121b014b7af702fcb"},"url":"https://jobsearcher.com/jobs/0cac3aa121b014b7af702fcb"}}