{"schemaVersion":"jobsearcher.job.v1","id":"79b3fb6281d0ad523f694657","url":"https://jobsearcher.com/jobs/79b3fb6281d0ad523f694657","canonicalUrl":"https://jobsearcher.com/jobs/79b3fb6281d0ad523f694657","title":"Performance Tools Intern","description":"About Etched\nEtched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.\n\nJob Summary\nJoin our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.\n\nDuring your internship, you may:\n\nBuild components of our performance analysis and profiling infrastructure.\n\nCollect and analyze performance data from our custom ML accelerators, including hardware counters, execution traces, and memory behavior.\n\nDevelop tooling to trace host-side runtime activity, system behavior, and accelerator execution.\n\nHelp correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads.\n\nBuild analysis and visualization tools that help engineers identify performance bottlenecks and optimize models.\n\nWork alongside hardware, compiler, firmware, and inference engineers to understand performance challenges and develop tools that improve developer productivity.\n\nRepresentative projects\n\nImplement the data collection framework for hardware performance counters on a custom PCIe-based accelerator.\n\nDevelop a user-space service for low-overhead tracing of accelerator activity.\n\nDesign and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.\n\nCreate an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.\n\nYou may be a good fit if you have\n\nStrong programming skills in C++ or Rust . Experience with Python is a plus.\n\nSolid understanding of computer architecture, including CPUs, GPUs or AI accelerators, memory hierarchies, and parallel programming.\n\nExperience or strong interest in low-level performance analysis, profiling, and performance optimization.\n\nFamiliarity with performance analysis tools such as Nsight, VTune, Xprof, Perfetto , or similar tools is a plus.\n\nExperience or strong interest in operating systems, compilers, firmware, drivers, or other low-level systems software.\n\nPassion for understanding how complex systems behave under real workloads and building tools that help other engineers optimize performance.\n\nStrong problem-solving skills and curiosity to learn quickly in a fast-paced engineering environment.\n\nStrong candidates may also have experience with (Nice-to-have qualifications)\n\nDirect experience developing performance analysis or debugging tools.\n\nExperience with ML accelerator architectures (GPUs, TPUs, etc.).\n\nExperience with kernel-mode driver development (Linux or Windows).\n\nHow we’re different\nEtched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.\n\nWe are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.\n\n#J-18808-Ljbffr","company":"Consensus","rawCompany":"consensus","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-02T03:50:44.275Z","occupations":[{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"334111","title":"Electronic Computer Manufacturing","slug":"electronic-computer-manufacturing"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Performance Tools Intern","description":"About Etched\nEtched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.\n\nJob Summary\nJoin our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.\n\nDuring your internship, you may:\n\nBuild components of our performance analysis and profiling infrastructure.\n\nCollect and analyze performance data from our custom ML accelerators, including hardware counters, execution traces, and memory behavior.\n\nDevelop tooling to trace host-side runtime activity, system behavior, and accelerator execution.\n\nHelp correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads.\n\nBuild analysis and visualization tools that help engineers identify performance bottlenecks and optimize models.\n\nWork alongside hardware, compiler, firmware, and inference engineers to understand performance challenges and develop tools that improve developer productivity.\n\nRepresentative projects\n\nImplement the data collection framework for hardware performance counters on a custom PCIe-based accelerator.\n\nDevelop a user-space service for low-overhead tracing of accelerator activity.\n\nDesign and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.\n\nCreate an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.\n\nYou may be a good fit if you have\n\nStrong programming skills in C++ or Rust . Experience with Python is a plus.\n\nSolid understanding of computer architecture, including CPUs, GPUs or AI accelerators, memory hierarchies, and parallel programming.\n\nExperience or strong interest in low-level performance analysis, profiling, and performance optimization.\n\nFamiliarity with performance analysis tools such as Nsight, VTune, Xprof, Perfetto , or similar tools is a plus.\n\nExperience or strong interest in operating systems, compilers, firmware, drivers, or other low-level systems software.\n\nPassion for understanding how complex systems behave under real workloads and building tools that help other engineers optimize performance.\n\nStrong problem-solving skills and curiosity to learn quickly in a fast-paced engineering environment.\n\nStrong candidates may also have experience with (Nice-to-have qualifications)\n\nDirect experience developing performance analysis or debugging tools.\n\nExperience with ML accelerator architectures (GPUs, TPUs, etc.).\n\nExperience with kernel-mode driver development (Linux or Windows).\n\nHow we’re different\nEtched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.\n\nWe are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.\n\n#J-18808-Ljbffr","datePosted":"2026-09-02T03:50:44.275Z","dateModified":"2026-09-02T03:50:44.275Z","hiringOrganization":{"@type":"Organization","name":"Consensus","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"79b3fb6281d0ad523f694657"},"url":"https://jobsearcher.com/jobs/79b3fb6281d0ad523f694657"}}