{"schemaVersion":"jobsearcher.job.v1","id":"f79de79573098f852f75ad56","url":"https://jobsearcher.com/jobs/f79de79573098f852f75ad56","canonicalUrl":"https://jobsearcher.com/jobs/f79de79573098f852f75ad56","title":"Software Engineer – Performance Profiling","description":"About Etched\nEtched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.\nJob Summary\nJoin our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.\nKey responsibilities\nLead the design and architecture of a comprehensive performance analysis suite, including data collection mechanisms, data processing pipelines, analysis engines, and user interfaces (CLI and/or GUI).\nDevelop robust methods to capture performance data directly from our custom ML accelerator hardware (e.g., hardware performance counters, execution unit status, memory access patterns) via driver interfaces or other mechanisms.\nImplement tracing for host-side API calls (runtime libraries, driver interactions) and system-level events (CPU activity, PCIe traffic, memory usage, network contention) related to our workloads.\nDesign and implement techniques to accurately correlate performance events across the host CPU, device driver, PCIe bus, multiple accelerators, and multiple hosts, ensuring precise time synchronization.\nBuild analysis modules to automatically interpret collected trace and counter data, identifying key performance limiters (e.g., compute-bound, memory bandwidth-bound, latency-bound, PCIe-bound, specific hardware bottlenecks).\nDevelop intuitive visualizations (timelines, dependency graphs, resource utilization charts, statistical summaries) to clearly communicate performance characteristics and bottlenecks to users.\nWork closely with hardware architects, firmware engineers, driver developers, compiler engineers, and ML application engineers to understand their needs, define tool requirements, and provide expert guidance on performance analysis and optimization using the tool.\nRepresentative projects\nArchitect and implement the core data collection framework for hardware performance counters on a custom PCIe-based accelerator.\nDevelop a kernel driver module or user-space service for low-overhead tracing of accelerator activity.\nDesign and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.\nCreate an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.\nYou may be a good fit if you have\nStrong proficiency in C++ or Rust\nProficiency in Python is a plus\nDeep understanding of computer architecture (CPU, GPU, accelerators), memory hierarchies (caches, DRAM), and interconnects (especially PCIe).\nProven experience in low-level performance analysis, profiling, and bottleneck identification on complex hardware systems (GPUs, CPUs, FPGAs, or custom accelerators).\nExperience with performance analysis tools (e.g., NVIDIA Nsight, AMD uProf, Intel VTune, perf, Tracy, ETW).\nExperience working close to hardware, potentially reading performance counters or interacting directly with device drivers.\n\nStrong candidates may also have experience with (Nice-to-have qualifications)\nDirect experience developing performance analysis or debugging tools.\nExperience with ML accelerator architectures (GPUs, TPUs, etc.).\nExperience with kernel-mode driver development (Linux or Windows).\nUnderstanding of compiler internals, code generation, and optimization.\nIn-depth knowledge of the PCIe protocol and analysis tools (PCIe analyzers).\nExperience with multi-chip or multi-host accelerator systems (e.g., TPU pods, or NVidia DGX clusters)\nExperience with firmware or embedded systems development.\nExperience with hardware description languages (Verilog, VHDL) or hardware verification.\n\nBenefits\nMedical, dental, and vision packages with generous premium coverage\n$500 per month credit for waiving medical benefits\nHousing subsidy of $2k per month for those living within walking distance of the office\nRelocation support for those moving to San Jose (Santana Row)\nVarious wellness benefits covering fitness, mental health, and more\nDaily lunch + dinner in our office\nUnlimited compute budget subject to ROI justification\nHow we’re different\nEtched believes in the Bitter Lesson. We are the first inference-focused frontier AI system, betting early on transformer and transformer-like architectures and on increasing model sizes. Our addressable market is the entirety of inference, unlike many of our competitors.\nWe are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.\nCompensation Range: $150K - $275K\n\nLocation\nSan Jose\nEmployment Type\nFull time\nLocation Type\nOn-site\nDepartment\nSoftware\nCompensation\n$150K – $275K • Plus Significant Equity","company":"Etched","rawCompany":"etched","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-05T12:42:57.556Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"17-2061.00","title":"Computer Hardware Engineers","slug":"computer-hardware-engineers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer – Performance Profiling","description":"About Etched\nEtched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.\nJob Summary\nJoin our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.\nKey responsibilities\nLead the design and architecture of a comprehensive performance analysis suite, including data collection mechanisms, data processing pipelines, analysis engines, and user interfaces (CLI and/or GUI).\nDevelop robust methods to capture performance data directly from our custom ML accelerator hardware (e.g., hardware performance counters, execution unit status, memory access patterns) via driver interfaces or other mechanisms.\nImplement tracing for host-side API calls (runtime libraries, driver interactions) and system-level events (CPU activity, PCIe traffic, memory usage, network contention) related to our workloads.\nDesign and implement techniques to accurately correlate performance events across the host CPU, device driver, PCIe bus, multiple accelerators, and multiple hosts, ensuring precise time synchronization.\nBuild analysis modules to automatically interpret collected trace and counter data, identifying key performance limiters (e.g., compute-bound, memory bandwidth-bound, latency-bound, PCIe-bound, specific hardware bottlenecks).\nDevelop intuitive visualizations (timelines, dependency graphs, resource utilization charts, statistical summaries) to clearly communicate performance characteristics and bottlenecks to users.\nWork closely with hardware architects, firmware engineers, driver developers, compiler engineers, and ML application engineers to understand their needs, define tool requirements, and provide expert guidance on performance analysis and optimization using the tool.\nRepresentative projects\nArchitect and implement the core data collection framework for hardware performance counters on a custom PCIe-based accelerator.\nDevelop a kernel driver module or user-space service for low-overhead tracing of accelerator activity.\nDesign and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.\nCreate an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.\nYou may be a good fit if you have\nStrong proficiency in C++ or Rust\nProficiency in Python is a plus\nDeep understanding of computer architecture (CPU, GPU, accelerators), memory hierarchies (caches, DRAM), and interconnects (especially PCIe).\nProven experience in low-level performance analysis, profiling, and bottleneck identification on complex hardware systems (GPUs, CPUs, FPGAs, or custom accelerators).\nExperience with performance analysis tools (e.g., NVIDIA Nsight, AMD uProf, Intel VTune, perf, Tracy, ETW).\nExperience working close to hardware, potentially reading performance counters or interacting directly with device drivers.\n\nStrong candidates may also have experience with (Nice-to-have qualifications)\nDirect experience developing performance analysis or debugging tools.\nExperience with ML accelerator architectures (GPUs, TPUs, etc.).\nExperience with kernel-mode driver development (Linux or Windows).\nUnderstanding of compiler internals, code generation, and optimization.\nIn-depth knowledge of the PCIe protocol and analysis tools (PCIe analyzers).\nExperience with multi-chip or multi-host accelerator systems (e.g., TPU pods, or NVidia DGX clusters)\nExperience with firmware or embedded systems development.\nExperience with hardware description languages (Verilog, VHDL) or hardware verification.\n\nBenefits\nMedical, dental, and vision packages with generous premium coverage\n$500 per month credit for waiving medical benefits\nHousing subsidy of $2k per month for those living within walking distance of the office\nRelocation support for those moving to San Jose (Santana Row)\nVarious wellness benefits covering fitness, mental health, and more\nDaily lunch + dinner in our office\nUnlimited compute budget subject to ROI justification\nHow we’re different\nEtched believes in the Bitter Lesson. We are the first inference-focused frontier AI system, betting early on transformer and transformer-like architectures and on increasing model sizes. Our addressable market is the entirety of inference, unlike many of our competitors.\nWe are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.\nCompensation Range: $150K - $275K\n\nLocation\nSan Jose\nEmployment Type\nFull time\nLocation Type\nOn-site\nDepartment\nSoftware\nCompensation\n$150K – $275K • Plus Significant Equity","datePosted":"2026-08-05T12:42:57.556Z","dateModified":"2026-08-05T12:42:57.556Z","hiringOrganization":{"@type":"Organization","name":"Etched","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f79de79573098f852f75ad56"},"url":"https://jobsearcher.com/jobs/f79de79573098f852f75ad56"}}