{"schemaVersion":"jobsearcher.job.v1","id":"de2d251a94fd178c8c367c7a","url":"https://jobsearcher.com/jobs/de2d251a94fd178c8c367c7a","canonicalUrl":"https://jobsearcher.com/jobs/de2d251a94fd178c8c367c7a","title":"Site Reliability Engineer","description":"The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE Office of Science programs. NERSC provides critical HPC and data systems and support for NERSC’s 11,000 plus users researching energy, physics, materials science, and chemistry and other DOE mission areas. As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work ensures that NERSC’s computational power remains an uninterrupted catalyst supporting fundamental scientific research related to energy.\n\nESSENTIAL DUTIES & RESPONSIBILITIES\n\nDescribe the duties/functions essential to performing the job.\n\nWorks an onsite 5-day weekly schedule consisting of Owl (midnight–8 am) shifts to monitor the NERSC HPC Facility.\nReview and respond to alerts from computer systems, storage, network, and other data center/facility-related systems by triaging or calling the appropriate on-call staff.\nCreate appropriate solutions to improve processes, prevent issue recurrence, and automate responses to all routine service conditions.\nIdentify issues and propose solutions that will improve monitoring capabilities or provide better automation for triage.\nPossess expertise in ServiceNow and its usage to develop and implement customized service management solutions.\nRespond to alerts from multiple systems to ensure that data collection continues 24/7, providing real-time information for diagnoses.\nDevelop and maintain tools within the monitoring pipeline in collaboration with the Operations Team.\nCreate new software to provide alerts and notifications from HPC system APIs into the monitoring pipeline.\nBuilds and maintains application/tool configurations to ensure software runs reliably as data and user demands grow.\nCollaborate with other groups at NERSC to ensure that communication and workflows are clearly understood.\nWork closely with other NERSC groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periods.\nPerform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure to ensure peak operational efficiency.\nProvide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents so that workflows and protocols can be appropriately tracked by others.\nWork on and resolve problems of diverse scope where data analysis requires the evaluation of identifiable factors.\nDemonstrate good judgment in selecting methods and techniques for obtaining solutions.\nWork on and resolve complex issues where the analysis of situations or data requires an in-depth evaluation of variable factors.\n\nPOSITION REQUIREMENTS\nExperience in or willingness to work within a 24/7 onsite team environment to support large-scale data centers or critical installations.\nExperience on Linux shell and working in a command-line (e.g. SSH) environment.\nExperience with developing tools using various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.\nMotivated, self-starter who can learn technologies that improve data center management in areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative\n\ncooling, and power utilization.\n\nExperience with network security: configuring/maintaining ACLs, knowledge of firewalls\nExperience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.\nGood to Have : Practical experience in developing and deploying Agentic AI or autonomous automation tools to streamline technical tasks.\nExperience with ServiceNow implementation is a plus\nFamiliarity with ITSM best practices and an understanding of how to align service lifecycles with business goals is preferred.\n\nKnowledge, Skills & Abilities\nStrong hands-on knowledge of the Linux shell and working in a command-line (e.g. SSH) environment.\nStrong hands-on knowledge on various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.\nKnowledge of and ability to work on large data communications networks/ Network Protocols and IT infrastructure supporting highly available systems and applications.\nStrong communication skills and ability to work effectively across multiple technical teams.\nGood to Have : Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to automate decision-making, optimize complex workflows, and enhance proactive system monitoring.","company":"Bay Systems Consulting","rawCompany":"bay systems consulting","city":"Berkeley","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-09T08:57:39.166Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1232.00","title":"Computer User Support Specialists","slug":"computer-user-support-specialists"}],"industries":[{"code":"541513","title":"Computer Facilities Management Services","slug":"computer-facilities-management-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer","description":"The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE Office of Science programs. NERSC provides critical HPC and data systems and support for NERSC’s 11,000 plus users researching energy, physics, materials science, and chemistry and other DOE mission areas. As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work ensures that NERSC’s computational power remains an uninterrupted catalyst supporting fundamental scientific research related to energy.\n\nESSENTIAL DUTIES & RESPONSIBILITIES\n\nDescribe the duties/functions essential to performing the job.\n\nWorks an onsite 5-day weekly schedule consisting of Owl (midnight–8 am) shifts to monitor the NERSC HPC Facility.\nReview and respond to alerts from computer systems, storage, network, and other data center/facility-related systems by triaging or calling the appropriate on-call staff.\nCreate appropriate solutions to improve processes, prevent issue recurrence, and automate responses to all routine service conditions.\nIdentify issues and propose solutions that will improve monitoring capabilities or provide better automation for triage.\nPossess expertise in ServiceNow and its usage to develop and implement customized service management solutions.\nRespond to alerts from multiple systems to ensure that data collection continues 24/7, providing real-time information for diagnoses.\nDevelop and maintain tools within the monitoring pipeline in collaboration with the Operations Team.\nCreate new software to provide alerts and notifications from HPC system APIs into the monitoring pipeline.\nBuilds and maintains application/tool configurations to ensure software runs reliably as data and user demands grow.\nCollaborate with other groups at NERSC to ensure that communication and workflows are clearly understood.\nWork closely with other NERSC groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periods.\nPerform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure to ensure peak operational efficiency.\nProvide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents so that workflows and protocols can be appropriately tracked by others.\nWork on and resolve problems of diverse scope where data analysis requires the evaluation of identifiable factors.\nDemonstrate good judgment in selecting methods and techniques for obtaining solutions.\nWork on and resolve complex issues where the analysis of situations or data requires an in-depth evaluation of variable factors.\n\nPOSITION REQUIREMENTS\nExperience in or willingness to work within a 24/7 onsite team environment to support large-scale data centers or critical installations.\nExperience on Linux shell and working in a command-line (e.g. SSH) environment.\nExperience with developing tools using various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.\nMotivated, self-starter who can learn technologies that improve data center management in areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative\n\ncooling, and power utilization.\n\nExperience with network security: configuring/maintaining ACLs, knowledge of firewalls\nExperience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.\nGood to Have : Practical experience in developing and deploying Agentic AI or autonomous automation tools to streamline technical tasks.\nExperience with ServiceNow implementation is a plus\nFamiliarity with ITSM best practices and an understanding of how to align service lifecycles with business goals is preferred.\n\nKnowledge, Skills & Abilities\nStrong hands-on knowledge of the Linux shell and working in a command-line (e.g. SSH) environment.\nStrong hands-on knowledge on various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.\nKnowledge of and ability to work on large data communications networks/ Network Protocols and IT infrastructure supporting highly available systems and applications.\nStrong communication skills and ability to work effectively across multiple technical teams.\nGood to Have : Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to automate decision-making, optimize complex workflows, and enhance proactive system monitoring.","datePosted":"2026-09-09T08:57:39.166Z","dateModified":"2026-09-09T08:57:39.166Z","hiringOrganization":{"@type":"Organization","name":"Bay Systems Consulting","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Berkeley","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"de2d251a94fd178c8c367c7a"},"url":"https://jobsearcher.com/jobs/de2d251a94fd178c8c367c7a"}}