{"schemaVersion":"jobsearcher.job.v1","id":"47e2029bcbb0d6b69f8c35c1","url":"https://jobsearcher.com/jobs/47e2029bcbb0d6b69f8c35c1","canonicalUrl":"https://jobsearcher.com/jobs/47e2029bcbb0d6b69f8c35c1","title":"Manager, Core Infrastructure Engineering","description":"For a single team delivering components of distributed systems. Translates goals into a 1–2 quarter execution plan, sets coding, testing, and scalability practices, and provides hands-on oversight of performance tuning and load/perf testing. Guides the team in building fault-tolerant, in-service-upgradable components (redundancy, replication, failover) and in applying resiliency patterns (retries, circuit breakers, timeouts). Ensures robust observability (tests, alarms, dashboards, telemetry) and operational readiness via reviewed runbooks and standard procedures. Manages delivery of scoped features and correctness testing (including fault-injection/brownouts), and directs implementation of data replication/synchronization to maintain integrity and availability. Leads team incident response and root-cause efforts, enforces no-customer-downtime practices, and drives use of automation/IaC for troubleshooting. Oversees team security implementation (encryption, access controls), tracks remediation plans, verifies compliance documentation, and coaches adherence to change-management plans for safe patching, updates, and rollbacks.\nKey Responsibilities\n\nSystem Design & Architecture – System Scalability:\n\nProvides oversight to the team on the development of components of distributed systems, including the use of distributed state management tools.\nMonitors and enables optimization of code and/or systems for large-scale data processing.\nGuides team to implement scalability requirements for assigned components.\nCoaches team to leverage data plane platforms to effectively handle large-scale data retrieval, storage, and processing.\nEnsures team accurately implements and executes performance and load testing.\n\nSystem Design & Architecture – System Reliability Design:\n\nProvides guidance to the team in building fault-tolerant systems capable of withstanding in-service updates by leading the implementation of redundancy, replication, and automatic failover mechanisms.\nGuides the design of components to effectively handle service disruptions.\nCoaches team on various approaches to handle network unreliability, including retry mechanisms, circuit breakers, and timeouts.\n\nSystem Design & Architecture – System Reliability Performance:\n\nGuides team to implement testing and alarming configurations for detecting and addressing issues/failures.\nEnsures team effectively supports recovery efforts by reviewing and aligning on runbooks and operational procedures.\nCoaches team on building and customizing dashboards, telemetry systems, and alerting mechanisms to monitor component health.\n\nSystem Design & Architecture – Correctness / Availability:\n\nManages the implementation and design of functional requirements and testing for features within an existing system.\nCoaches team on implementing test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.\nGuides the implementation of data replication and synchronization techniques within the team to maintain data integrity and availability.\n\nOperational Troubleshooting & Incident Management:\n\nLeads team efforts in diagnosing, debugging, and resolving issues in system components to support ongoing operation.\nEnsures teams are following protocols for preventing interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.\nManages team in effectively implementing automation scripts and tooling when troubleshooting operational issues.\nCreates schedules and manages operational support rotations.\n\nCompliance & Security:\n\nManages team implementation of robust security measures to protect data and applications in multi-tenant environments, overseeing encryption techniques and access controls.\nManages execution of remediation plans to address identified security gaps, ensuring continuous improvement of security measures.\nReviews documentation and ensures cloud infrastructure compliance with industry standards and regulations.\n\nAutomation & Change Management:\n\nLeads the maintenance of automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.\nCoaches team on change management plans for patching, updating, and rolling back applications.\n\nCore Responsibilities\n\nPlanning & Execution:\n\nCreates and owns the execution plan for the team’s work and multiple projects or initiatives, monitoring timelines and budgets (when applicable) to ensure projects are completed on time and in adherence with requirements.\nDelegates work across the team and helps them prioritize their work.\nIdentifies resource needs and adapts plans based on changing priorities and business needs.\n\nCollaboration & Partnership:\n\nStrengthens collaborative partnerships across teams to align on expectations and shared objectives.\nGuides team members to build relationships with business leaders, stakeholders, and/or customers to ensure effective collaboration.\nPractices active listening and asks insightful questions to promote an inclusive culture.\n\nProblem Solving:\n\nLeads team to identify and address moderately complex operational and/or technical issues in accordance with standard practices, providing guidance as appropriate.\nDirects team to analyze data and/or information from multiple sources to troubleshoot moderately complex errors.\n\nContinuous Learning:\n\nActively seeks learning opportunities for self and team to enhance knowledge and skills in key areas and remain current with industry advancements.\nUses feedback and training to elevate personal and team skills, modeling a commitment to learning.\nIdentifies skill gaps within the team and supports team members in learning opportunities by providing resources and fostering an environment that encourages knowledge-sharing.\n\nContinuous Improvement:\n\nIdentifies and recommends improvements and encourages the team to share ideas to increase the efficiency and effectiveness of processes, protocols, and workflows within a team.\nReviews and provides feedback on team members’ ideas and seeks input from others on alternative approaches and methods for improving work.\n\nPerformance and Development:\n\nCoaches and provides direction to team members in alignment with performance management processes, guidelines, and expectations.\nFacilitates goal-setting discussions with team members, ensuring individual goals are aligned with broader team goals.\nIdentifies development opportunities to help team members improve performance.\nContributes to the talent acquisition pipeline by leading candidate interviews, assessing promotion eligibility, and managing talent resources.","company":"Oracle","rawCompany":"oracle","city":"Seattle","state":"WA","isRemote":false,"isActive":false,"createdAt":"2026-08-22T12:04:36.076Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Manager, Core Infrastructure Engineering","description":"For a single team delivering components of distributed systems. Translates goals into a 1–2 quarter execution plan, sets coding, testing, and scalability practices, and provides hands-on oversight of performance tuning and load/perf testing. Guides the team in building fault-tolerant, in-service-upgradable components (redundancy, replication, failover) and in applying resiliency patterns (retries, circuit breakers, timeouts). Ensures robust observability (tests, alarms, dashboards, telemetry) and operational readiness via reviewed runbooks and standard procedures. Manages delivery of scoped features and correctness testing (including fault-injection/brownouts), and directs implementation of data replication/synchronization to maintain integrity and availability. Leads team incident response and root-cause efforts, enforces no-customer-downtime practices, and drives use of automation/IaC for troubleshooting. Oversees team security implementation (encryption, access controls), tracks remediation plans, verifies compliance documentation, and coaches adherence to change-management plans for safe patching, updates, and rollbacks.\nKey Responsibilities\n\nSystem Design & Architecture – System Scalability:\n\nProvides oversight to the team on the development of components of distributed systems, including the use of distributed state management tools.\nMonitors and enables optimization of code and/or systems for large-scale data processing.\nGuides team to implement scalability requirements for assigned components.\nCoaches team to leverage data plane platforms to effectively handle large-scale data retrieval, storage, and processing.\nEnsures team accurately implements and executes performance and load testing.\n\nSystem Design & Architecture – System Reliability Design:\n\nProvides guidance to the team in building fault-tolerant systems capable of withstanding in-service updates by leading the implementation of redundancy, replication, and automatic failover mechanisms.\nGuides the design of components to effectively handle service disruptions.\nCoaches team on various approaches to handle network unreliability, including retry mechanisms, circuit breakers, and timeouts.\n\nSystem Design & Architecture – System Reliability Performance:\n\nGuides team to implement testing and alarming configurations for detecting and addressing issues/failures.\nEnsures team effectively supports recovery efforts by reviewing and aligning on runbooks and operational procedures.\nCoaches team on building and customizing dashboards, telemetry systems, and alerting mechanisms to monitor component health.\n\nSystem Design & Architecture – Correctness / Availability:\n\nManages the implementation and design of functional requirements and testing for features within an existing system.\nCoaches team on implementing test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.\nGuides the implementation of data replication and synchronization techniques within the team to maintain data integrity and availability.\n\nOperational Troubleshooting & Incident Management:\n\nLeads team efforts in diagnosing, debugging, and resolving issues in system components to support ongoing operation.\nEnsures teams are following protocols for preventing interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.\nManages team in effectively implementing automation scripts and tooling when troubleshooting operational issues.\nCreates schedules and manages operational support rotations.\n\nCompliance & Security:\n\nManages team implementation of robust security measures to protect data and applications in multi-tenant environments, overseeing encryption techniques and access controls.\nManages execution of remediation plans to address identified security gaps, ensuring continuous improvement of security measures.\nReviews documentation and ensures cloud infrastructure compliance with industry standards and regulations.\n\nAutomation & Change Management:\n\nLeads the maintenance of automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.\nCoaches team on change management plans for patching, updating, and rolling back applications.\n\nCore Responsibilities\n\nPlanning & Execution:\n\nCreates and owns the execution plan for the team’s work and multiple projects or initiatives, monitoring timelines and budgets (when applicable) to ensure projects are completed on time and in adherence with requirements.\nDelegates work across the team and helps them prioritize their work.\nIdentifies resource needs and adapts plans based on changing priorities and business needs.\n\nCollaboration & Partnership:\n\nStrengthens collaborative partnerships across teams to align on expectations and shared objectives.\nGuides team members to build relationships with business leaders, stakeholders, and/or customers to ensure effective collaboration.\nPractices active listening and asks insightful questions to promote an inclusive culture.\n\nProblem Solving:\n\nLeads team to identify and address moderately complex operational and/or technical issues in accordance with standard practices, providing guidance as appropriate.\nDirects team to analyze data and/or information from multiple sources to troubleshoot moderately complex errors.\n\nContinuous Learning:\n\nActively seeks learning opportunities for self and team to enhance knowledge and skills in key areas and remain current with industry advancements.\nUses feedback and training to elevate personal and team skills, modeling a commitment to learning.\nIdentifies skill gaps within the team and supports team members in learning opportunities by providing resources and fostering an environment that encourages knowledge-sharing.\n\nContinuous Improvement:\n\nIdentifies and recommends improvements and encourages the team to share ideas to increase the efficiency and effectiveness of processes, protocols, and workflows within a team.\nReviews and provides feedback on team members’ ideas and seeks input from others on alternative approaches and methods for improving work.\n\nPerformance and Development:\n\nCoaches and provides direction to team members in alignment with performance management processes, guidelines, and expectations.\nFacilitates goal-setting discussions with team members, ensuring individual goals are aligned with broader team goals.\nIdentifies development opportunities to help team members improve performance.\nContributes to the talent acquisition pipeline by leading candidate interviews, assessing promotion eligibility, and managing talent resources.","datePosted":"2026-08-22T12:04:36.076Z","dateModified":"2026-08-22T12:04:36.076Z","hiringOrganization":{"@type":"Organization","name":"Oracle","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Seattle","addressRegion":"WA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"47e2029bcbb0d6b69f8c35c1"},"url":"https://jobsearcher.com/jobs/47e2029bcbb0d6b69f8c35c1"}}