{"schemaVersion":"jobsearcher.job.v1","id":"c5bbfd3b948ae12afbe239b1","url":"https://jobsearcher.com/jobs/c5bbfd3b948ae12afbe239b1","canonicalUrl":"https://jobsearcher.com/jobs/c5bbfd3b948ae12afbe239b1","title":"Software Engineering Manager-Production Support Operations","description":"Job Description Summary\r\nThe Manager of Production Support leads teams responsible for ensuring the stability, resilience, and operational excellence of critical technology platforms supporting core lines of business. This role owns end-to-end production support operations while driving maturity toward engineering-first, site reliability–focused practices. The Director identifies and resolves complex technical, operational, risk, and organizational challenges, while building high-performing, accountable teams across onshore and offshore locations.\r\nThis position carries full people management responsibility, including hiring, coaching, performance management, and disciplinary actions, and serves as a key partner to Technology, Risk, and Business leadership.\r\nEssential Duties And Responsibilities\r\nFollowing is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time.\r\nProduction Support Leadership & Accountability\r\nOwn end-to-end production support operations for multiple mission-critical applications supporting key lines of business, ensuring availability, stability, and performance meet defined SLAs and SLOs. Provide accountable, visible leadership for 24x7 operational support, including on-call models, escalation paths, and incident response effectiveness. Act as the senior escalation point for major incidents, ensuring swift recovery, accurate root cause analysis, and durable remediation.\r\nIncident & Problem Management\r\nLead cross-functional incident recovery efforts in partnership with Incident Management, engineering teams, infrastructure, and business stakeholders. Ensure timely root cause analysis (RCA), post-incident reviews, and corrective actions that prevent recurrence. Establish and mature a production knowledge base, documenting known issues, recovery procedures, and architectural insights.\r\nEngineering-First & SRE Practices\r\nDrive adoption of Site Reliability Engineering (SRE) and lean engineering principles, including: Reduction of toil through automation; Engineering-based reliability metrics (error budgets, SLIs/SLOs); Proactive resilience and failure prevention practices; Champion automation of repetitive and manual operational tasks, including incident detection, response, validation, and recovery where feasible. Promote a culture of preventative engineering, partnering with development teams to improve system reliability upstream.\r\nMonitoring, Observability & AI Enablement\r\nImplement and continuously improve real-time monitoring, alerting, and observability across applications and infrastructure. Measure and optimize the effectiveness of monitoring and alerting to eliminate noise and accelerate mean-time-to-detect and mean-time-to-recover. Leverage AI and advanced analytics to correlate telemetry data (logs, metrics, traces) and proactively identify emerging risks and root causes. Champion the safe and responsible use of AI within production operations by adhering to enterprise guardrails and protecting sensitive data and system integrity.\r\nOperational Readiness & Change Enablement\r\nOversee operational readiness across releases, disaster recovery and failover testing and certificate and dependency lifecycle management. Ensure production support is actively embedded in change planning, minimizing risk from releases and infrastructure changes.\r\nPeople, Vendor & Financial Management\r\nLead one or more Agile teams (Scrum, Kanban), including onshore and offshore engineers, fostering high performance and accountability. Manage workforce vendors and partners, setting expectations, reviewing performance, and ensuring delivery quality. Own budget and staffing plan aligned to application criticality, operational risk, and business growth objectives.\r\nRisk Management & Governance\r\nAct as the first line of defense in production operations by proactively identifying and mitigating technology, operational, and resiliency risks. Partner effectively with second-line Risk, Audit, and Regulatory teams, ensuring findings are addressed and controls are continuously improved. Ensure compliance with internal policies, regulatory requirements, and external audit expectations. Own and drive remediation plans for risk, audit, and regulatory findings, ensuring timely, effective and sustainable resolution. Lead responses to audit and regulatory inquiries, including providing evidence, clarifying controls, and appropriately challenging findings based on documented compliance.\r\nStrategy, Influence & Continuous Improvement\r\nServe as a trusted advisor to senior Technology and Business leaders, communicating operational health, risk posture, and improvement roadmaps. Lead or contribute significantly to large-scale initiatives, platform transformations, or regulatory-driven efforts. Continuously assess organizational maturity and lead initiatives to improve reliability, efficiency, and talent capability.\r\nManagement Responsibilities\r\nFull people management accountability, including:\r\nHiring and succession planning\r\nCoaching and performance management\r\nCompensation input and talent development\r\nDisciplinary action and terminations as necessary\r\nAgile & Operating Model Expectations\r\nAct as an Agile and DevOps champion, embedding production support within fast-moving delivery models.\r\nBalance \"keep-the-lights-on\" operational excellence with continuous engineering improvement.\r\nDrive measurable outcomes such as improved uptime, reduced incident volume, faster recovery, and improved customer experience.\r\nQualifications\r\nRequired Qualifications\r\nBachelor's degree in Computer Science, Software Engineering, or a related technical field, or equivalent practical experience.\r\nA minimum of 5 years of professional software engineering experience, including team leadership or supervisory responsibilities.\r\nPreferred Qualifications\r\nUnderstanding of multiple approaches to production support and software engineering delivery.\r\nFull understanding of Agile methodology.\r\nExperience leading teams in an Agile organization, particularly those practicing Site Reliability Engineering.\r\nExperience using AI agents in day-to-day activities, particularly in regard to enabling software delivery and production support operations.\r\nBanking or financial services experience.\r\nBachelor's degree and twelve years of experience in software development, production support, including five years of management experience\r\nGeneral Description of Available Benefits for Eligible Employees of Truist Financial Corporation\r\nAll regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position.\r\nTruist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates.\r\nTeammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays.\r\nFor more details on Truist's generous benefit plans, please visit our Benefits site.\r\nDepending on the position and division, this job may also be eligible for Truist's defined benefit pension plan, restricted stock units, and/or a deferred compensation plan.\r\nAs you advance through the hiring process, you will also learn more about the specific benefits available for any non-temporary position for which you apply, based on full-time or part-time status, position, and division of work.\r\nTruist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace.\r\nEEO is the Law E-Verify IER Right to Work\r\nRegular or Temporary\r\nRegular\r\nLanguage Fluency\r\nEnglish (Required)\r\nWork Shift\r\n1st shift (United States of America)\r\nNeed Help? If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).\r\nJ-18808-Ljbffr","company":"Socket","rawCompany":"socket","city":"Charlotte","state":"NC","isRemote":false,"isActive":false,"createdAt":"2026-08-08T01:59:27.473Z","occupations":[{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineering Manager-Production Support Operations","description":"Job Description Summary\r\nThe Manager of Production Support leads teams responsible for ensuring the stability, resilience, and operational excellence of critical technology platforms supporting core lines of business. This role owns end-to-end production support operations while driving maturity toward engineering-first, site reliability–focused practices. The Director identifies and resolves complex technical, operational, risk, and organizational challenges, while building high-performing, accountable teams across onshore and offshore locations.\r\nThis position carries full people management responsibility, including hiring, coaching, performance management, and disciplinary actions, and serves as a key partner to Technology, Risk, and Business leadership.\r\nEssential Duties And Responsibilities\r\nFollowing is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time.\r\nProduction Support Leadership & Accountability\r\nOwn end-to-end production support operations for multiple mission-critical applications supporting key lines of business, ensuring availability, stability, and performance meet defined SLAs and SLOs. Provide accountable, visible leadership for 24x7 operational support, including on-call models, escalation paths, and incident response effectiveness. Act as the senior escalation point for major incidents, ensuring swift recovery, accurate root cause analysis, and durable remediation.\r\nIncident & Problem Management\r\nLead cross-functional incident recovery efforts in partnership with Incident Management, engineering teams, infrastructure, and business stakeholders. Ensure timely root cause analysis (RCA), post-incident reviews, and corrective actions that prevent recurrence. Establish and mature a production knowledge base, documenting known issues, recovery procedures, and architectural insights.\r\nEngineering-First & SRE Practices\r\nDrive adoption of Site Reliability Engineering (SRE) and lean engineering principles, including: Reduction of toil through automation; Engineering-based reliability metrics (error budgets, SLIs/SLOs); Proactive resilience and failure prevention practices; Champion automation of repetitive and manual operational tasks, including incident detection, response, validation, and recovery where feasible. Promote a culture of preventative engineering, partnering with development teams to improve system reliability upstream.\r\nMonitoring, Observability & AI Enablement\r\nImplement and continuously improve real-time monitoring, alerting, and observability across applications and infrastructure. Measure and optimize the effectiveness of monitoring and alerting to eliminate noise and accelerate mean-time-to-detect and mean-time-to-recover. Leverage AI and advanced analytics to correlate telemetry data (logs, metrics, traces) and proactively identify emerging risks and root causes. Champion the safe and responsible use of AI within production operations by adhering to enterprise guardrails and protecting sensitive data and system integrity.\r\nOperational Readiness & Change Enablement\r\nOversee operational readiness across releases, disaster recovery and failover testing and certificate and dependency lifecycle management. Ensure production support is actively embedded in change planning, minimizing risk from releases and infrastructure changes.\r\nPeople, Vendor & Financial Management\r\nLead one or more Agile teams (Scrum, Kanban), including onshore and offshore engineers, fostering high performance and accountability. Manage workforce vendors and partners, setting expectations, reviewing performance, and ensuring delivery quality. Own budget and staffing plan aligned to application criticality, operational risk, and business growth objectives.\r\nRisk Management & Governance\r\nAct as the first line of defense in production operations by proactively identifying and mitigating technology, operational, and resiliency risks. Partner effectively with second-line Risk, Audit, and Regulatory teams, ensuring findings are addressed and controls are continuously improved. Ensure compliance with internal policies, regulatory requirements, and external audit expectations. Own and drive remediation plans for risk, audit, and regulatory findings, ensuring timely, effective and sustainable resolution. Lead responses to audit and regulatory inquiries, including providing evidence, clarifying controls, and appropriately challenging findings based on documented compliance.\r\nStrategy, Influence & Continuous Improvement\r\nServe as a trusted advisor to senior Technology and Business leaders, communicating operational health, risk posture, and improvement roadmaps. Lead or contribute significantly to large-scale initiatives, platform transformations, or regulatory-driven efforts. Continuously assess organizational maturity and lead initiatives to improve reliability, efficiency, and talent capability.\r\nManagement Responsibilities\r\nFull people management accountability, including:\r\nHiring and succession planning\r\nCoaching and performance management\r\nCompensation input and talent development\r\nDisciplinary action and terminations as necessary\r\nAgile & Operating Model Expectations\r\nAct as an Agile and DevOps champion, embedding production support within fast-moving delivery models.\r\nBalance \"keep-the-lights-on\" operational excellence with continuous engineering improvement.\r\nDrive measurable outcomes such as improved uptime, reduced incident volume, faster recovery, and improved customer experience.\r\nQualifications\r\nRequired Qualifications\r\nBachelor's degree in Computer Science, Software Engineering, or a related technical field, or equivalent practical experience.\r\nA minimum of 5 years of professional software engineering experience, including team leadership or supervisory responsibilities.\r\nPreferred Qualifications\r\nUnderstanding of multiple approaches to production support and software engineering delivery.\r\nFull understanding of Agile methodology.\r\nExperience leading teams in an Agile organization, particularly those practicing Site Reliability Engineering.\r\nExperience using AI agents in day-to-day activities, particularly in regard to enabling software delivery and production support operations.\r\nBanking or financial services experience.\r\nBachelor's degree and twelve years of experience in software development, production support, including five years of management experience\r\nGeneral Description of Available Benefits for Eligible Employees of Truist Financial Corporation\r\nAll regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position.\r\nTruist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates.\r\nTeammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays.\r\nFor more details on Truist's generous benefit plans, please visit our Benefits site.\r\nDepending on the position and division, this job may also be eligible for Truist's defined benefit pension plan, restricted stock units, and/or a deferred compensation plan.\r\nAs you advance through the hiring process, you will also learn more about the specific benefits available for any non-temporary position for which you apply, based on full-time or part-time status, position, and division of work.\r\nTruist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace.\r\nEEO is the Law E-Verify IER Right to Work\r\nRegular or Temporary\r\nRegular\r\nLanguage Fluency\r\nEnglish (Required)\r\nWork Shift\r\n1st shift (United States of America)\r\nNeed Help? If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).\r\nJ-18808-Ljbffr","datePosted":"2026-08-08T01:59:27.473Z","dateModified":"2026-08-08T01:59:27.473Z","hiringOrganization":{"@type":"Organization","name":"Socket","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Charlotte","addressRegion":"NC","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"c5bbfd3b948ae12afbe239b1"},"url":"https://jobsearcher.com/jobs/c5bbfd3b948ae12afbe239b1"}}