{"schemaVersion":"jobsearcher.job.v1","id":"e7fdde2d9f35a4e1235a3d0f","url":"https://jobsearcher.com/jobs/e7fdde2d9f35a4e1235a3d0f","canonicalUrl":"https://jobsearcher.com/jobs/e7fdde2d9f35a4e1235a3d0f","title":"Site Reliability Engineer","description":"Senior Site Reliability Engineer (SRE) nLocation: Pittsburgh, PA / Cleveland, OH / Dallas, TX nFTE nPosition Overview nWe are seeking an experienced Senior Site Reliability Engineer (SRE) to support production operations, application reliability, performance management, and continuous improvement initiatives. nThe selected candidate will work closely with production support and engineering teams to ensure critical internal and external applications maintain appropriate levels of availability, reliability, and uptime. nThis role requires strong experience in production support, incident management, monitoring, troubleshooting, log analysis, automation identification, infrastructure technologies, databases, and application servers. The SRE will also provide technical leadership and collaborate with geographically distributed teams. nKey Skills nnnSite Reliability Engineering (SRE) n nnProduction Support / Application Support n nnIncident & Problem Management n nnLinux n nnWindows Server n nnOracle / PL/SQL / DB2 n nnDynatrace / DT Managed n nnGlassBox / ITCAM / TrueSight / OEM n nnTomcat / Apache / WebSphere (WAS) / IIS n nnREST & SOAP Web Services n nnLog Analysis & Troubleshooting n nnAIOps / NLP n nnMonitoring & Performance Management n nnAutomation n nnRoot Cause Analysis n nnBusiness Analytics n nnAgile n nnTechnical Leadership n nnClient-Facing Production Support n n nResponsibilities nnnMonitor distributed systems and proactively identify potential production issues. n nnSupport troubleshooting and participate in on-call activities. n nnManage, track, and coordinate production incidents and application outages. n nnLead incident-analysis and problem-management meetings. n nnIdentify opportunities for operational and production-support automation. n nnMonitor applications and related infrastructure to maintain system reliability. n nnCoordinate follow-up activities through incident resolution and closure. n nnTroubleshoot complex application issues using system and application logs. n nnParticipate in critical incident calls and contribute technical expertise toward resolution. n nnPerform root cause analysis and recommend corrective actions. n nnResearch and reproduce user issues to validate solutions. n nnResolve technical problems that cannot be handled by junior team members. n nnProvide technical guidance and solutions to the production-support team. n nnIntroduce process improvements and innovative solutions for operational challenges. n nnDevelop and maintain SOPs, operational procedures, and knowledge documentation. n nnCollaborate with offshore and geographically distributed teams. n nnWork with client technical teams, SMEs, and leadership. n nnSupport extended or weekend hours when required during critical production events. n nnParticipate in overlapping business-hour shifts for critical meetings and activities. n n nRequired Qualifications nnn5+ years of overall IT experience. n nn2-3 years of business analytics and technical leadership experience. n nnStrong experience with production/application support in a client-facing environment. n nnStrong understanding of Site Reliability Engineering and production operations. n nnHands-on experience troubleshooting production applications and analyzing log files. n nnStrong knowledge of system-management, monitoring, and support analytics tools. n nnExperience with incident management, root cause analysis, and problem resolution. n nnStrong understanding of AIOps and NLP concepts. n nnExperience identifying opportunities for automation and process improvement. n nnStrong problem-solving and analytical capabilities. n nnAbility to recommend efficient and cost-effective technical solutions. n nnExperience working with geographically distributed/onshore-offshore teams. n nnExcellent client-facing verbal and written communication skills. n n nDatabase Technologies nStrong knowledge of: nnnOracle n nnPL/SQL n nnDB2 n n nWeb Services nExperience developing and consuming: nnnREST APIs n nnSOAP Web Services n n nExperience should preferably be within an operational/production environment. nApplication Servers / Web Servers nStrong knowledge of: nnnTomcat n nnApache n nnWebSphere (WAS) n nnIIS n n nOperating Systems nnnExtensive experience with Linux n nnGood understanding of Windows Server n nnLinux and Windows server configuration and troubleshooting n n nMonitoring & Support Tools nExperience with monitoring tools such as: nnnDynatrace n nnDynatrace Managed / DT Managed n nnGlassBox n nnITCAM / ITCAMS n nnTrueSight n nnOracle Enterprise Manager (OEM) n n nAdditional Skills nnnAgile methodology n nnSOP and technical documentation n nnPerformance management n nnSystem reliability and availability n nnProduction incident coordination n nnTechnical research and solution evaluation n nnProcess improvement n nnAutomation opportunity identification n nnStrong stakeholder and client communication n n n#M1 n#DI-CB2 n#L1 - KB1 nRef: #404-IT Pittsburgh nSystem One, and its subsidiaries including Joulé, ALTA IT Services, CM Access, TPGS, and MOUNTAIN, LTD., are leaders in delivering workforce solutions and integrated services across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible full-time employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan. nSystem One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.","company":"System One","rawCompany":"system one","city":"Pittsburgh","state":"PA","isRemote":false,"isActive":true,"createdAt":"2026-08-15T22:49:07.134Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer","description":"Senior Site Reliability Engineer (SRE) nLocation: Pittsburgh, PA / Cleveland, OH / Dallas, TX nFTE nPosition Overview nWe are seeking an experienced Senior Site Reliability Engineer (SRE) to support production operations, application reliability, performance management, and continuous improvement initiatives. nThe selected candidate will work closely with production support and engineering teams to ensure critical internal and external applications maintain appropriate levels of availability, reliability, and uptime. nThis role requires strong experience in production support, incident management, monitoring, troubleshooting, log analysis, automation identification, infrastructure technologies, databases, and application servers. The SRE will also provide technical leadership and collaborate with geographically distributed teams. nKey Skills nnnSite Reliability Engineering (SRE) n nnProduction Support / Application Support n nnIncident & Problem Management n nnLinux n nnWindows Server n nnOracle / PL/SQL / DB2 n nnDynatrace / DT Managed n nnGlassBox / ITCAM / TrueSight / OEM n nnTomcat / Apache / WebSphere (WAS) / IIS n nnREST & SOAP Web Services n nnLog Analysis & Troubleshooting n nnAIOps / NLP n nnMonitoring & Performance Management n nnAutomation n nnRoot Cause Analysis n nnBusiness Analytics n nnAgile n nnTechnical Leadership n nnClient-Facing Production Support n n nResponsibilities nnnMonitor distributed systems and proactively identify potential production issues. n nnSupport troubleshooting and participate in on-call activities. n nnManage, track, and coordinate production incidents and application outages. n nnLead incident-analysis and problem-management meetings. n nnIdentify opportunities for operational and production-support automation. n nnMonitor applications and related infrastructure to maintain system reliability. n nnCoordinate follow-up activities through incident resolution and closure. n nnTroubleshoot complex application issues using system and application logs. n nnParticipate in critical incident calls and contribute technical expertise toward resolution. n nnPerform root cause analysis and recommend corrective actions. n nnResearch and reproduce user issues to validate solutions. n nnResolve technical problems that cannot be handled by junior team members. n nnProvide technical guidance and solutions to the production-support team. n nnIntroduce process improvements and innovative solutions for operational challenges. n nnDevelop and maintain SOPs, operational procedures, and knowledge documentation. n nnCollaborate with offshore and geographically distributed teams. n nnWork with client technical teams, SMEs, and leadership. n nnSupport extended or weekend hours when required during critical production events. n nnParticipate in overlapping business-hour shifts for critical meetings and activities. n n nRequired Qualifications nnn5+ years of overall IT experience. n nn2-3 years of business analytics and technical leadership experience. n nnStrong experience with production/application support in a client-facing environment. n nnStrong understanding of Site Reliability Engineering and production operations. n nnHands-on experience troubleshooting production applications and analyzing log files. n nnStrong knowledge of system-management, monitoring, and support analytics tools. n nnExperience with incident management, root cause analysis, and problem resolution. n nnStrong understanding of AIOps and NLP concepts. n nnExperience identifying opportunities for automation and process improvement. n nnStrong problem-solving and analytical capabilities. n nnAbility to recommend efficient and cost-effective technical solutions. n nnExperience working with geographically distributed/onshore-offshore teams. n nnExcellent client-facing verbal and written communication skills. n n nDatabase Technologies nStrong knowledge of: nnnOracle n nnPL/SQL n nnDB2 n n nWeb Services nExperience developing and consuming: nnnREST APIs n nnSOAP Web Services n n nExperience should preferably be within an operational/production environment. nApplication Servers / Web Servers nStrong knowledge of: nnnTomcat n nnApache n nnWebSphere (WAS) n nnIIS n n nOperating Systems nnnExtensive experience with Linux n nnGood understanding of Windows Server n nnLinux and Windows server configuration and troubleshooting n n nMonitoring & Support Tools nExperience with monitoring tools such as: nnnDynatrace n nnDynatrace Managed / DT Managed n nnGlassBox n nnITCAM / ITCAMS n nnTrueSight n nnOracle Enterprise Manager (OEM) n n nAdditional Skills nnnAgile methodology n nnSOP and technical documentation n nnPerformance management n nnSystem reliability and availability n nnProduction incident coordination n nnTechnical research and solution evaluation n nnProcess improvement n nnAutomation opportunity identification n nnStrong stakeholder and client communication n n n#M1 n#DI-CB2 n#L1 - KB1 nRef: #404-IT Pittsburgh nSystem One, and its subsidiaries including Joulé, ALTA IT Services, CM Access, TPGS, and MOUNTAIN, LTD., are leaders in delivering workforce solutions and integrated services across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible full-time employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan. nSystem One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.","datePosted":"2026-08-15T22:49:07.134Z","dateModified":"2026-08-15T22:49:07.134Z","hiringOrganization":{"@type":"Organization","name":"System One","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Pittsburgh","addressRegion":"PA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e7fdde2d9f35a4e1235a3d0f"},"url":"https://jobsearcher.com/jobs/e7fdde2d9f35a4e1235a3d0f"}}