{"schemaVersion":"jobsearcher.job.v1","id":"1025f122489f7b58856d8d8b","url":"https://jobsearcher.com/jobs/1025f122489f7b58856d8d8b","canonicalUrl":"https://jobsearcher.com/jobs/1025f122489f7b58856d8d8b","title":"Java Trainer / Mentor / Educator","description":"Irving, TX Work schedule (days & times) - M-F 8 AM to 5 PM PST\nSite Reliability Engineer Position Description:\nThe Site Reliability Engineer designs, enhances, and operates highly reliable, scalable, and observable production systems in an Azure-based environment. This role blends software engineering with systems administration to build resilient infrastructure, automate operations, and improve system performance. The engineer applies strong engineering principles to operational challenges with a focus on reliability, automation, observability, and continuous improvement.\nCore responsibilities include engineering led incident response, implementing permanent corrective actions, reducing operational toil, and proactively preventing failures. The role contributes to code fixes, owns Dynatrace based observability, and delivers custom reliability and operational reporting to improve system health and availability. Participation in a scheduled-on call rotation is required.\nMinimum Requirement:\n4-year Computer Science, Information Systems, Engineering degree or relevant experience. 7+ Years of Site reliability experience. 8+ Years of overall experience.\nKey Responsibilities:\nDesign, implement, and maintain monitoring solutions to ensure system health and performance.\nDevelop and manage CI/CD pipelines using GitHub Actions.\nDeploy, manage, and troubleshoot containerized applications using Docker and Kubernetes.\nSupport and optimize Java-based applications in production environments.\nCollaborate with development teams to improve system reliability and reduce operational toil.\nImplement best practices for incident response, capacity planning, and disaster recovery.\nProvision and manage infrastructure using Azure cloud services.\nImprove system observability using tools such as Dynatrace (preferred).\nPerform Linux and Windows system administration, including patching, configuration, and troubleshooting.\nAutomate operational tasks using Ansible, Python, Bash, or similar tools.\nAdvanced SRE Leadership Responsibilities:\nProvide technical leadership for SRE practices across multiple services or platforms.\nDefine and evolve reliability standards, operational best practices, and incident response frameworks.\nInfluence system architecture and design decisions to ensure scalability, resilience, and operability.\nServe as a subject matter expert for reliability, availability, and production risk management.\nAct as the lead escalation point for complex and business critical production incidents.\nLead high severity incident response, coordinating across engineering, platform, and security teams.\nDrive blameless post incident reviews and ensure corrective actions are prioritized and completed.\nImprove call processes, escalation models, and incident response effectiveness.\nOwn the strategy and implementation of Dynatrace based observability, including dashboards and alerting standards.\nEstablish and monitor reliability signals (availability, latency, error rates) across critical systems.\nIdentify reliability risks and lead mitigation initiatives before customer impact occurs.\nDefine and maintain leadership level reliability and operational reporting.\nUse production data to drive prioritization of reliability investments and operational improvements.\nCommunicate reliability posture, risks, and recommendations to senior engineering leadership.\nMentor and guide senior and midlevel SREs and production support engineers.\nSupport hiring, onboarding, and technical evaluation of SRE talent.\nCollaborate with squad members to define iteration plans and commitments.\nEnsure compliance with HIPAA and other security regulations.\nCritical Skills:\nStrong experience with monitoring and observability tools (Dynatrace experience is a plus).\nHands-on experience with GitHub Actions for CI/CD automation.\nProficiency in Kubernetes and Docker for container orchestration.\nFamiliarity with Azure cloud services.\nExperience with Ansible.\nDemonstrated experience in automation of infrastructure and operational processes using scripting or configuration management tools.\nExperience supporting Java applications in production.\nSolid understanding of Linux and Windows system administration.\nKnowledge of SRE principles (SLIs, SLOs, error budgets).\nAdditional Skills:\nExperience working with onsite and offshore teams.\nStrong communication skills (written and verbal).\nStrong organizational skills, attention to detail, and ability to multitask.\nExperience in healthcare software or compliance solutions is a plus.\nStrong analytical and problem-solving skills with the ability to identify root causes and propose effective solutions.\nFor applications and inquiries, contact: hirings@openkyber.com","company":"Openkyber","rawCompany":"openkyber","city":"Texas","state":"PA","isRemote":false,"isActive":false,"createdAt":"2026-08-06T13:34:32.906Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Java Trainer / Mentor / Educator","description":"Irving, TX Work schedule (days & times) - M-F 8 AM to 5 PM PST\nSite Reliability Engineer Position Description:\nThe Site Reliability Engineer designs, enhances, and operates highly reliable, scalable, and observable production systems in an Azure-based environment. This role blends software engineering with systems administration to build resilient infrastructure, automate operations, and improve system performance. The engineer applies strong engineering principles to operational challenges with a focus on reliability, automation, observability, and continuous improvement.\nCore responsibilities include engineering led incident response, implementing permanent corrective actions, reducing operational toil, and proactively preventing failures. The role contributes to code fixes, owns Dynatrace based observability, and delivers custom reliability and operational reporting to improve system health and availability. Participation in a scheduled-on call rotation is required.\nMinimum Requirement:\n4-year Computer Science, Information Systems, Engineering degree or relevant experience. 7+ Years of Site reliability experience. 8+ Years of overall experience.\nKey Responsibilities:\nDesign, implement, and maintain monitoring solutions to ensure system health and performance.\nDevelop and manage CI/CD pipelines using GitHub Actions.\nDeploy, manage, and troubleshoot containerized applications using Docker and Kubernetes.\nSupport and optimize Java-based applications in production environments.\nCollaborate with development teams to improve system reliability and reduce operational toil.\nImplement best practices for incident response, capacity planning, and disaster recovery.\nProvision and manage infrastructure using Azure cloud services.\nImprove system observability using tools such as Dynatrace (preferred).\nPerform Linux and Windows system administration, including patching, configuration, and troubleshooting.\nAutomate operational tasks using Ansible, Python, Bash, or similar tools.\nAdvanced SRE Leadership Responsibilities:\nProvide technical leadership for SRE practices across multiple services or platforms.\nDefine and evolve reliability standards, operational best practices, and incident response frameworks.\nInfluence system architecture and design decisions to ensure scalability, resilience, and operability.\nServe as a subject matter expert for reliability, availability, and production risk management.\nAct as the lead escalation point for complex and business critical production incidents.\nLead high severity incident response, coordinating across engineering, platform, and security teams.\nDrive blameless post incident reviews and ensure corrective actions are prioritized and completed.\nImprove call processes, escalation models, and incident response effectiveness.\nOwn the strategy and implementation of Dynatrace based observability, including dashboards and alerting standards.\nEstablish and monitor reliability signals (availability, latency, error rates) across critical systems.\nIdentify reliability risks and lead mitigation initiatives before customer impact occurs.\nDefine and maintain leadership level reliability and operational reporting.\nUse production data to drive prioritization of reliability investments and operational improvements.\nCommunicate reliability posture, risks, and recommendations to senior engineering leadership.\nMentor and guide senior and midlevel SREs and production support engineers.\nSupport hiring, onboarding, and technical evaluation of SRE talent.\nCollaborate with squad members to define iteration plans and commitments.\nEnsure compliance with HIPAA and other security regulations.\nCritical Skills:\nStrong experience with monitoring and observability tools (Dynatrace experience is a plus).\nHands-on experience with GitHub Actions for CI/CD automation.\nProficiency in Kubernetes and Docker for container orchestration.\nFamiliarity with Azure cloud services.\nExperience with Ansible.\nDemonstrated experience in automation of infrastructure and operational processes using scripting or configuration management tools.\nExperience supporting Java applications in production.\nSolid understanding of Linux and Windows system administration.\nKnowledge of SRE principles (SLIs, SLOs, error budgets).\nAdditional Skills:\nExperience working with onsite and offshore teams.\nStrong communication skills (written and verbal).\nStrong organizational skills, attention to detail, and ability to multitask.\nExperience in healthcare software or compliance solutions is a plus.\nStrong analytical and problem-solving skills with the ability to identify root causes and propose effective solutions.\nFor applications and inquiries, contact: hirings@openkyber.com","datePosted":"2026-08-06T13:34:32.906Z","dateModified":"2026-08-06T13:34:32.906Z","hiringOrganization":{"@type":"Organization","name":"Openkyber","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Texas","addressRegion":"PA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"1025f122489f7b58856d8d8b"},"url":"https://jobsearcher.com/jobs/1025f122489f7b58856d8d8b"}}