{"schemaVersion":"jobsearcher.job.v1","id":"3cb40d967640a08d6b2c8ec4","url":"https://jobsearcher.com/jobs/3cb40d967640a08d6b2c8ec4","canonicalUrl":"https://jobsearcher.com/jobs/3cb40d967640a08d6b2c8ec4","title":"Senior Technical Lead - DevOps","description":"- Provide consulting services for improved system stability, availability, performance and reliability. - Assist in determining the impact of operational issues and provide input into their resolution via data extraction and quantification. - Work through day-to-day support issues, ensure effective and timely resolution of issues in production environment, troubleshoot customer impacting issues. - Support multiple applications, specifically running Kubernetes, Gloo, AWS, Apigee, PCF, GCP/Java based systems in an enterprise environment. - Support Gloo running on Kubernetes, Apigee opdk and saas, Grafana, Prometheus, Cassandra, Postgres, Spring Boot or Java based applications running on Kubernetes, PCF, and Java application servers. - Apply GitOps principles to manage infrastructure and application configurations. - Apply monitoring and create complex alerts and dashboards for production systems. - Provide capacity analysis and tuning analysis for Apigee and Java applications hosted on LINUX and container platform. - Available to provide 24X7 on-call support on a rotating basis with other team members. - Lead efforts in troubleshooting, recovery, and root cause investigation. - Perform analysis of user requirements and problems to automate or improve systems and review system capabilities, workflow, and scheduling limitations. - Able to follow and develop detailed work plans, schedules, project estimates, resource plans, and status reports. - Facilitate HA (High Availability) / DR (Disaster Recovery) exercises to ensure that the team is fully prepared for any event. - Lead root cause analysis sessions to understand what causes issues in Production and come up with RCA Report along with solutions that will prevent them from happening in the future. - Ensure documentation is created and remains updated for any related work. - Strong understanding of UNIX operating systems and any scripting language. - Forecast and plan for a rapidly growing environment. - Evaluate new software product and service solutions. Skill Requirements: Expertise in analyzing and troubleshooting large-scale distributed systems. Strong experience with Kubernetes – Container Orchestration Tool, Gloo, AWS, Apigee API Gateway. Experience with REST, SOAP, and GraphQL API support. Experience with tools like: Git, Gitlab, Docker, Postman, Splunk, App Dynamics, Imperva WAF and CI/CD tools. Good experience in GitOps process, performance measurement tuning, capacity planning and management, contingency, and disaster recovery. Good understanding and strong experience with Unix/Linux operating systems. Ability to debug, optimize code, and automate routine tasks. Systematic problem-solving approach coupled with effective communication skills. Strong scripting knowledge and experience. Good understanding of networking, routing, and TLS/SSL. #J-18808-Ljbffr","company":"TechDigital Group","rawCompany":"techdigital group","city":"Seattle","state":"WA","isRemote":false,"isActive":false,"createdAt":"2026-07-04T00:13:00.590Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Technical Lead - DevOps","description":"- Provide consulting services for improved system stability, availability, performance and reliability. - Assist in determining the impact of operational issues and provide input into their resolution via data extraction and quantification. - Work through day-to-day support issues, ensure effective and timely resolution of issues in production environment, troubleshoot customer impacting issues. - Support multiple applications, specifically running Kubernetes, Gloo, AWS, Apigee, PCF, GCP/Java based systems in an enterprise environment. - Support Gloo running on Kubernetes, Apigee opdk and saas, Grafana, Prometheus, Cassandra, Postgres, Spring Boot or Java based applications running on Kubernetes, PCF, and Java application servers. - Apply GitOps principles to manage infrastructure and application configurations. - Apply monitoring and create complex alerts and dashboards for production systems. - Provide capacity analysis and tuning analysis for Apigee and Java applications hosted on LINUX and container platform. - Available to provide 24X7 on-call support on a rotating basis with other team members. - Lead efforts in troubleshooting, recovery, and root cause investigation. - Perform analysis of user requirements and problems to automate or improve systems and review system capabilities, workflow, and scheduling limitations. - Able to follow and develop detailed work plans, schedules, project estimates, resource plans, and status reports. - Facilitate HA (High Availability) / DR (Disaster Recovery) exercises to ensure that the team is fully prepared for any event. - Lead root cause analysis sessions to understand what causes issues in Production and come up with RCA Report along with solutions that will prevent them from happening in the future. - Ensure documentation is created and remains updated for any related work. - Strong understanding of UNIX operating systems and any scripting language. - Forecast and plan for a rapidly growing environment. - Evaluate new software product and service solutions. Skill Requirements: Expertise in analyzing and troubleshooting large-scale distributed systems. Strong experience with Kubernetes – Container Orchestration Tool, Gloo, AWS, Apigee API Gateway. Experience with REST, SOAP, and GraphQL API support. Experience with tools like: Git, Gitlab, Docker, Postman, Splunk, App Dynamics, Imperva WAF and CI/CD tools. Good experience in GitOps process, performance measurement tuning, capacity planning and management, contingency, and disaster recovery. Good understanding and strong experience with Unix/Linux operating systems. Ability to debug, optimize code, and automate routine tasks. Systematic problem-solving approach coupled with effective communication skills. Strong scripting knowledge and experience. Good understanding of networking, routing, and TLS/SSL. #J-18808-Ljbffr","datePosted":"2026-07-04T00:13:00.590Z","dateModified":"2026-07-04T00:13:00.590Z","hiringOrganization":{"@type":"Organization","name":"TechDigital Group","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Seattle","addressRegion":"WA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"3cb40d967640a08d6b2c8ec4"},"url":"https://jobsearcher.com/jobs/3cb40d967640a08d6b2c8ec4"}}