{"schemaVersion":"jobsearcher.job.v1","id":"d672d6eafefe0db1e29c70fc","url":"https://jobsearcher.com/jobs/d672d6eafefe0db1e29c70fc","canonicalUrl":"https://jobsearcher.com/jobs/d672d6eafefe0db1e29c70fc","title":"Site Reliability Engineer IV","description":"Candescent is the leading cloud-based digital banking solutions provider for financial institutions. We are transforming digital banking with intelligent, cloud-powered solutions that connect account opening, digital banking, and branch experiences for financial institutions. Our advanced technology and developer tools enable seamless, differentiated customer journeys that elevate trust, service, and innovation. Success here requires flexibility in a fast-paced environment, a client-first mindset, and a commitment to delivering consistent, reliable results as part of a performance-driven, values-led team. With team members around the world, Candescent is an equal opportunity employer.\n\nPosition: Site Reliability Engineer IV\n\nExperience: 9-12 Years\n\nLocation: Bangalore (Ecospace)\n\nCandescent Site Reliability Engineering (SRE) mission is to proactively ensure the reliability, availability and performance of our Digital First banking applications. As a member of the SRE team, you will focus on building and operating highly reliable application platforms by applying SRE principles such as automation, observability, resilience and continuous improvement.\n\nYou will partner closely with application and platform teams to define reliability standards, implement monitoring, alerting and incident response practices and embed scalability and performance considerations into application design and delivery. Through tooling, automation, and best practices, you will help development teams build and operate services that meet agreed reliability objectives.\n\nAs a senior engineer in the organization, you will also provide mentorship within the SRE team and across peer engineering teams, helping elevate operational maturity, drive adoption of SRE practices, and strengthen reliability culture across our core initiatives.\n\nResponsibilities\n\nSupport and operate production applications running on Kubernetes and AWS\n\nTroubleshoot application-level issues using logs, metrics, traces, and runtime signals\n\nParticipate in incident response, root cause analysis, and post-incident reviews\n\nWork closely with development teams to understand application architecture, dependencies, and data flows\n\nImprove application observability by defining meaningful alerts, dashboards, and SLOs\n\nAutomate repetitive operational tasks to reduce toil\n\nSupport application deployments, rollbacks, and runtime configuration changes\n\nIdentify reliability, performance, and scalability gaps in application behavior\n\nDrive continuous improvements in operational readiness, runbooks, and on-call practices\n\nInfluence application teams to adopt shift-left reliability practices\n\nMust-Have Skills & Experience\n\nHands-on experience supporting Java applications in production\n\nStrong understanding of JVM fundamentals (heap/memory management, garbage collection, OOM issues, thread analysis)\n\nProven experience with SRE practices, including:\n\nIncident response and on-call support\n\nRoot cause analysis and postmortems\n\nSLIs, SLOs, and reliability-driven operations\n\nStrong experience troubleshooting using application logs, metrics, and monitoring tools\n\nExperience operating Java applications on Kubernetes (EKS) from an application/runtime perspective\n\nExperience with deployment strategies (rolling, blue/green, canary)\n\nAbility to write automation and scripts (Python or any) to reduce operational toil\n\nSolid understanding of application architecture and service dependencies (databases, messaging systems, external APIs)\n\nStrong collaboration and communication skills; ability to work closely with development teams\n\nDemonstrates accountability and sound judgment when responding to high-pressure incidents\n\nGood-to-Have Skills & Experience\n\nExposure to platform or infrastructure concepts supporting application workloads\n\nExperience with AWS services such as EKS, RDS/Aurora, S3, EFS, and CloudWatch\n\nCI/CD pipeline experience (GitHub Actions, Jenkins)\n\nFamiliarity with GitOps practices\n\nExperience with cloud migrations or modernization efforts\n\n#J-18808-Ljbffr","company":"Capitolis","rawCompany":"capitolis","city":"Sterling","state":"VA","isRemote":false,"isActive":false,"createdAt":"2026-06-17T04:16:56.806Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer IV","description":"Candescent is the leading cloud-based digital banking solutions provider for financial institutions. We are transforming digital banking with intelligent, cloud-powered solutions that connect account opening, digital banking, and branch experiences for financial institutions. Our advanced technology and developer tools enable seamless, differentiated customer journeys that elevate trust, service, and innovation. Success here requires flexibility in a fast-paced environment, a client-first mindset, and a commitment to delivering consistent, reliable results as part of a performance-driven, values-led team. With team members around the world, Candescent is an equal opportunity employer.\n\nPosition: Site Reliability Engineer IV\n\nExperience: 9-12 Years\n\nLocation: Bangalore (Ecospace)\n\nCandescent Site Reliability Engineering (SRE) mission is to proactively ensure the reliability, availability and performance of our Digital First banking applications. As a member of the SRE team, you will focus on building and operating highly reliable application platforms by applying SRE principles such as automation, observability, resilience and continuous improvement.\n\nYou will partner closely with application and platform teams to define reliability standards, implement monitoring, alerting and incident response practices and embed scalability and performance considerations into application design and delivery. Through tooling, automation, and best practices, you will help development teams build and operate services that meet agreed reliability objectives.\n\nAs a senior engineer in the organization, you will also provide mentorship within the SRE team and across peer engineering teams, helping elevate operational maturity, drive adoption of SRE practices, and strengthen reliability culture across our core initiatives.\n\nResponsibilities\n\nSupport and operate production applications running on Kubernetes and AWS\n\nTroubleshoot application-level issues using logs, metrics, traces, and runtime signals\n\nParticipate in incident response, root cause analysis, and post-incident reviews\n\nWork closely with development teams to understand application architecture, dependencies, and data flows\n\nImprove application observability by defining meaningful alerts, dashboards, and SLOs\n\nAutomate repetitive operational tasks to reduce toil\n\nSupport application deployments, rollbacks, and runtime configuration changes\n\nIdentify reliability, performance, and scalability gaps in application behavior\n\nDrive continuous improvements in operational readiness, runbooks, and on-call practices\n\nInfluence application teams to adopt shift-left reliability practices\n\nMust-Have Skills & Experience\n\nHands-on experience supporting Java applications in production\n\nStrong understanding of JVM fundamentals (heap/memory management, garbage collection, OOM issues, thread analysis)\n\nProven experience with SRE practices, including:\n\nIncident response and on-call support\n\nRoot cause analysis and postmortems\n\nSLIs, SLOs, and reliability-driven operations\n\nStrong experience troubleshooting using application logs, metrics, and monitoring tools\n\nExperience operating Java applications on Kubernetes (EKS) from an application/runtime perspective\n\nExperience with deployment strategies (rolling, blue/green, canary)\n\nAbility to write automation and scripts (Python or any) to reduce operational toil\n\nSolid understanding of application architecture and service dependencies (databases, messaging systems, external APIs)\n\nStrong collaboration and communication skills; ability to work closely with development teams\n\nDemonstrates accountability and sound judgment when responding to high-pressure incidents\n\nGood-to-Have Skills & Experience\n\nExposure to platform or infrastructure concepts supporting application workloads\n\nExperience with AWS services such as EKS, RDS/Aurora, S3, EFS, and CloudWatch\n\nCI/CD pipeline experience (GitHub Actions, Jenkins)\n\nFamiliarity with GitOps practices\n\nExperience with cloud migrations or modernization efforts\n\n#J-18808-Ljbffr","datePosted":"2026-06-17T04:16:56.806Z","dateModified":"2026-06-17T04:16:56.806Z","hiringOrganization":{"@type":"Organization","name":"Capitolis","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sterling","addressRegion":"VA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"d672d6eafefe0db1e29c70fc"},"url":"https://jobsearcher.com/jobs/d672d6eafefe0db1e29c70fc"}}