JOBSEARCHER

Site Reliability Engineer - (SRE ) with Java

Key Responsibilities: • Design, build, and operate highly available, fault-tolerant distributedsystems supporting enterprise retail platforms.* Lead Site Reliability Engineering (SRE) initiatives including SLIs, SLOs, errorbudgets, capacity planning, and operational excellence.* Provide L3/L4 production support, incident management, root causeanalysis (RCA), and post-incident remediation.* Support large-scale Java/Spring Boot microservices deployed onKubernetes.* Design and manage production workloads on Google Kubernetes Engine(GKE).* Implement Infrastructure-as-Code using Terraform.* Automate infrastructure provisioning, deployment, scaling, and policyenforcement.* Design and maintain enterprise CI/CD pipelines using GitLab CI/CD,Jenkins, Cloud Build, and GitOps methodologies.* Build deployment automation for zero-downtime releases using Blue-Green and Canary deployment strategies.* Implement advanced observability using Prometheus, Grafana, Datadog,Splunk, OpenTelemetry, and Cloud Monitoring.* Develop operational automation using Python, Go, Bash, and cloud-nativetooling.* Optimize JVM performance including Garbage Collection tuning, heapanalysis, thread dumps, and application profiling.* Support Kafka-based event-driven retail systems handling inventory, orderprocessing, customer events, and payment processing.* Manage Kubernetes networking, Ingress Controllers, service mesh (Istio),storage, autoscaling, and multi-cluster deployments.* Build resilient disaster recovery and business continuity strategies acrossmultiple GCP regions.* Enforce enterprise security controls including IAM, Secret Manager,Workload Identity, Binary Authorization, and least-privilege access.* Support compliance initiatives including PCI-DSS, SOC2, ISO 27001, GDPR,and internal security standards.* Participate in production support rotation, weekend deployments,maintenance windows, and on-call responsibilities.Required TechnicalExpertise:* 10+ years of hands-on experience in Java application development,production engineering, and Site Reliability Engineering (SRE).* Strong expertise in Core Java, Java 11/17+, Spring Boot, Spring Cloud,Microservices, JVM Internals, Garbage Collection (GC) tuning, andmultithreading.* Extensive experience with Google Cloud Platform (GCP), including GoogleKubernetes Engine (GKE), Compute Engine, Cloud Storage, Cloud SQL, Pub/Sub, IAM, VPC, Cloud Monitoring, Cloud Logging, Secret Manager, andCloud Load Balancing.* Expert-level experience administering and troubleshooting Kubernetesproduction environments, including cluster management, networking,autoscaling, RBAC, storage, Helm, Ingress Controllers, and service meshtechnologies.* Hands-on experience with Docker, container orchestration, andInfrastructure as Code (IaC) using Terraform.* Strong experience building and maintaining enterprise CI/CD pipelinesusing GitLab CI/CD, Jenkins, GitOps, and cloud-native deployment tools.* Experience implementing and supporting zero-downtime deploymentstrategies, including Blue-Green and Canary deployments.* Strong proficiency in Linux/Unix administration, shell scripting (Bash), andoperational automation using Python or Go.* Hands-on experience with Kafka, Kafka Streams, Pub/Sub, or other event-driven messaging platforms supporting high-volume distributed systems.* Experience implementing enterprise monitoring, logging, and observabilitysolutions using Prometheus, Grafana, Datadog, Splunk, OpenTelemetry,Cloud Monitoring, and Kiali.* Strong understanding of networking concepts, including TCP/IP, DNS, LoadBalancers, NGINX, Ingress Controllers, TLS/SSL, and Service Mesh (Istio).* Experience supporting highly available, scalable, and mission-criticalproduction environments with 24x7 operational responsibilities.* Strong experience performing production incident management, RCA(Root Cause Analysis), performance tuning, capacity planning, andreliability engineering.* Experience with enterprise security best practices, including IAM, SecretManager, RBAC, least-privilege access, vulnerability remediation, andcontainer security.* Experience supporting environments compliant with PCI-DSS, SOC2, SOX,ISO 27001, or similar regulatory frameworks.* Prior experience supporting large-scale retail, eCommerce, omnichannel,supply chain, order management, inventory management, or paymentprocessing platforms.* Excellent troubleshooting, analytical, communication, and stakeholdermanagement skills.* Experience with Anthos, ArgoCD, Cloud Build, Cloud Deploy, Redis,PostgreSQL, MongoDB, and Cassandra.* Experience with eBPF, distributed tracing, performance engineering, andadvanced observability.* Exposure to VMware, disaster recovery planning, chaos engineering, andmulti-region Kubernetes deployments.* Experience leading SRE initiatives, mentoring engineering teams, anddriving operational excellence in enterprise environments.CertificationsRequired:* Google Professional Cloud Architect OR Google Professional Cloud DevOpsEngineer* Certified Kubernetes Administrator (CKA) OR Certified Kubernetes Security