SRE - Java
Everest Global Solutions Hiring for Below Role.Candidates can apply to jyothik@everestglobalsolutionsinc.comRole: SRE - JavaLocation: Brentwood, TNEngagement Type: FulltimeRequired Technical Expertise:• 10+ years of hands-on experience in Java application development,production engineering, and Site Reliability Engineering (SRE).• Strong expertise in Core Java, Java 11/17+, Spring Boot, Spring Cloud,Microservices, JVM Internals, Garbage Collection (GC) tuning, andmultithreading.• Extensive experience with Google Cloud Platform (GCP), including GoogleKubernetes Engine (GKE), Compute Engine, Cloud Storage, Cloud SQL,Pub/Sub, IAM, VPC, Cloud Monitoring, Cloud Logging, Secret Manager, andCloud Load Balancing.• Expert-level experience administering and troubleshooting Kubernetesproduction environments, including cluster management, networking,autoscaling, RBAC, storage, Helm, Ingress Controllers, and service meshtechnologies.• Hands-on experience with Docker, container orchestration, andInfrastructure as Code (IaC) using Terraform.• Strong experience building and maintaining enterprise CI/CD pipelinesusing GitLab CI/CD, Jenkins, GitOps, and cloud-native deployment tools.• Experience implementing and supporting zero-downtime deploymentstrategies, including Blue-Green and Canary deployments.• Strong proficiency in Linux/Unix administration, shell scripting (Bash), andoperational automation using Python or Go.• Hands-on experience with Kafka, Kafka Streams, Pub/Sub, or other event-driven messaging platforms supporting high-volume distributed systems.• Experience implementing enterprise monitoring, logging, and observabilitysolutions using Prometheus, Grafana, Datadog, Splunk, OpenTelemetry,Cloud Monitoring, and Kiali.• Strong understanding of networking concepts, including TCP/IP, DNS, LoadBalancers, NGINX, Ingress Controllers, TLS/SSL, and Service Mesh (Istio).• Experience supporting highly available, scalable, and mission-criticalproduction environments with 24x7 operational responsibilities.• Strong experience performing production incident management, RCA(Root Cause Analysis), performance tuning, capacity planning, andreliability engineering.• Experience with enterprise security best practices, including IAM, SecretManager, RBAC, least-privilege access, vulnerability remediation, andcontainer security.• Experience supporting environments compliant with PCI-DSS, SOC2, SOX,ISO 27001, or similar regulatory frameworks.• Prior experience supporting large-scale retail, eCommerce, omnichannel,supply chain, order management, inventory management, or paymentprocessing platforms.• Excellent troubleshooting, analytical, communication, and stakeholdermanagement skills.• Experience with Anthos, ArgoCD, Cloud Build, Cloud Deploy, Redis,PostgreSQL, MongoDB, and Cassandra.• Experience with eBPF, distributed tracing, performance engineering, andadvanced observability.• Exposure to VMware, disaster recovery planning, chaos engineering, andmulti-region Kubernetes deployments.• Experience leading SRE initiatives, mentoring engineering teams, anddriving operational excellence in enterprise environments.Certifications Required:• Google Professional Cloud Architect OR Google Professional Cloud DevOpsEngineer• Certified Kubernetes Administrator (CKA) OR Certified Kubernetes Security