Lead Cloud Platform & Infrastructure Engineer
Core Platform Engineer (L1, breadth-first)Sunnyvale, CA or San Jose, CA (5 days Onsite, Final round F2F)Contract & FTE (Both)Job Description:-The engineering first line of defense: incident response, triage, reliability, and automation across the full infrastructure stack.Day-to-day:Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence.Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies.Participates in team stand-ups on projects, incidents, and daily priorities.Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring.Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment.Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset.Must have:Architecture, design patterns, reliability, and scaling of new and existing systems.Incident command experience — driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through.Observability built from the ground up — defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do.Linux kernel internals — scheduler, memory allocation, driver subsystems.High-quality code in at least one language (Python, Go, or similar).System-level debugging — kdump, kernel panic analysis.IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.TCP/IP and network programming.Distributed storage systems — object, block, and/or file storage paradigms.Strong communication skills.Nice to have:Hardware and GPU troubleshooting.OVN/OVS-based networking stack exposure.Direct: 469-421-5604 , Ext- 218 • nitesh.j@tekgence.comLinkedin: linkedin.com/in/nitesh-ch-a378b52226655 Deseo Dr, Suite 104,Irving, TX , 75039 • www.tekgence.com