Site Reliability Engineer
Cloud Native Lab AdministratorSeeking a candidate who can administer a cloud native lab environment focusing on private cloud systems supporting 5G wireless systems. This position will focus on platform monitoring, logging, and reliability aspects supporting the Mobile Core team. A critical goal is to gather metrics of the platform during stress and load events to ensure the cloud native environment is operating as expected and how to improve upon it over time. Assist in troubleshooting systems and network level issues preventing onboarding of third party software into a Kubernetes environment. Collaborate with a highly skilled self-sufficient team to make 5G concepts a reality.Required Experience:Network Fundamentals - CCNA Certification (or equivalent knowledge)Ability to document (text, visual, and ladder diagrams) (LucidChart, Confluence, JIRA, etc)Working knowledge of Kubernetes solutions (Plan, Deploy, Operate, Scale, and Monitor)Experience deploying to and orchestrating containers (Docker, Kubernetes, etc.)Experience with automation and creating Ansible playbooksWorked with common infrastructure tools like Docker, Terraform, HelmOptional Experience:Have a strong understanding of Data-streaming-systems (e.g. Fluentd, Kafka), monitoring-tools (e.g. Splunk, ELK, Prometheus, Datadog)Should have working knowledge on Prometheus, Grafana monitoring tools and developing custom Grafana dashboards and developing Prometheus queries to analyze dataAbility to converse with application owners, architects, performance testers to pinpoint application performance bottlenecks via the monitoring & observability toolsScripting experience a plus