Linux Systems Engineer/ SRE (RHEL, Linux, BIOS) - PST Working Hours
Linux Systems Engineer/ SRE (RHEL, Linux, BIOS) - PST Working Hours100% RemoteFull-Time PermanentJob Description:Hands-on Linux Systems/Site Reliability Engineer responsible for maintaining the availability, reliability, performance, and operational health of production compute and API services. The role focuses on RHEL/Linux administration, production troubleshooting, incident response, automation, monitoring, and operational tooling. The engineer will diagnose infrastructure and service issues independently while improving scalability, reliability, and operational efficiency.Experience:4–8 years of relevant Linux administration, production operations, SRE, or infrastructure support experience.Hands-on experience supporting mission-critical Tier-1 production services.Experience with incident response, pager/on-call support, debugging, and root cause analysis.Core Skills:Operating Systems: RHEL 7, RHEL 8, RHEL 9, Linux, UnixSystem Administration: Linux boot process, BIOS, UEFI, systemd, systemctl, rescue mode, emergency mode, root-password recoveryStorage & Filesystems: Filesystems, disk utilization, inodes, SWAP, permanent mounts, /etc/fstab, Ext4, XFSProcesses & Networking: Memory and process analysis, zombie processes, orphan processes, listening ports, service troubleshootingProgramming & Scripting: Python, Bash, JavaScriptMonitoring & Observability: Service metrics, dashboards, KPIs, alarmsProduction Operations: Incident triage, ticket management, runbooks, root cause analysis, operational toil reductionEngineering: Services, operational tools, CI/CD, scalability, reliability, API availabilityKey Responsibilities:Maintain the operational health, availability, reliability, and low latency of core compute and API services.Administer and troubleshoot RHEL/Linux systems in production environments.Manage and triage incidents and tickets based on business and service impact.Diagnose filesystem, disk, memory, process, boot, service, and network-related issues.Build automation and operational tooling using Python, Bash, or JavaScript.Develop dashboards, service KPIs, monitoring systems, and actionable alerts.Create and automate frequently used runbooks to reduce incident triage time and operational toil.Collaborate with developers to improve system scalability, reliability, and development velocity.Participate in on-call support, incident response, debugging, and root cause analysis.