{"schemaVersion":"jobsearcher.job.v1","id":"08a9411a6fc37b634793f1e0","url":"https://jobsearcher.com/jobs/08a9411a6fc37b634793f1e0","canonicalUrl":"https://jobsearcher.com/jobs/08a9411a6fc37b634793f1e0","title":"Senior Site Reliability / DevOps Engineer (Cloud & Platform Reliability)","description":"Location: Austin, TX (Hybrid – 2 days onsite, 3 days remote)\r\nDuration: May 2026 – August 2026 (Extension Possible)\r\nSchedule: Monday–Friday | 8:00 AM – 5:00 PM CST\r\nHours: Up to 780 hours\r\nWork Authorization: U.S.-based candidates only\r\nLocal Candidates Only: Must reside within 50 miles of Austin, TX\r\nOverview We are seeking a highly experienced Site Reliability / DevOps Engineer to support enterprise production systems and cloud infrastructure.\r\nThis role focuses on ensuring system reliability, scalability, and performance by applying software engineering principles to infrastructure and operations. The ideal candidate will partner with development teams to build resilient, observable, and automated platforms aligned with service level objectives (SLOs).\r\nKey Responsibilities Platform Reliability & Engineering Design, build, and maintain highly available, scalable distributed systems\r\nEnsure system reliability, performance, and uptime across production environments\r\nDefine and manage SLIs, SLOs, and error budgets\r\nInfrastructure & Cloud Operations Manage and optimize cloud environments (AWS or GCP)\r\nImplement infrastructure automation and configuration management\r\nSupport containerized environments using Docker and Kubernetes\r\nMonitoring, Observability & Incident Management Implement monitoring, logging, and alerting solutions\r\nPerform incident response, root cause analysis (RCA), and postmortems\r\nDevelop and maintain dashboards, runbooks, and operational standards\r\nDevOps & Automation Develop scripts and tools using languages such as Python, Go, Java, or Bash\r\nEnable CI/CD pipelines and improve deployment reliability\r\nSupport progressive delivery practices (canary releases, feature flags)\r\nSecurity & Compliance Integrate security best practices into operational workflows\r\nEnsure compliance and reliability standards are maintained across systems\r\nRequired Qualifications 8+ years of experience in Site Reliability Engineering, DevOps, or Systems Engineering\r\nStrong experience with Linux/Unix systems and system internals\r\nProficiency in at least one programming/scripting language ( Python, Go, Java, or Bash )\r\nExperience designing and operating distributed, highly available systems\r\nHands-on experience with cloud platforms (AWS or GCP)\r\nExperience with Docker and Kubernetes\r\nStrong understanding of monitoring, logging, and alerting systems\r\nExperience with SLIs, SLOs, and error budgets\r\nProven experience in incident management and root cause analysis\r\nPreferred Qualifications Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk, or Application Insights\r\nExperience supporting 24/7 production environments and on-call rotations\r\nFamiliarity with chaos engineering and resiliency testing\r\nExperience with canary deployments and progressive delivery strategies\r\nHybrid role with mandatory onsite days (Monday & Thursday)\r\nOccasional after-hours or weekend support may be required\r\nAll travel or relocation expenses are the responsibility of the candidate\r\nJ-18808-Ljbffr","company":"Crowdplat","rawCompany":"crowdplat","city":"Austin","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-04-09T09:08:27.948Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Site Reliability / DevOps Engineer (Cloud & Platform Reliability)","description":"Location: Austin, TX (Hybrid – 2 days onsite, 3 days remote)\r\nDuration: May 2026 – August 2026 (Extension Possible)\r\nSchedule: Monday–Friday | 8:00 AM – 5:00 PM CST\r\nHours: Up to 780 hours\r\nWork Authorization: U.S.-based candidates only\r\nLocal Candidates Only: Must reside within 50 miles of Austin, TX\r\nOverview We are seeking a highly experienced Site Reliability / DevOps Engineer to support enterprise production systems and cloud infrastructure.\r\nThis role focuses on ensuring system reliability, scalability, and performance by applying software engineering principles to infrastructure and operations. The ideal candidate will partner with development teams to build resilient, observable, and automated platforms aligned with service level objectives (SLOs).\r\nKey Responsibilities Platform Reliability & Engineering Design, build, and maintain highly available, scalable distributed systems\r\nEnsure system reliability, performance, and uptime across production environments\r\nDefine and manage SLIs, SLOs, and error budgets\r\nInfrastructure & Cloud Operations Manage and optimize cloud environments (AWS or GCP)\r\nImplement infrastructure automation and configuration management\r\nSupport containerized environments using Docker and Kubernetes\r\nMonitoring, Observability & Incident Management Implement monitoring, logging, and alerting solutions\r\nPerform incident response, root cause analysis (RCA), and postmortems\r\nDevelop and maintain dashboards, runbooks, and operational standards\r\nDevOps & Automation Develop scripts and tools using languages such as Python, Go, Java, or Bash\r\nEnable CI/CD pipelines and improve deployment reliability\r\nSupport progressive delivery practices (canary releases, feature flags)\r\nSecurity & Compliance Integrate security best practices into operational workflows\r\nEnsure compliance and reliability standards are maintained across systems\r\nRequired Qualifications 8+ years of experience in Site Reliability Engineering, DevOps, or Systems Engineering\r\nStrong experience with Linux/Unix systems and system internals\r\nProficiency in at least one programming/scripting language ( Python, Go, Java, or Bash )\r\nExperience designing and operating distributed, highly available systems\r\nHands-on experience with cloud platforms (AWS or GCP)\r\nExperience with Docker and Kubernetes\r\nStrong understanding of monitoring, logging, and alerting systems\r\nExperience with SLIs, SLOs, and error budgets\r\nProven experience in incident management and root cause analysis\r\nPreferred Qualifications Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk, or Application Insights\r\nExperience supporting 24/7 production environments and on-call rotations\r\nFamiliarity with chaos engineering and resiliency testing\r\nExperience with canary deployments and progressive delivery strategies\r\nHybrid role with mandatory onsite days (Monday & Thursday)\r\nOccasional after-hours or weekend support may be required\r\nAll travel or relocation expenses are the responsibility of the candidate\r\nJ-18808-Ljbffr","datePosted":"2026-04-09T09:08:27.948Z","dateModified":"2026-04-09T09:08:27.948Z","hiringOrganization":{"@type":"Organization","name":"Crowdplat","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Austin","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"08a9411a6fc37b634793f1e0"},"url":"https://jobsearcher.com/jobs/08a9411a6fc37b634793f1e0"}}