JOBSEARCHER

Senior DevOps Engineer

HcssHouston, TXL6 LeadOctober 5th, 2026
We are HCSS. For the last 40 years, we have been developing software to help construction companies streamline their operations. Based in Sugar Land, TX, our mission is helping customers achieve excellence through our proven customer-centric, end-to-end solutions and exceptionally helpful service, while providing a great life for our employees. With this mission at the core of everything we do, HCSS is a pioneer and leader in the construction software space and a consistently recognized employer. We have earned Best Companies to Work for in Texas honors for 18 consecutive years and have been named a USA Today Top Workplace. HCSS has also been recognized by Built In as a Best Place to Work in Greater Houston and by Construction Executive for our technology innovation, reflecting our strong culture, industry leadership, and commitment to excellence.Who We NeedAs a Senior DevOps Engineer with a focus on Site Reliability Engineering (SRE), you will play a key role in driving infrastructure resilience, availability, and operational excellence across our cloud environments. A core responsibility of this role is managing and optimizing Azure SQL Elastic Pools, ensuring performance, cost-efficiency, observability, and automation are aligned with business objectives. You will lead initiatives around high availability, disaster recovery, incident response, and reliability automation. Your expertise in Azure or AWS, observability tools such as Grafana, and scalable infrastructure will be essential to ensuring our systems are robust, performant, and recoverable.Qualifications8+ years of experience in DevOps or SRE roles with a strong focus on cloud infrastructure and systems reliability3+ years of hands-on experience with managing Azure SQL Elastic Pools, including performance tuning, scaling, and automation5+ years of expertise in Azure cloud services including networking, compute, databases, and identity3+ years of experience applying SRE principles including SLIs, SLOs, and incident management best practicesExtensive experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templatesExtensive experience building and managing CI/CD pipelines with tools like Azure DevOps or GitHub ActionsStrong scripting skills using Azure CLI and PowerShell for automation and operational tasksExperience with monitoring and observability platforms, ideally Grafana, or a strong foundation in similar toolsSoft SkillsStrong troubleshooting and problem-solving abilities.Excellent communication skills and a collaborative mindset to work with cross-functional teams.Ability to work independently, manage multiple tasks, and prioritize efficiently.A proactive attitude toward continuous improvement and learning.Preferred QualificationsManaged 10+ Elastic pools and 100+ databases in AzureAdvanced level certifications on cloud infrastructure like Az-400 or equivalentRole ResponsibilitiesAzure Elastic Pool Management:Take ownership of the design, scaling, and optimization of Azure SQL Elastic PoolsMonitor and tune pool performance to ensure efficiency and SLA complianceEstablish observability and alerting for SQL resource consumption, errors, and performance anomaliesAutomate provisioning, scaling, and failover using infrastructure and scripting toolsCollaborate with database and application teams to align on resource usage strategiesHigh Availability And Disaster RecoveryDesign and implement highly available and fault tolerant systemsDevelop and maintain disaster recovery strategies across critical servicesPerform regular failover testing, documentation, and validation of recovery proceduresWork closely with infrastructure and development teams to ensure business continuity objectives are metMonitoring, Observability, And AutomationImplement and manage observability stacks with logs, metrics, traces, and alertingCreate dashboards and alerts in Grafana or similar platforms to track key system indicatorsDevelop automated solutions for provisioning, monitoring, and maintenance tasksContinuously improve system visibility and reduce time to detect and resolve issuesAssist in the automation of performance testing to proactively identify bottlenecks, validate scalability and ensure reliable system behaviorIncident Management And Operational ExcellenceEstablish and refine incident response processes, escalation workflows, and resolution protocolsLead root cause analysis, post-incident reviews, and continuous improvement efforts, including following up to ensure identified improvements to the application or process are implemented.Develop and maintain runbooks, diagnostic tools, and automated remediation solutionsChampion a blameless culture of reliability and operational readiness across engineering teamsCloud Infrastructure ManagementArchitect and manage scalable and secure cloud infrastructure in Azure or AWSProvision and manage services including compute, networking, storage, and containerized workloadsContinuously monitor performance, latency, and uptime to ensure system healthApply cost optimization practices while aligning infrastructure with business goalsInfrastructure As Code (IaC)Define and implement infrastructure using tools such as TerraformMaintain modular, version-controlled infrastructure code that supports environment consistencyApply automation and policy enforcement to reduce drift and improve auditabilityEnsure IaC best practices are embedded in the development lifecycleCollaboration And LeadershipMentor and support junior DevOps engineers through code reviews, knowledge sharing, and technical guidancePartner with cross-functional teams including development, security, and database operations to drive initiativesAct as a subject matter expert in SRE practices and reliability-driven engineeringLead continuous improvement efforts across infrastructure and operations processesTravel RequirementsRemote RequirementsEmployees will be expected to come into the office on a periodic basis.Baseline expectations for roles are as follows but may fluctuate based on manager’s discretion:Individual Contributor - Up to 2x per yearEmployees will be expected to attend HCSS sponsored events per manager discretion (ex. UGM)Benefits & PerksPart of our mission is to provide a great life for our employees. We believe that when our people are happy, they do their best work. Some of the benefits and perks we offer include:Flexibility to work RemotelyMedical, dental, and vision coverage with company-paid and employee-paid optionsPaid holidays, sick days, and personal time offEmployee Resource Groups (ERGs) that foster connection and inclusionOn-site amenities including a covered basketball court, soccer field, track, pickleball/tennis courts, gym, etc.Dog-friendly campus and WiFi-accessible courtyards401(k) with a 5% company matchCoverage for employee professional development and wellnessAnd more!