{"schemaVersion":"jobsearcher.job.v1","id":"61d792cb019d5f58cd272ee7","url":"https://jobsearcher.com/jobs/61d792cb019d5f58cd272ee7","canonicalUrl":"https://jobsearcher.com/jobs/61d792cb019d5f58cd272ee7","title":"Cloud Solutions Engineer","description":"DescriptionPlatform Operations EngineerThe Platform Operations Engineer is a technical role responsible for strengthening the reliability, scalability, performance, and operational maturity of cloud-based applications and services. This role partners closely with engineering, product support, cloud operations, and infrastructure teams to proactively identify reliability risks, reduce operational toil through automation, improve observability, and implement preventative measures that help avoid production issues before they impact customers. The Platform Operations Engineer supports SLA commitments, improves incident readiness and response, and drives continuous improvement across systems, processes, and teams to deliver a dependable, high-quality customer experience.ResponsibilitiesDesign, build, and maintain highly available, scalable, secure, and cost-effective cloud infrastructure and supporting services.Implement, maintain, and continuously improve monitoring, alerting, logging, and observability practices to detect reliability risks early and ensure alerts are actionable.Proactively assess production systems for operational risk, performance bottlenecks, resiliency gaps, and recurring failure patterns; recommend and implement preventative improvements.Define, track, and improve service reliability indicators, objectives, and error budgets in partnership with engineering, product, and operations teams to balance reliability, customer impact, and delivery priorities.Partner with engineering, product support, cloud operations, and infrastructure teams to strengthen incident response readiness, improve communication, and drive timely resolution of production issues.Lead root cause analysis and cross-functional incident retrospectives, translating findings into preventative actions that reduce recurring incidents and improve system resiliency.Develop automation, operational tooling, deployment support, and infrastructure provisioning practices to reduce manual effort, improve consistency, and minimize operational toil.Conduct regular performance tuning, capacity analysis, reliability reviews, and system health assessments to support SLA commitments and customer expectations.Drive continuous improvement initiatives that improve reliability engineering practices, operational maturity, and productivity across teams.Participate in a structured on-call rotation during primary business hours and contribute to incident management best practices.Stay current with industry trends, best practices, and emerging technologies in site reliability engineering, DevOps, cloud operations, automation, and observability.QualificationsBS/BA degree in Computer Science, Computer Engineering, Information Systems, or a related field, or equivalent practical experience.3+ years of experience in Site Reliability Engineering, Cloud Operations, DevOps, Software Engineering, or a related technical role supporting cloud-based production systems.5+ years of hands-on software engineering experience, including application design, code review, testing, debugging, release processes, and supporting production software systems.Proficiency with cloud computing platforms such as AWS, Azure, or GCP, including experience with infrastructure as code tools such as Terraform.Hands-on experience with monitoring, logging, alerting, and observability tools such as Datadog, AWS CloudWatch, SolarWinds Database Performance Analyzer, or similar platforms.Experience participating in or improving incident response processes, including use of tools such as PagerDuty, JSM Operations, or similar incident management platforms.Experience conducting root cause analysis, identifying systemic risks, and implementing preventative measures to reduce production incidents.Familiarity with SRE practices such as service level indicators, service level objectives, error budgets, toil reduction, blameless post-incident reviews, and reliability-focused engineering.Expertise in scripting or programming languages such as PowerShell, Python, Bash, C#, Go, or .NET for automation and tooling.Strong understanding of database concepts and administration, including T-SQL scripting, indexing, query performance, and performance tuning for MS SQL Server and/or PostgreSQL.Solid understanding of networking concepts, security best practices, and system administration in Windows and Linux environments.Experience with CI/CD practices, deployment automation, configuration management, and modern DevOps methodologies.Strong analytical and problem-solving skills, with the ability to troubleshoot complex issues and drive timely, sustainable resolution.Clear communicator who thrives in collaborative, cross-functional environments and can explain technical risks, tradeoffs, and recommendations to varied audiences.Demonstrated curiosity, ownership, and a continuous improvement mindset focused on reliability, prevention, and operational excellence.Required Skills and TechnologiesAWS or comparable cloud platform experienceMS SQL Server and/or PostgreSQLWindows Server OS and Linux OSPowerShell and at least one additional scripting or programming language such as Python, Bash, C#, or .NETIIS and web application hosting conceptsMonitoring, alerting, observability, and incident management practicesInfrastructure as Code and automation practicesPreferred Skills and TechnologiesKnowledge of .NET Framework and languages such as C# or VB.NETExperience with monitoring and observability tools such as Datadog, AWS CloudWatch, and SolarWinds Database Performance AnalyzerExperience with PagerDuty, JSM Operations, or similar incident management platformsKnowledge of Agile development practices and tools such as JIRAKnowledge of web development practices and technologiesExperience with ticketing systems such as Microsoft CRMAdvanced knowledge of AWS hosting technologiesAdvanced knowledge of DevOps practices, release automation, and operational readinessExperience with security and endpoint protection tools such as CrowdStrike Anti-Virus","company":"Tyler Technologies","rawCompany":"tyler technologies","city":"Plano","state":"TX","isRemote":false,"isActive":true,"createdAt":"2026-07-29T02:36:57.034Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Cloud Solutions Engineer","description":"DescriptionPlatform Operations EngineerThe Platform Operations Engineer is a technical role responsible for strengthening the reliability, scalability, performance, and operational maturity of cloud-based applications and services. This role partners closely with engineering, product support, cloud operations, and infrastructure teams to proactively identify reliability risks, reduce operational toil through automation, improve observability, and implement preventative measures that help avoid production issues before they impact customers. The Platform Operations Engineer supports SLA commitments, improves incident readiness and response, and drives continuous improvement across systems, processes, and teams to deliver a dependable, high-quality customer experience.ResponsibilitiesDesign, build, and maintain highly available, scalable, secure, and cost-effective cloud infrastructure and supporting services.Implement, maintain, and continuously improve monitoring, alerting, logging, and observability practices to detect reliability risks early and ensure alerts are actionable.Proactively assess production systems for operational risk, performance bottlenecks, resiliency gaps, and recurring failure patterns; recommend and implement preventative improvements.Define, track, and improve service reliability indicators, objectives, and error budgets in partnership with engineering, product, and operations teams to balance reliability, customer impact, and delivery priorities.Partner with engineering, product support, cloud operations, and infrastructure teams to strengthen incident response readiness, improve communication, and drive timely resolution of production issues.Lead root cause analysis and cross-functional incident retrospectives, translating findings into preventative actions that reduce recurring incidents and improve system resiliency.Develop automation, operational tooling, deployment support, and infrastructure provisioning practices to reduce manual effort, improve consistency, and minimize operational toil.Conduct regular performance tuning, capacity analysis, reliability reviews, and system health assessments to support SLA commitments and customer expectations.Drive continuous improvement initiatives that improve reliability engineering practices, operational maturity, and productivity across teams.Participate in a structured on-call rotation during primary business hours and contribute to incident management best practices.Stay current with industry trends, best practices, and emerging technologies in site reliability engineering, DevOps, cloud operations, automation, and observability.QualificationsBS/BA degree in Computer Science, Computer Engineering, Information Systems, or a related field, or equivalent practical experience.3+ years of experience in Site Reliability Engineering, Cloud Operations, DevOps, Software Engineering, or a related technical role supporting cloud-based production systems.5+ years of hands-on software engineering experience, including application design, code review, testing, debugging, release processes, and supporting production software systems.Proficiency with cloud computing platforms such as AWS, Azure, or GCP, including experience with infrastructure as code tools such as Terraform.Hands-on experience with monitoring, logging, alerting, and observability tools such as Datadog, AWS CloudWatch, SolarWinds Database Performance Analyzer, or similar platforms.Experience participating in or improving incident response processes, including use of tools such as PagerDuty, JSM Operations, or similar incident management platforms.Experience conducting root cause analysis, identifying systemic risks, and implementing preventative measures to reduce production incidents.Familiarity with SRE practices such as service level indicators, service level objectives, error budgets, toil reduction, blameless post-incident reviews, and reliability-focused engineering.Expertise in scripting or programming languages such as PowerShell, Python, Bash, C#, Go, or .NET for automation and tooling.Strong understanding of database concepts and administration, including T-SQL scripting, indexing, query performance, and performance tuning for MS SQL Server and/or PostgreSQL.Solid understanding of networking concepts, security best practices, and system administration in Windows and Linux environments.Experience with CI/CD practices, deployment automation, configuration management, and modern DevOps methodologies.Strong analytical and problem-solving skills, with the ability to troubleshoot complex issues and drive timely, sustainable resolution.Clear communicator who thrives in collaborative, cross-functional environments and can explain technical risks, tradeoffs, and recommendations to varied audiences.Demonstrated curiosity, ownership, and a continuous improvement mindset focused on reliability, prevention, and operational excellence.Required Skills and TechnologiesAWS or comparable cloud platform experienceMS SQL Server and/or PostgreSQLWindows Server OS and Linux OSPowerShell and at least one additional scripting or programming language such as Python, Bash, C#, or .NETIIS and web application hosting conceptsMonitoring, alerting, observability, and incident management practicesInfrastructure as Code and automation practicesPreferred Skills and TechnologiesKnowledge of .NET Framework and languages such as C# or VB.NETExperience with monitoring and observability tools such as Datadog, AWS CloudWatch, and SolarWinds Database Performance AnalyzerExperience with PagerDuty, JSM Operations, or similar incident management platformsKnowledge of Agile development practices and tools such as JIRAKnowledge of web development practices and technologiesExperience with ticketing systems such as Microsoft CRMAdvanced knowledge of AWS hosting technologiesAdvanced knowledge of DevOps practices, release automation, and operational readinessExperience with security and endpoint protection tools such as CrowdStrike Anti-Virus","datePosted":"2026-07-29T02:36:57.034Z","dateModified":"2026-07-29T02:36:57.034Z","hiringOrganization":{"@type":"Organization","name":"Tyler Technologies","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Plano","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"61d792cb019d5f58cd272ee7"},"url":"https://jobsearcher.com/jobs/61d792cb019d5f58cd272ee7"}}