{"schemaVersion":"jobsearcher.job.v1","id":"a2fc9250dbf2d734bd231ef2","url":"https://jobsearcher.com/jobs/a2fc9250dbf2d734bd231ef2","canonicalUrl":"https://jobsearcher.com/jobs/a2fc9250dbf2d734bd231ef2","title":"Cloud Systems Engineer - Site Reliability","description":"About UsTherapyNotes is the go-to superhero for behavioral health Practice Management and EHR software! Our top-notch SaaS solution handles scheduling, billing, documenting, telehealth, and more so clinicians can focus on awesome patient care.We're a dynamic team of pros who love to innovate and push the envelope, keeping our software cutting-edge. Join us, and let's revolutionize behavioral health software together while making a real difference!About The PositionWe are seeking a Site Reliability Engineer to improve the reliability and operability of the production services and shared platforms supporting our growing 24×7 SaaS environment. In this role, you will apply software and systems engineering practices to improve availability, performance, scalability, resilience, observability, incident response, and operational automation. You will partner with software development, infrastructure, database, security, and other technology teams to establish measurable reliability goals, reduce operational toil, and ensure services are supportable throughout their lifecycle. If you are passionate about building reliable systems, solving complex production problems, and driving continuous improvement, we want to hear from you.RequirementsBS degree in Information Systems, Engineering, or equivalent experience5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SREExperience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferredStrong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in productionExpertise with an observability platform; Datadog experience strongly preferred. Experience with Prometheus, Grafana, New Relic, or equivalent platforms is also valuableExperience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practicesExperience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvementExperience working in Agile/DevOps environments and operating production services using ITSM practices where applicablePrior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plusResponsibilitiesOwn and continuously improve how we use Datadog to make reliability visible and actionable across metrics, logs, traces, dashboards, monitors, alerts, and service-level viewsDesign, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24×7 SaaS platformPartner with service owners to define and improve reliability through meaningful SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practicesParticipate in and help drive incident management for production events, serving as an incident commander or technical responder as needed. Coordinate triage, service restoration, escalation, communication, incident documentation, root cause analysis, and completion of corrective actionsPartner with development teams to investigate issues across the infrastructure and application layers using metrics, logs, distributed traces, and code-level contextImprove deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and analysis of system failure modesPartner with other technical leaders to ensure all newly introduced systems are supportable and maintainable by both development and operationsProvide escalated technical guidance and support to other technology teams throughout the organizationProvide on-call coverage for production support and other duties as requiredEnsure supported systems and operational activities comply with organizational security, HIPAA, and operating policiesIdentify and eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible. Manage infrastructure as code using Terraform/OpenTofu and configuration automation using AnsibleBenefitsCompetitive salary - $110,000-$150,000Employer sponsored health, dental, vision, life, and disability insuranceRetirement plan with company contributionAnnual company profit sharingPersonal development/training budgetOpen, collaborative work environmentExtensive 2-week onboarding planComprehensive mentorship programEqual Opportunity Employer Statement & Applicant RightsTherapyNotes LLC is an Equal Opportunity Employer and does not discriminate based on race, color, religion, sex, national origin, age, disability, genetic information, or any other protected status under federal, state, or local law. We are committed to providing a workplace free of discrimination and harassment.For more information about your rights under federal employment laws, please review the following{{:}}Know Your Rights{{:}} Workplace Discrimination is IllegalFamily and Medical Leave Act (FMLA){{:}} Employee Rights Under FMLAIf you require a reasonable accommodation during the application process, please contact humanresources@therapynotes.com.9/17/2026","company":"Therapynotes","rawCompany":"therapynotes","city":"Coatesville","state":"PA","isRemote":false,"isActive":false,"createdAt":"2026-09-18T09:16:54.860Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Cloud Systems Engineer - Site Reliability","description":"About UsTherapyNotes is the go-to superhero for behavioral health Practice Management and EHR software! Our top-notch SaaS solution handles scheduling, billing, documenting, telehealth, and more so clinicians can focus on awesome patient care.We're a dynamic team of pros who love to innovate and push the envelope, keeping our software cutting-edge. Join us, and let's revolutionize behavioral health software together while making a real difference!About The PositionWe are seeking a Site Reliability Engineer to improve the reliability and operability of the production services and shared platforms supporting our growing 24×7 SaaS environment. In this role, you will apply software and systems engineering practices to improve availability, performance, scalability, resilience, observability, incident response, and operational automation. You will partner with software development, infrastructure, database, security, and other technology teams to establish measurable reliability goals, reduce operational toil, and ensure services are supportable throughout their lifecycle. If you are passionate about building reliable systems, solving complex production problems, and driving continuous improvement, we want to hear from you.RequirementsBS degree in Information Systems, Engineering, or equivalent experience5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SREExperience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferredStrong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in productionExpertise with an observability platform; Datadog experience strongly preferred. Experience with Prometheus, Grafana, New Relic, or equivalent platforms is also valuableExperience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practicesExperience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvementExperience working in Agile/DevOps environments and operating production services using ITSM practices where applicablePrior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plusResponsibilitiesOwn and continuously improve how we use Datadog to make reliability visible and actionable across metrics, logs, traces, dashboards, monitors, alerts, and service-level viewsDesign, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24×7 SaaS platformPartner with service owners to define and improve reliability through meaningful SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practicesParticipate in and help drive incident management for production events, serving as an incident commander or technical responder as needed. Coordinate triage, service restoration, escalation, communication, incident documentation, root cause analysis, and completion of corrective actionsPartner with development teams to investigate issues across the infrastructure and application layers using metrics, logs, distributed traces, and code-level contextImprove deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and analysis of system failure modesPartner with other technical leaders to ensure all newly introduced systems are supportable and maintainable by both development and operationsProvide escalated technical guidance and support to other technology teams throughout the organizationProvide on-call coverage for production support and other duties as requiredEnsure supported systems and operational activities comply with organizational security, HIPAA, and operating policiesIdentify and eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible. Manage infrastructure as code using Terraform/OpenTofu and configuration automation using AnsibleBenefitsCompetitive salary - $110,000-$150,000Employer sponsored health, dental, vision, life, and disability insuranceRetirement plan with company contributionAnnual company profit sharingPersonal development/training budgetOpen, collaborative work environmentExtensive 2-week onboarding planComprehensive mentorship programEqual Opportunity Employer Statement & Applicant RightsTherapyNotes LLC is an Equal Opportunity Employer and does not discriminate based on race, color, religion, sex, national origin, age, disability, genetic information, or any other protected status under federal, state, or local law. We are committed to providing a workplace free of discrimination and harassment.For more information about your rights under federal employment laws, please review the following{{:}}Know Your Rights{{:}} Workplace Discrimination is IllegalFamily and Medical Leave Act (FMLA){{:}} Employee Rights Under FMLAIf you require a reasonable accommodation during the application process, please contact humanresources@therapynotes.com.9/17/2026","datePosted":"2026-09-18T09:16:54.860Z","dateModified":"2026-09-18T09:16:54.860Z","hiringOrganization":{"@type":"Organization","name":"Therapynotes","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Coatesville","addressRegion":"PA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a2fc9250dbf2d734bd231ef2"},"url":"https://jobsearcher.com/jobs/a2fc9250dbf2d734bd231ef2"}}