{"schemaVersion":"jobsearcher.job.v1","id":"c4ccb03e7c5e86236041e624","url":"https://jobsearcher.com/jobs/c4ccb03e7c5e86236041e624","canonicalUrl":"https://jobsearcher.com/jobs/c4ccb03e7c5e86236041e624","title":"Senior Software Engineer, Reliability Engineering - REMOTE","description":"Job DescriptionSenior Software Engineer, Reliability EngineeringPosition SummaryOur Healthcare Client is seeking a Senior Software Engineer, Reliability Engineering who combines strong software development expertise with a passion for building highly reliable, scalable, observable, and operationally excellent systems.This role sits at the intersection of software engineering, cloud architecture, and site reliability engineering. You will design, build, and support customer-facing applications and APIs while ensuring reliability, scalability, security, performance, and operational excellence are built into every stage of the software lifecycle. Unlike traditional SRE or operations-focused roles, this position places equal emphasis on software engineering and reliability engineering. You will develop business capabilities, while helping teams improve observability, resiliency, deployment safety, and operational maturity.You will work closely with Product Engineering, Architecture, Platform Engineering, Security, and Product teams to enable rapid innovation without compromising system reliability. This role is designed to support mobility across engineering disciplines, allowing engineers to rotate between product development and reliability engineering teams.This is an engineering-first role with end-to-end production ownership, including participation in an on-call rotation for critical customer-facing applications and services.Hiring PhilosophyWe believe reliability is a shared engineering responsibility, not a separate operational function. The ideal candidate is a strong software engineer who enjoys building resilient systems, automating operational processes, improving developer productivity, and owning services throughout their entire lifecycle. Engineers in this role are expected to contribute to production ready code, influence system design, and help create a culture where reliability, scalability, and operational excellence are integral parts of software development.What You'll DoSoftware EngineeringDesign, develop, test, deploy, and maintain cloud-native resilient distributed systems using APIs, microservices, messaging, caching, and event-driven architectures.Build scalable and resilient services using Node.js and TypeScriptParticipate in architecture and design discussions with a focus on scalability, resiliency, maintainability, security, and performance.Apply modern software engineering practices, including domain-driven design, test automation, code reviews, and continuous integration.Create and maintain technical documentation, architecture diagrams, and engineering standards.Reliability EngineeringDesign systems with reliability, recoverability, observability, and operational excellence built in from inception.Define and improve Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and operational health metrics.Participate in incident response and lead the investigation of complex production issues.Participate in the team's on-call rotation, providing timely triage, escalation, communication, and resolution support for production incidents impacting customer-facing applications.Conduct root cause analysis and drive blameless post-incident reviews.Partner with engineering teams to eliminate recurring operational issues through engineering solutions and automation.Improve release safety, deployment reliability, and production readiness across services.Observability & Operational ExcellenceBuild and maintain monitoring, logging, alerting, and distributed tracing solutions.Implement modern observability practices using tools such as OpenTelemetry, New Relic, Grafana, Splunk, Datadog, CloudWatch, etc.Develop actionable dashboards, reliability scorecards, and service health metrics.Improve detection, diagnosis, and recovery processes while reducing alert fatigue and improving signal quality.Automation EngineeringEliminate operational toil through software engineering and automation.Contribute to CI/CD pipelines and deployment automation.AI-Assisted EngineeringLeverage AI-assisted engineering capabilities to accelerate software delivery, operational workflows, incident triage, and root cause analysis.Evaluate and adopt emerging AI-enabled developer productivity and reliability engineering tools.Technical LeadershipMentor engineers in both software engineering and reliability engineering best practices.Promote a culture of ownership, accountability, automation, and continuous improvement.Influence engineering standards, platform direction, reliability objectives, and software development practices across the organization.We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.Skills and RequirementsRequired Qualifications7+ years of professional software engineering experience.3+ years of experience building APIs, services, or distributed systems using Node.js and TypeScript3+ years of cloud experience with AWS3+ years' experience with Splunk, New Relic and Cloud WatchExperience using AI ToolsExperience designing and building scalable microservice architectures.Experience developing, deploying, and operating production systems at scale.Strong understanding of distributed systems, resiliency, scalability, observability, and performance engineering.Experience with monitoring, logging, tracing, and operational tooling.Experience participating in an on-call rotation or supporting production incidents for business-critical systems.Experience with CI/CD pipelines, deployment automation, and modern software delivery practices.Strong troubleshooting, debugging, and root cause analysis skills.Excellent communication, collaboration, and leadership skills.Preferred QualificationsExperience in Site Reliability Engineering, Production Engineering, Platform Engineering, DevOps, or Internal Developer Platforms.Experience with Kubernetes, EKS, AKS, GKE, ECS, or other container orchestration platforms.Experience implementing SLOs, SLIs, error budgets, and reliability scorecards.Experience with OpenTelemetry, New Relic, Splunk, Grafana, Prometheus, Datadog, CloudWatch, Azure Monitor, or Google Cloud Operations Suite.Experience supporting customer-facing web, mobile, and API platforms at scale.Experience building AI-enabled applications, operational tooling, or developer productivity solutions.Experience working within healthcare or other highly regulated industries.","company":"Insight Global","rawCompany":"insight global","isRemote":true,"isActive":false,"createdAt":"2026-09-04T02:38:12.561Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Software Engineer, Reliability Engineering - REMOTE","description":"Job DescriptionSenior Software Engineer, Reliability EngineeringPosition SummaryOur Healthcare Client is seeking a Senior Software Engineer, Reliability Engineering who combines strong software development expertise with a passion for building highly reliable, scalable, observable, and operationally excellent systems.This role sits at the intersection of software engineering, cloud architecture, and site reliability engineering. You will design, build, and support customer-facing applications and APIs while ensuring reliability, scalability, security, performance, and operational excellence are built into every stage of the software lifecycle. Unlike traditional SRE or operations-focused roles, this position places equal emphasis on software engineering and reliability engineering. You will develop business capabilities, while helping teams improve observability, resiliency, deployment safety, and operational maturity.You will work closely with Product Engineering, Architecture, Platform Engineering, Security, and Product teams to enable rapid innovation without compromising system reliability. This role is designed to support mobility across engineering disciplines, allowing engineers to rotate between product development and reliability engineering teams.This is an engineering-first role with end-to-end production ownership, including participation in an on-call rotation for critical customer-facing applications and services.Hiring PhilosophyWe believe reliability is a shared engineering responsibility, not a separate operational function. The ideal candidate is a strong software engineer who enjoys building resilient systems, automating operational processes, improving developer productivity, and owning services throughout their entire lifecycle. Engineers in this role are expected to contribute to production ready code, influence system design, and help create a culture where reliability, scalability, and operational excellence are integral parts of software development.What You'll DoSoftware EngineeringDesign, develop, test, deploy, and maintain cloud-native resilient distributed systems using APIs, microservices, messaging, caching, and event-driven architectures.Build scalable and resilient services using Node.js and TypeScriptParticipate in architecture and design discussions with a focus on scalability, resiliency, maintainability, security, and performance.Apply modern software engineering practices, including domain-driven design, test automation, code reviews, and continuous integration.Create and maintain technical documentation, architecture diagrams, and engineering standards.Reliability EngineeringDesign systems with reliability, recoverability, observability, and operational excellence built in from inception.Define and improve Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and operational health metrics.Participate in incident response and lead the investigation of complex production issues.Participate in the team's on-call rotation, providing timely triage, escalation, communication, and resolution support for production incidents impacting customer-facing applications.Conduct root cause analysis and drive blameless post-incident reviews.Partner with engineering teams to eliminate recurring operational issues through engineering solutions and automation.Improve release safety, deployment reliability, and production readiness across services.Observability & Operational ExcellenceBuild and maintain monitoring, logging, alerting, and distributed tracing solutions.Implement modern observability practices using tools such as OpenTelemetry, New Relic, Grafana, Splunk, Datadog, CloudWatch, etc.Develop actionable dashboards, reliability scorecards, and service health metrics.Improve detection, diagnosis, and recovery processes while reducing alert fatigue and improving signal quality.Automation EngineeringEliminate operational toil through software engineering and automation.Contribute to CI/CD pipelines and deployment automation.AI-Assisted EngineeringLeverage AI-assisted engineering capabilities to accelerate software delivery, operational workflows, incident triage, and root cause analysis.Evaluate and adopt emerging AI-enabled developer productivity and reliability engineering tools.Technical LeadershipMentor engineers in both software engineering and reliability engineering best practices.Promote a culture of ownership, accountability, automation, and continuous improvement.Influence engineering standards, platform direction, reliability objectives, and software development practices across the organization.We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.Skills and RequirementsRequired Qualifications7+ years of professional software engineering experience.3+ years of experience building APIs, services, or distributed systems using Node.js and TypeScript3+ years of cloud experience with AWS3+ years' experience with Splunk, New Relic and Cloud WatchExperience using AI ToolsExperience designing and building scalable microservice architectures.Experience developing, deploying, and operating production systems at scale.Strong understanding of distributed systems, resiliency, scalability, observability, and performance engineering.Experience with monitoring, logging, tracing, and operational tooling.Experience participating in an on-call rotation or supporting production incidents for business-critical systems.Experience with CI/CD pipelines, deployment automation, and modern software delivery practices.Strong troubleshooting, debugging, and root cause analysis skills.Excellent communication, collaboration, and leadership skills.Preferred QualificationsExperience in Site Reliability Engineering, Production Engineering, Platform Engineering, DevOps, or Internal Developer Platforms.Experience with Kubernetes, EKS, AKS, GKE, ECS, or other container orchestration platforms.Experience implementing SLOs, SLIs, error budgets, and reliability scorecards.Experience with OpenTelemetry, New Relic, Splunk, Grafana, Prometheus, Datadog, CloudWatch, Azure Monitor, or Google Cloud Operations Suite.Experience supporting customer-facing web, mobile, and API platforms at scale.Experience building AI-enabled applications, operational tooling, or developer productivity solutions.Experience working within healthcare or other highly regulated industries.","datePosted":"2026-09-04T02:38:12.561Z","dateModified":"2026-09-04T02:38:12.561Z","hiringOrganization":{"@type":"Organization","name":"Insight Global","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"c4ccb03e7c5e86236041e624"},"url":"https://jobsearcher.com/jobs/c4ccb03e7c5e86236041e624"}}