{"schemaVersion":"jobsearcher.job.v1","id":"221e1c95f9f094282bb12720","url":"https://jobsearcher.com/jobs/221e1c95f9f094282bb12720","canonicalUrl":"https://jobsearcher.com/jobs/221e1c95f9f094282bb12720","title":"Principal DevOps Architect","description":"“Be part of a company that is influential and the standard for a rapidly evolving industry!”\nWHO ARE WE?\nRealTime eClinical Solutions is a Global Leader and rapidly growing SaaS technology company that provides comprehensive Software Solutions to the clinical research industry.\nOur Vision is to reshape the global clinical research industry with innovative solutions that help advance medicine and save lives. Our cloud-based solutions are dedicated to solving complex problems and simplifying clinical research processes to be more organized, efficient, and cost-effective. We are based out of San Antonio, TX but are truly a remote and telecommuting company.\nWHAT ARE WE LOOKING FOR?\nThe Principal DevOps Architect is the senior technical authority for the cloud platform that runs RealTime’s global, multi-tenant clinical-trial services, and drives RealTime’s adoption of infrastructure as code and AI tooling. This is a hands-on individual-contributor architect role rather than a people-management role: you design the platform architecture, set the standards and reference implementations other teams build on, and still write Terraform, build pipelines, and stand up the AI platform yourself. You define how reliability is measured against SLOs, how releases ship, and how the AI/ML platform is built and governed, and you keep the environment HIPAA-compliant and SOC 2 Type 2 audit-ready. You influence products, software, and QA through architecture and example rather than direct authority.\nWHAT WILL YOU BE DOING?\nTechnical Leadership & Architecture\nOwn the platform architecture and technical roadmap for infrastructure, deployment, observability, and the AI platform, and advise the VP of Software Architecture and engineering leadership.\nSet the engineering standards, patterns, and golden paths for infrastructure as code, CI/CD, and AI tooling, and drive their adoption through reference implementations and architecture reviews.\nAct as hands-on technical authority and mentor to engineers across teams, leading by example in code, infrastructure, and incident response, without direct management responsibility.\nDrive RealTime’s move to infrastructure as code and AI tooling and prevent uncontrolled spread of unvetted AI tools.\nRecommend build/buy decisions for platform tooling and evaluate vendors, with final decisions owned by the VP and CTO.\nInfrastructure as Code & Cloud Platform\nManage all cloud infrastructure as code in Terraform as the single source of truth: reusable modules, remote state, peer-reviewed infrastructure pull requests, drift detection, and automated plan/apply in CI/CD.\nEnforce policy-as-code (for example, OPA or Sentinel) so infrastructure changes meet security and cost guardrails before they merge, with separation between the author and approver of a change.\nDesign and deploy AWS infrastructure for performance, availability, recoverability, and security across development, UAT, staging, and production, meeting the CIS Critical Security Controls.\nMaintain and improve the multi-tenant database and hosting architecture in a cloud-hosted environment, including replication and per-tenant isolation.\nCI/CD & Release Engineering\nBuild and operate CI/CD pipelines for large-scale applications on AWS using GitHub Actions or equivalent, with automated build, test, and deployment.\nOwn release management and rollback: safe deployment strategies (blue/green, canary), fast and reliable rollbacks, and release gates for QA and customer acceptance.\nPackage and run containerized workloads on Docker and Kubernetes (EKS), and automate configuration with tools such as Ansible and Packer.\nReliability & Observability\nLead the SLI/SLO/SLA program and manage error budgets to balance reliability against delivery speed.\nOperate modern observability using cloud-native tooling and OpenTelemetry (metrics, logs, and distributed traces) to drive down MTTD and MTTR.\nAnalyze production events to improve reliability, operability, and customer experience, and lead blameless post-incident reviews.\nParticipate in on-call rotations, triage and resolve incidents, and provide workarounds or escalation to service owners.\nAI / ML Platform Operations\nProvision and operate the AI/ML platform, including Anthropic Claude models served through AWS Bedrock and, where used, the Anthropic API, along with inference endpoints and RealTime’s internal MCP services, all managed as infrastructure as code.\nBuild guardrails and observability for AI workloads: prompt and response logging, rate limits, model and token cost controls, and usage auditing across Claude and any other models in use.\nEnforce the data boundary for AI systems so Protected Health Information is handled only through approved, BAA-covered providers (AWS Bedrock under the AWS BAA, and Anthropic under an Anthropic BAA where the Anthropic API is used), and confirm no PHI leaves the compliant boundary.\nOperate and support the agentic and MCP tooling used across engineering and operations (for example Claude Code and internal MCP services) and govern how that tooling accesses code and data with least-privilege tool access and human-in-the-loop escalation.\nStand up LLMOps practices: prompt versioning and regression detection, evaluation frameworks for non-deterministic output, token-level cost attribution, and audit logging of agent actions.\nRoute inference across a multi-provider model portfolio and manage protocol choices (such as MCP) to avoid vendor lock-in.\nUse AI-assisted tooling to accelerate operations where appropriate, such as incident triage, runbook generation, and log analysis.\nSecurity, Compliance & Cost\nOperate and evidence the platform controls required for SOC 2 Type 2 and HIPAA continuously and support external audits.\nOwn secrets management so no credentials, keys, or tokens live in source code, container images, or Terraform state.\nRun supply-chain security for the pipeline: image and dependency scanning, SBOM generation, and vulnerability remediation to defined SLAs.\nManage cloud cost (FinOps): tagging, budgets, and right-sizing across compute, storage, and AI/Bedrock spend.\nCollaboration\nCoordinate with product, development, support, operations, and QA so installation and integration are automated and well-documented.\nCommunicate and work effectively across a distributed, multi-time-zone team, and cultivate cross-team collaboration and trust.\nWHAT DO YOU NEED?\nBachelor’s degree in software engineering or an equivalent combination of technical education and work experience.\n10+ years in SRE/DevOps/platform engineering delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability, including time at a senior individual-contributor or architect level (Staff, Principal, or Architect).\nProven technical authority across teams: you set architecture and standards and influence delivery through expertise and example rather than direct management.\nDemonstrated experience driving adoption of a new practice or platform (infrastructure as code, a CI/CD overhaul, or an AI/ML platform) across multiple teams.\nHands-on experience with Terraform, including writing reusable modules that other teams consume through self-service, remote state, and changing management in a CI/CD pipeline.\nExperience building and operating CI/CD pipelines for large-scale applications on AWS (GitHub Actions, Jenkins, GitLab, or AWS-native).\nExperience running containerized workloads on Docker and Kubernetes.\nExperience with monitoring and troubleshooting using cloud-native tooling and OpenTelemetry; New Relic experience a plus.\nLinux system administration, Unix scripting, and automation.\n2+ years in one or more PHP, MySQL, and SQL is a plus, matching RealTime’s stack.\nExperience working in a HIPAA / HITECH / HITRUST / PHI / PII or PCI DSS environment.\nWHAT SETS YOU APART?\nExperience operating an AI/ML or GenAI platform in production, including Anthropic Claude via AWS Bedrock or the Anthropic API, or comparable model-serving infrastructure, with cost and guardrail controls.\nLLMOps maturity: evaluation pipelines for non-deterministic output, prompt versioning, token cost attribution, and incident response for AI-specific failures (data modification, unwanted workflow triggering, or exfiltration by an agent).\nExperience supporting LLM applications, agents, or MCP services (for example Claude Code or custom MCP servers), and governing how they access regulated data.\nFamiliarity with AI security and governance: OWASP LLM Top 10, key management (BYOK), and alignment to frameworks such as the NIST AI RMF.\nPlatform-as-product mindset with a focus on internal developer experience and self-service.\nExperience leading or supporting a SOC 2 Type 2 examination and familiarity with security benchmarks such as CIS, OWASP, PCI DSS, and FedRAMP.\nExperience with policy-as-code, secrets management, and supply-chain security (SBOM, image scanning).\nExperience with FinOps and managing large AWS infrastructure and its cost.\nExperience in clinical research or healthcare technology.\nSelf-starter who establishes best practices, ships on time, and communicates complex technical information clearly to any audience.\nWHAT IS IN IT FOR YOU?\nThe company sponsors health insurance, long-term disability, and life insurance.\nUnlimited Paid Time Off.\n10 paid Holidays.\nPaid Parental Leave.\nWork Anniversary Bonus.\nParticipation in the Employee of the Quarter Program.\nMonthly $100 Connectivity Stipend Reimbursement.\nRealTime matches employee 401K contributions at 100% of the first 3% invested and 50% of the next 2% invested.\nAll successful candidates must complete and pass reference and background checks.\nThe desired salary must be indicated for the application to be considered.\nThe pay rate is commensurate with experience and is determined on an individual basis after an interview has occurred.\nEqual Opportunity Employer – RealTime eClinical Solutions strongly values diversity and is committed to equal opportunity and non-discrimination in all of its policies and practices, including employment.\nYour Right to Work – In compliance with federal law, all persons hired will be required to verify identity and eligibility.\nThank you for your interest in RealTime eClinical Solutions.","company":"Realtime Software Solutions","rawCompany":"realtime software solutions","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-03T23:33:06.949Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Principal DevOps Architect","description":"“Be part of a company that is influential and the standard for a rapidly evolving industry!”\nWHO ARE WE?\nRealTime eClinical Solutions is a Global Leader and rapidly growing SaaS technology company that provides comprehensive Software Solutions to the clinical research industry.\nOur Vision is to reshape the global clinical research industry with innovative solutions that help advance medicine and save lives. Our cloud-based solutions are dedicated to solving complex problems and simplifying clinical research processes to be more organized, efficient, and cost-effective. We are based out of San Antonio, TX but are truly a remote and telecommuting company.\nWHAT ARE WE LOOKING FOR?\nThe Principal DevOps Architect is the senior technical authority for the cloud platform that runs RealTime’s global, multi-tenant clinical-trial services, and drives RealTime’s adoption of infrastructure as code and AI tooling. This is a hands-on individual-contributor architect role rather than a people-management role: you design the platform architecture, set the standards and reference implementations other teams build on, and still write Terraform, build pipelines, and stand up the AI platform yourself. You define how reliability is measured against SLOs, how releases ship, and how the AI/ML platform is built and governed, and you keep the environment HIPAA-compliant and SOC 2 Type 2 audit-ready. You influence products, software, and QA through architecture and example rather than direct authority.\nWHAT WILL YOU BE DOING?\nTechnical Leadership & Architecture\nOwn the platform architecture and technical roadmap for infrastructure, deployment, observability, and the AI platform, and advise the VP of Software Architecture and engineering leadership.\nSet the engineering standards, patterns, and golden paths for infrastructure as code, CI/CD, and AI tooling, and drive their adoption through reference implementations and architecture reviews.\nAct as hands-on technical authority and mentor to engineers across teams, leading by example in code, infrastructure, and incident response, without direct management responsibility.\nDrive RealTime’s move to infrastructure as code and AI tooling and prevent uncontrolled spread of unvetted AI tools.\nRecommend build/buy decisions for platform tooling and evaluate vendors, with final decisions owned by the VP and CTO.\nInfrastructure as Code & Cloud Platform\nManage all cloud infrastructure as code in Terraform as the single source of truth: reusable modules, remote state, peer-reviewed infrastructure pull requests, drift detection, and automated plan/apply in CI/CD.\nEnforce policy-as-code (for example, OPA or Sentinel) so infrastructure changes meet security and cost guardrails before they merge, with separation between the author and approver of a change.\nDesign and deploy AWS infrastructure for performance, availability, recoverability, and security across development, UAT, staging, and production, meeting the CIS Critical Security Controls.\nMaintain and improve the multi-tenant database and hosting architecture in a cloud-hosted environment, including replication and per-tenant isolation.\nCI/CD & Release Engineering\nBuild and operate CI/CD pipelines for large-scale applications on AWS using GitHub Actions or equivalent, with automated build, test, and deployment.\nOwn release management and rollback: safe deployment strategies (blue/green, canary), fast and reliable rollbacks, and release gates for QA and customer acceptance.\nPackage and run containerized workloads on Docker and Kubernetes (EKS), and automate configuration with tools such as Ansible and Packer.\nReliability & Observability\nLead the SLI/SLO/SLA program and manage error budgets to balance reliability against delivery speed.\nOperate modern observability using cloud-native tooling and OpenTelemetry (metrics, logs, and distributed traces) to drive down MTTD and MTTR.\nAnalyze production events to improve reliability, operability, and customer experience, and lead blameless post-incident reviews.\nParticipate in on-call rotations, triage and resolve incidents, and provide workarounds or escalation to service owners.\nAI / ML Platform Operations\nProvision and operate the AI/ML platform, including Anthropic Claude models served through AWS Bedrock and, where used, the Anthropic API, along with inference endpoints and RealTime’s internal MCP services, all managed as infrastructure as code.\nBuild guardrails and observability for AI workloads: prompt and response logging, rate limits, model and token cost controls, and usage auditing across Claude and any other models in use.\nEnforce the data boundary for AI systems so Protected Health Information is handled only through approved, BAA-covered providers (AWS Bedrock under the AWS BAA, and Anthropic under an Anthropic BAA where the Anthropic API is used), and confirm no PHI leaves the compliant boundary.\nOperate and support the agentic and MCP tooling used across engineering and operations (for example Claude Code and internal MCP services) and govern how that tooling accesses code and data with least-privilege tool access and human-in-the-loop escalation.\nStand up LLMOps practices: prompt versioning and regression detection, evaluation frameworks for non-deterministic output, token-level cost attribution, and audit logging of agent actions.\nRoute inference across a multi-provider model portfolio and manage protocol choices (such as MCP) to avoid vendor lock-in.\nUse AI-assisted tooling to accelerate operations where appropriate, such as incident triage, runbook generation, and log analysis.\nSecurity, Compliance & Cost\nOperate and evidence the platform controls required for SOC 2 Type 2 and HIPAA continuously and support external audits.\nOwn secrets management so no credentials, keys, or tokens live in source code, container images, or Terraform state.\nRun supply-chain security for the pipeline: image and dependency scanning, SBOM generation, and vulnerability remediation to defined SLAs.\nManage cloud cost (FinOps): tagging, budgets, and right-sizing across compute, storage, and AI/Bedrock spend.\nCollaboration\nCoordinate with product, development, support, operations, and QA so installation and integration are automated and well-documented.\nCommunicate and work effectively across a distributed, multi-time-zone team, and cultivate cross-team collaboration and trust.\nWHAT DO YOU NEED?\nBachelor’s degree in software engineering or an equivalent combination of technical education and work experience.\n10+ years in SRE/DevOps/platform engineering delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability, including time at a senior individual-contributor or architect level (Staff, Principal, or Architect).\nProven technical authority across teams: you set architecture and standards and influence delivery through expertise and example rather than direct management.\nDemonstrated experience driving adoption of a new practice or platform (infrastructure as code, a CI/CD overhaul, or an AI/ML platform) across multiple teams.\nHands-on experience with Terraform, including writing reusable modules that other teams consume through self-service, remote state, and changing management in a CI/CD pipeline.\nExperience building and operating CI/CD pipelines for large-scale applications on AWS (GitHub Actions, Jenkins, GitLab, or AWS-native).\nExperience running containerized workloads on Docker and Kubernetes.\nExperience with monitoring and troubleshooting using cloud-native tooling and OpenTelemetry; New Relic experience a plus.\nLinux system administration, Unix scripting, and automation.\n2+ years in one or more PHP, MySQL, and SQL is a plus, matching RealTime’s stack.\nExperience working in a HIPAA / HITECH / HITRUST / PHI / PII or PCI DSS environment.\nWHAT SETS YOU APART?\nExperience operating an AI/ML or GenAI platform in production, including Anthropic Claude via AWS Bedrock or the Anthropic API, or comparable model-serving infrastructure, with cost and guardrail controls.\nLLMOps maturity: evaluation pipelines for non-deterministic output, prompt versioning, token cost attribution, and incident response for AI-specific failures (data modification, unwanted workflow triggering, or exfiltration by an agent).\nExperience supporting LLM applications, agents, or MCP services (for example Claude Code or custom MCP servers), and governing how they access regulated data.\nFamiliarity with AI security and governance: OWASP LLM Top 10, key management (BYOK), and alignment to frameworks such as the NIST AI RMF.\nPlatform-as-product mindset with a focus on internal developer experience and self-service.\nExperience leading or supporting a SOC 2 Type 2 examination and familiarity with security benchmarks such as CIS, OWASP, PCI DSS, and FedRAMP.\nExperience with policy-as-code, secrets management, and supply-chain security (SBOM, image scanning).\nExperience with FinOps and managing large AWS infrastructure and its cost.\nExperience in clinical research or healthcare technology.\nSelf-starter who establishes best practices, ships on time, and communicates complex technical information clearly to any audience.\nWHAT IS IN IT FOR YOU?\nThe company sponsors health insurance, long-term disability, and life insurance.\nUnlimited Paid Time Off.\n10 paid Holidays.\nPaid Parental Leave.\nWork Anniversary Bonus.\nParticipation in the Employee of the Quarter Program.\nMonthly $100 Connectivity Stipend Reimbursement.\nRealTime matches employee 401K contributions at 100% of the first 3% invested and 50% of the next 2% invested.\nAll successful candidates must complete and pass reference and background checks.\nThe desired salary must be indicated for the application to be considered.\nThe pay rate is commensurate with experience and is determined on an individual basis after an interview has occurred.\nEqual Opportunity Employer – RealTime eClinical Solutions strongly values diversity and is committed to equal opportunity and non-discrimination in all of its policies and practices, including employment.\nYour Right to Work – In compliance with federal law, all persons hired will be required to verify identity and eligibility.\nThank you for your interest in RealTime eClinical Solutions.","datePosted":"2026-08-03T23:33:06.949Z","dateModified":"2026-08-03T23:33:06.949Z","hiringOrganization":{"@type":"Organization","name":"Realtime Software Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"221e1c95f9f094282bb12720"},"url":"https://jobsearcher.com/jobs/221e1c95f9f094282bb12720"}}