🐶🐱 Staff Platform Engineer (Remote)
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
$175k–$225k base • Full-timeRadimal is a veterinary radiology and AI diagnostics platform delivering 24/7 imaging insights to hospitals nationwide. We combine board-certified radiologists with advanced AI to support real-time clinical decision-making for patients when it matters most.Our platform spans high-throughput medical imaging, GPU-backed inference, global distribution, and enterprise-grade reliability. As we scale, we’re investing in senior platform ownership to make the system safer, more predictable, and easier for engineers to build on.Why This Role ExistsRadimal has grown quickly. While the platform is working, operational ownership has been too diffuse. Reliability, on-call clarity, and platform standards need a single senior owner who can reduce noise, establish guardrails, and make the system more predictable as we scale.This role exists to bring focus, ownership, and calm to the platform layer.The RoleWe’re hiring a Staff Platform Engineer to own the technical foundations that enable Radimal’s engineering teams to move quickly and reliably.This is a senior, hands-on role with real accountability for platform architecture, infrastructure, and production systems. You’ll own DevOps, reliability, and on-call systems, with authority to investigate and diagnose issues across the full stack.You will not be expected to do everything at once. Success comes from establishing ownership, setting priorities, and making the system more predictable over time.A core part of this role is operational containment and reliability ownership. You’ll reduce operational burden on product and AI teams by owning platform standards, tooling, and reliability so others can focus on building.You’ll work closely with the CEO and VP of Engineering on platform strategy, architectural tradeoffs, and operational risk, while maintaining clear ownership of production systems.This role is for someone who wants true ownership and influence through execution, not advisory distance.What You Will OwnPlatform EngineeringOwn the core platform foundations that support all product and AI developmentBuild shared infrastructure, libraries, and patterns that make it easier to ship safelyEstablish clear interfaces and ownership boundaries so teams can move independentlyImprove developer experience through better CI/CD, local tooling, and observabilityRaise the overall operational maturity of the engineering organizationInfrastructure and CloudOwn and evolve Radimal’s AWS and Terraform footprintLead deployments across ECS, Fargate, EC2, containerized services, and GPU workloadsManage and improve workloads running on Render and ModalMake architectural decisions for scale, reliability, and cost efficiencyEngineering EnablementReduce operational burden on product and AI engineers by owning reliability and toolingCreate guardrails that increase safety without slowing developmentEnable engineers to self-serve infrastructure and diagnostics where appropriateReliability, On-Call, and OperationsOwn production uptime, SLOs, and operational healthDesign and own on-call coverage and escalation modelsServe as senior escalation during incidents while building systems that minimize the need for escalationLead incident response and post-incident reviews with clear accountabilityEliminate ambiguity around who owns production at all timesOver time, success in this role means fewer incidents, fewer escalations, and a platform that largely runs without heroics.Observability and PerformanceOperate and extend Grafana and Prometheus monitoring stacksImprove alerting, diagnostics, and operational visibilityBuild high-availability and fault-tolerant architecturesImplement caching, CDN, and performance strategies for global scaleFull-Stack InvestigationInvestigate production issues across infrastructure, backend services, data pipelines, AI inference workflows, and frontend behaviorTrace request flow end to end across GraphQL APIs, Python services, and React applicationsRead and debug React code as needed to understand client-side behavior and API usageForm and test hypotheses during incidents to drive fast, accurate resolutionKnow when to dive deep personally and when to pull in specialistsML Ops and AI Platform SupportUnderstand ML Ops fundamentals including model deployment, versioning, and monitoringSupport GPU-backed inference workloads and AI service reliabilityPartner with AI engineers to ensure models are observable, debuggable, and production-readyIdentify and mitigate operational risks related to model performance, latency, and failuresLeadership and CollaborationPartner with the CEO and VP of Engineering on platform strategy and architectural tradeoffsProvide clear, grounded assessments of platform risk and readinessAct as a trusted technical owner during high-impact decisions and incidentsAlign platform reliability with product and business goalsSecurity and ComplianceStrengthen infrastructure security and access controlsSupport enterprise security reviews, penetration testing, and SOC 2 readinessImprove auditability, monitoring, and operational hygieneWhat We’re Looking For7+ years operating production systems at scaleStrong Python experience for automation and backend toolingDeep AWS experience (ECS, Fargate, EC2, ECR, RDS, CloudFront, IAM)Strong Terraform, Docker, CI/CD, and infrastructure-as-code expertiseHands-on experience with Grafana and PrometheusExperience with Postgres and modern backend architecturesStrong understanding of distributed systems, caching, and performanceComfort debugging across GraphQL APIs, Python services, and React frontendsWorking knowledge of ML Ops concepts and AI inference systemsClear communicator with a strong ownership mindsetBonus ExperienceMedical imaging or DICOM workflowsGPU compute, AI inference, or ML pipeline integrationEnterprise security reviews or penetration testingGraphQL or Hasura-based platformsWhy Join RadimalHigh ownership of a mission-critical clinical platformDeep technical challenges with real patient impactClear mandate and authority to shape how engineering worksFully remote team with high trust and low bureaucracyOpportunity to grow into broader technical leadership over time