Senior Software Engineer (Backend/Fullstack)-Reliability Focused
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
Latin America | 100% RemoteAbout the RoleWe're looking for a senior engineer who can design, build, and operate production systems end-to-end. You'll work across our core stack building features and services, while also owning the reliability, observability, and performance of what you ship. You won't be handing your code off to a separate ops team — you'll be on-call for it, instrumenting it, and improving it based on how it behaves in production.What You'll DoDesign and build backend/fullstack features across Java 21, Spring/Spring boot, JPA/Hibernate, NodeJS, React/React NativeInstrument services with logging, metrics, and tracing (e.g., Splunk, Prometheus, Grafana, Datadog, OpenTelemetry) as a standard part of development, not an afterthoughtDefine and monitor SLIs/SLOs for the services you own; use error budgets to guide prioritization between feature work and reliability workParticipate in an on-call rotation; respond to and resolve production incidents affecting your servicesWrite and maintain postmortems/RCAs, and drive follow-up fixes to prevent recurrenceContribute to infrastructure-as-code (Terraform, CloudFormation, etc.) for the systems you buildPerform capacity planning and load testing for services ahead of scale eventsCollaborate with platform/infra teams on shared tooling, but take primary ownership of your service's healthParticipate in code reviews, architecture discussions, and mentor junior engineersWhat We're Looking For5+ years of professional software engineering experience, with deep expertise in Java 21, Spring/Spring boot, JPA/Hibernate, NodeJS, React/React NativeDemonstrated experience owning services in production — not just writing code, but debugging, scaling, and maintaining it liveComfort with observability tooling and reading dashboards/logs/traces to diagnose issues under pressureExperience with incident response processes (on-call, paging, postmortems)Working knowledge of cloud infrastructure with AWS and containerization (Docker, Kubernetes) — enough to reason about deployment and scaling, even if you're not a dedicated platform engineerStrong communication skills — you can explain a production issue to both engineers and stakeholdersBonus: experience with CI/CD pipelines, infrastructure-as-code, or chaos engineering practicesWhat This Role Is NotThis is not a dedicated SRE/DevOps/Platform Engineering role. You won't be building the observability platform itself or managing infrastructure for other teams — you'll be a strong practitioner of reliability engineering within your own feature work.What You’ll Love100% RemoteHolidays offPaid Time OffHealth insurance assistanceCompetitive USD compensationGrowth opportunities