Senior Database Reliability Engineer (DBRE)
Overview
In this Senior DBRE role, you will design and operate scalable PostgreSQL-backed data platforms that power mission-critical systems. You’ll collaborate with SRE, Platform, and Engineering teams to boost performance, reliability, and automation. This hands-on role focuses on resilient data infrastructure, not just administration. You’ll tackle high-availability, performance, and incident response at scale, contributing to secure, enterprise-grade data services. You’ll join a mission-driven, fast-paced environment engineering impactful, reliable systems.
Compensation / Benefitshealth, dental and vision insurance401(k)flexible spending accountpaid leave including PTO and parental leaveequity (where applicable)bonus (where applicable)
ResponsibilitiesDesign, implement, and operate highly available PostgreSQL clusters (physical and logical replication, sharding/partitioning, failover automation)Optimize queries, indexing, schema design, and storage enginesPerform capacity planning, growth forecasting, and workload modelingOwn high-availability strategies including automatic failover, multi-AZ/multi-region setups, and disaster recoveryDevelop automation for provisioning, configuration, backups, failovers, vacuum tuning, and schema management using Terraform, Ansible, Kubernetes Operators, or custom toolingBuild monitoring, alerting, and self-healing systems for PostgreSQL and MySQLLead incident response for database issues and perform root-cause analysis and permanent fixesCollaborate with software engineers to review SQL, optimize schemas, and promote best practices
Key requirements4+ years of hands-on PostgreSQL experience in high-volume or large-scale production environmentsStrong knowledge of PostgreSQL internals (WAL, MVCC, bloat/vacuum tuning, query planner, indexing)Production experience with MySQL (InnoDB internals, replication, performance tuning)Advanced SQL and good grasp of schema design and query optimizationExperience with Linux, networking, and systems troubleshootingExperience building automation with Go or PythonExperience with monitoring tools (Prometheus, Grafana, Datadog, PMM, pg_stat_statements)Hands-on experience with cloud environments (AWS or GCP)cross-functional collaborationproblem-solving under pressurestrong communication for incident responsesPostgreSQL at scale (replication, failover, tuning)MySQL InnoDB internals and replicationSQL optimization and schema design