JOBSEARCHER

Infrastructure Engineer, Database

Database Infra EngineerWe're building a database specifically designed for AI observability and evaluation, and we need someone to own the infrastructure layer that keeps it running reliably at scale. As a Database Infra Engineer on the SmithDB team, you won't be designing the storage engine — you'll be making sure the engine never goes down, scales seamlessly as our customer base grows, and is operationally excellent across cloud environments.What you'll doOwn the deployment and operations of SmithDB across cloud environments — including cluster lifecycle management, blue/green and rolling upgrades, and automated failoverBuild and maintain the infrastructure tooling (Terraform, Kubernetes, Helm, or equivalent) that provisions, configures, and scales SmithDB nodesOwn the Kubernetes infrastructure that runs our distributed database services (multi-tenant, high throughput, low latency)Build and improve deployment pipelines, rollout strategies, and infrastructure-as-code for the storage layerDrive reliability engineering efforts: incident response, postmortems, SLOs, and disaster recovery for a system operating at massive scaleManage capacity planning and cost efficiency — model growth, rightsize resources, and ensure SmithDB can absorb traffic spikes from our largest customers without manual interventionBuild the CI/CD pipeline for database infrastructure changes — safe, tested, and fast promotion from dev through staging to productionCollaborate closely with SmithDB internals engineers to translate new engine features into production-ready infrastructure changes and ensure safe, low-risk rolloutsWhat you'll bring5+ years of experience in infrastructure, platform engineering, or SRE with hands-onStrong hands-on experience with Kubernetes and cloud infrastructure (AWS/GCP/Azure)Solid scripting/systems programming ability (Go, Python, or similar);Experience with infrastructure-as-code and CI/CD tooling (Terraform, Helm, ArgoCD, or similar)Deep familiarity with at least one major cloud provider (AWS, GCP, or Azure) and the primitives used to run stateful workloads reliably — persistent volumes, managed node groups, cloud storage, etc.Infrastructure-as-code fluency — you write Terraform (or Pulumi/CDK) as your primary language, not an afterthoughtStrong operational instincts — you've been on-call for high-traffic data systems, you know how to triage under pressure, and you write runbooks that actually get usedExperience with container orchestration (Kubernetes) and deploying stateful workloads in productionA bias for automation — if you've done something manual twice, you're already thinking about how to make it never happen againStrong written and oral communication skills, with the ability to translate infrastructure health into language product and business stakeholders understandThe DNA to thrive in a fast-moving, high-autonomy environment — you see gaps as opportunities and own them end to endNice to HaveOwnership of production database systems (Postgres, ClickHouse, Redis, or similar)Comfort reading and reasoning about Rust is a plus, as it's the language our database is written inUnderstanding of database reliability concepts — replication, backups, point-in-time recovery, connection pooling, and graceful degradation under loadCompensationSalary Range: $180,000-$230,000 USDCompensation Philosophy:We offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.