JOBSEARCHER

MTS · Engineering · Distributed Systems

Nebula is the system of state for enterprise agent workflows. This role owns the storage and execution systems beneath it: object-storage-native architecture, distributed execution, partitioning, multi-region deployment, and customer-managed infrastructure. It's a product engineering role, not general DevOps or internal-platform work.What you'll work onEvolve Nebula's S3-backed storage architecture, including logs, manifests, snapshots, indexes, checkpoints, caching, and compaction.Build durable ingestion and workflow execution on top of queues, including idempotency, retries, backpressure, replay, dependency scheduling, and partial-failure recovery.Design partitioning and routing across tenants, workers, Kubernetes clusters, storage partitions, and regions.Define clear consistency, durability, and recovery guarantees across asynchronous pipelines and replicated object storage.Improve read-path performance and keep p95 and p99 latency predictable as datasets and workloads grow.Build multi-region replication, routing, failover, and recovery mechanisms.Detect and recover from queue buildup, retry storms, stale indexes, data drift, hotspots, noisy tenants, and silent correctness failures.Operate the same core data plane across our SaaS, BYOC, on-premises, and air-gapped deployments.Work directly with the founders and early customers to turn scale, reliability, and deployment constraints into product architecture.What we're looking forProduction experience building distributed systems, storage systems, databases, workflow engines, search or indexing infrastructure, or high-throughput backend systems.Experience owning stateful systems where correctness, latency, durability, and cost mattered simultaneously.Strong judgment around consistency, concurrency, persistence, caching, partitioning, replication, and failure recovery.Experience with queues and asynchronous execution, including retries, idempotency, ordering, replay, and backpressure.Hands-on experience with Kubernetes and cloud infrastructure.Experience debugging difficult production problems such as tail latency, queue buildup, retry storms, storage bottlenecks, data drift, hotspots, and partial failures.High ownership and comfort turning ambiguous systems problems into explicit invariants, guarantees, and measurable outcomes.We care more about demonstrated systems judgment than a particular number of years of experience.Nice to haveExperience with Rust, Go, C++, or another systems-oriented language.Experience with object-storage-native databases, distributed databases, event streaming, workflow orchestration, vector indexes, graph systems, or search infrastructure.Experience with multi-region systems and asynchronous replication.Experience supporting BYOC, on-premises, air-gapped, or regulated customer environments.Experience building infrastructure for AI agents or other data-intensive AI applications.Why ZerosetOwn core systems that directly determine the reliability, performance, and economics of the product.Help define a new architecture for durable agent state rather than maintaining a mature system with fixed assumptions.Work with a small, technical team with high trust, low bureaucracy, and direct access to customers and founders.Take on difficult storage and distributed-systems problems at the center of production enterprise AI.