JOBSEARCHER

Staff AI Platform & Reliability Engineer

Location: Santa Ana, California — on-siteEmployment type: Full-timeSalary: $172,000 - $220,000About UsCreated.ai is an AI-powered creative platform combining text, image and video generation into production tools for commercial creative work. The platform is operated by eJam, a multi-brand consumer products and technology company.We are expanding our engineering team with two appointments covering our AI platform infrastructure and our product engineering surface. Both roles carry substantial technical ownership and report directly into engineering leadership.This position is AI-native by design. Our engineers work with agentic coding tools as a primary part of their daily workflow, which allows a small team to operate at a scope that would conventionally require a much larger one. We are looking for engineers whose judgement, architectural thinking and review discipline scale that leverage rather than being replaced by it.Role SummaryWe are seeking a Staff Engineer to take ownership of the AI generation services, provider integrations, billing integrity, tenant security and overall platform reliability underpinning created.ai. The role is Python and GCP-first, with sufficient TypeScript proficiency to trace and modify cross-service contracts.This is a senior individual contributor position with architectural authority over the platform layer.Key responsibilitiesOwn the design, delivery and operation of AI generation services and all third-party provider integrationsEnsure billing and usage-metering integrity across the platform, including reconciliation and credit accountingMaintain tenant isolation and platform security controlsLead reliability engineering: failure handling, degradation strategy, capacity and incident responseStrengthen deployment safeguards, observability and cost attributionSet technical standards for the Python services and mentor engineers working within themRequirementsPython 3.12 with FastAPI, Pydantic, asyncio and strict typing in productionGoogle Cloud Platform: Cloud Run, Pub/Sub, Cloud Tasks, GCS, Firestore, Cloud SQL/PostgresDemonstrated experience with webhooks, queues, retries, idempotency, dead-letter queues and durable asynchronous jobsSQLAlchemy and Alembic, together with billing or usage-metering experienceInfrastructure and delivery tooling: Terraform, IAM/OIDC, Secret Manager, CI/CDProduction integrations with image, video or LLM providersWorking proficiency in NestJS/TypeScript sufficient to modify cross-service contractsStrong background in observability, cost tracking and incident debuggingFluency with agentic coding tools (Claude Code, Codex, Cursor or comparable), including the ability to scope work for them, review their output critically and maintain architectural coherence across AI-assisted changesBenefitsHealth, Dental, Vision401k PlanPTO Plan14 observed local holidaysStock optionsAmazing, pet-friendly office environmentEquipment budget and learning allowanceReal technical ownership within a small engineering team, and visible impact on a commercially active AI product with real users and real scale constraintsDirect involvement in product decisions — the team is small enough that there is nowhere to hide, in both directions