JOBSEARCHER

Software Engineer, AI Runtime & Platform Services

About CrewAICrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production.The RoleYou'll work on the enterprise runtime layer that turns CrewAI's open-source Crews and Flows into secure, observable, remotely executable production systems. This is the layer between the framework and the platform: APIs, workers, checkpoints, webhooks, auth, deployment behavior, telemetry, and enterprise extensions that make CrewAI run reliably in real customer environments.You'll partner closely with the open-source, product, and infrastructure teams, but your center of gravity is production execution: making agent workflows resumable, inspectable, authenticated, observable, and safe to operate at scale.What You'll DoBuild and maintain the Python enterprise runtime around CrewAI: FastAPI services, Celery workers, Redis-backed state, execution APIs, and deployment-facing toolsExtend open-source CrewAI behavior for enterprise environments while preserving compatibility with upstream framework changesOwn production execution flows: crew and flow kickoff, status, retries, cancellation, checkpoint restore and fork, chat/session state, and human-in-the-loop resume pathsBuild secure integration surfaces: JWT auth, signed webhooks, token refresh, file handling, secret fetching, and workload identity across AWS, GCP, and AzureImprove observability across distributed execution: OpenTelemetry traces, structured logs, Sentry, event tracking, and debuggability across API, worker, and platform boundariesMaintain strong test coverage for async/runtime behavior using pytest, mypy, ruff, mocks/fakes, and e2e deployment harnessesPartner with the Agent Management Platform team on API contracts, versioning, enterprise client behavior, deployment status, and failure reportingRequirementsWhat We're Looking ForStrong Python backend/platform engineering experience, especially building production services rather than only librariesExperience with FastAPI or similar API frameworks, Celery or other job systems, Redis, Pydantic, and typed PythonGood instincts for distributed systems: retries, idempotency, async execution, status tracking, race conditions, and failure recoveryComfort with auth and security-sensitive systems: JWTs, webhooks, signatures, secrets, IAM/workload identity, and least-privilege thinkingPractical observability experience: tracing, structured logging, metrics, Sentry/OpenTelemetry, and debugging multi-service failuresAbility to work at the boundary between an open-source framework and a hosted enterprise platform without creating brittle couplingStrong testing habits and comfort with CI, package/version management, and release disciplineBonusExperience operating AI/agent runtimes, workflow engines, or distributed task systemsCloud platform experience with AWS ECS/ECR, Kubernetes, Helm, GCP/Azure identity, or secret managersExperience with enterprise SaaS constraints: auditability, tenant isolation, customer environments, deployment rollbacks, and supportabilityFamiliarity with Rails/SaaS platforms is useful, but not required