Software Engineer, Infrastructure & Reliability
About CrewAICrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.The RoleYou'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You'll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer's production environments safer.This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.What You'll DoOwn and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related servicesBuild and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safetyImprove reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAsPartner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behaviorManage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systemsHarden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege accessBuild the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over timeReduce operational toil by automating recurring workflows and making deployments boringRequirementsWhat We're Looking ForStrong infrastructure/platform engineering experience in production SaaS environmentsDeep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized servicesExperience with ECS and/or Kubernetes; Helm experience is a strong plusComfort operating PostgreSQL, Redis, background job systems, queues, and web services in productionStrong debugging instincts across app, infra, network, deploy, and dependency layersSecurity-minded approach to IAM, secrets, workload identity, vulnerability management, and production accessAbility to write reliable automation in Python, Ruby, Go, Bash, or similarCalm, rigorous approach to incidents, rollbacks, migrations, and production change managementBonusExperience with AI/agent platforms, workflow runtimes, or high-volume async execution systemsExperience supporting enterprise/self-hosted deploymentsTerraform or other IaC experienceSRE background: SLOs, incident review, capacity planning, load testingFamiliarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability