{"schemaVersion":"jobsearcher.job.v1","id":"46bc3c43f69909b38fc53c61","url":"https://jobsearcher.com/jobs/46bc3c43f69909b38fc53c61","canonicalUrl":"https://jobsearcher.com/jobs/46bc3c43f69909b38fc53c61","title":"Software Engineer, Infrastructure & Reliability","description":"About CrewAI\nCrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.\nThe Role\nYou'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer’s production environments safer.\nThis is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.\nWhat You'll Do\nOwn and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.\nBuild and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.\nImprove reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.\nPartner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.\nManage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.\nHarden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.\nBuild the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.\nReduce operational toil by automating recurring workflows and making deployments boring.\n\nRequirements\n\nWhat We're Looking For\nStrong infrastructure/platform engineering experience in production SaaS environments.\nDeep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.\nExperience with ECS and/or Kubernetes; Helm experience is a strong plus.\nComfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.\nStrong debugging instincts across app, infra, network, deploy, and dependency layers.\nSecurity-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.\nAbility to write reliable automation in Python, Ruby, Go, Bash, or similar.\nCalm, rigorous approach to incidents, rollbacks, migrations, and production change management.\nBonus\nExperience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.\nExperience supporting enterprise/self-hosted deployments.\nTerraform or other IaC experience.\nSRE background: SLOs, incident review, capacity planning, load testing.\nFamiliarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.","company":"Crewai","rawCompany":"crewai","city":"Denver","state":"CO","isRemote":false,"isActive":false,"createdAt":"2026-08-14T13:22:55.794Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer, Infrastructure & Reliability","description":"About CrewAI\nCrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.\nThe Role\nYou'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer’s production environments safer.\nThis is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.\nWhat You'll Do\nOwn and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.\nBuild and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.\nImprove reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.\nPartner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.\nManage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.\nHarden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.\nBuild the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.\nReduce operational toil by automating recurring workflows and making deployments boring.\n\nRequirements\n\nWhat We're Looking For\nStrong infrastructure/platform engineering experience in production SaaS environments.\nDeep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.\nExperience with ECS and/or Kubernetes; Helm experience is a strong plus.\nComfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.\nStrong debugging instincts across app, infra, network, deploy, and dependency layers.\nSecurity-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.\nAbility to write reliable automation in Python, Ruby, Go, Bash, or similar.\nCalm, rigorous approach to incidents, rollbacks, migrations, and production change management.\nBonus\nExperience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.\nExperience supporting enterprise/self-hosted deployments.\nTerraform or other IaC experience.\nSRE background: SLOs, incident review, capacity planning, load testing.\nFamiliarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.","datePosted":"2026-08-14T13:22:55.794Z","dateModified":"2026-08-14T13:22:55.794Z","hiringOrganization":{"@type":"Organization","name":"Crewai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"46bc3c43f69909b38fc53c61"},"url":"https://jobsearcher.com/jobs/46bc3c43f69909b38fc53c61"}}