{"schemaVersion":"jobsearcher.job.v1","id":"4dd31a05c5a16bd8b7d8726b","url":"https://jobsearcher.com/jobs/4dd31a05c5a16bd8b7d8726b","canonicalUrl":"https://jobsearcher.com/jobs/4dd31a05c5a16bd8b7d8726b","title":"Software Engineer - Infrastructure","description":"Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at global scale and are used to build millions of real applications.\nSince our public launch, we've crossed $100M in ARR and grown to over 10M users across 190+ countries. We're backed by Khosla Ventures, SoftBank, Google, Lightspeed, Prosus, Together, and Y Combinator.\nWe're solving the hard part of AI-driven software creation: correctness, reliability, security, and scale in real production systems. The team is built by repeat founders, Olympiad medalists, IIT & IIM alumni, and leaders from Google, Amazon, and Dropbox.\nWe're hiring builders who want ownership, speed, and impact at global scale.\nWhat You'll Do:\nPlatform and Infrastructure\nMaintain stability of our platform consisting of distributed microservices closely interacting with Kubernetes and cloud providers (GCP, AWS)\nManage Kubernetes workloads with ArgoCD (GitOps), deploy, monitor, and troubleshoot application syncs, resource trees, and rollouts\nDebug and resolve complex Kubernetes issues across clusters\nManage CDN and edge infrastructure (Cloudflare) for performance, caching, and traffic management\nAutomate infrastructure lifecycle operations and workflows\nObservability and Incident Response\nOwn the observability stack: Grafana (dashboards, Loki logs, Prometheus metrics) and New Relic (APM, golden metrics, transaction analysis)\nEnhance monitoring, alerting, and distributed tracing across services\nParticipate in on-call rotation via PagerDuty, handle incident response, and perform root cause analysis\nProactively identify reliability risks before they become incidents\nAI Agent Infrastructure\nSupport the platform that runs AI agent workloads including job scheduling, trajectory tracking, environment provisioning, deployments, and cost attribution\nDevelop Kubernetes controllers and operators to extend platform capabilities for agent orchestration\nCollaboration and Internal Tooling\nWork closely with product and backend teams to ensure platform scalability and reliability\nBuild internal tools, automate workflows, and integrate systems to improve team productivity\nStay current with Kubernetes releases, CNCF ecosystem updates, and cloud-native best practices\nWhat We're Looking For:\nCore Requirements\n3+ years of software/platform engineering experience with production systems\nStrong proficiency in Go or Python, you write production code in at least one daily\nHands-on experience building and deploying services on Kubernetes, not just YAML, you've developed something that runs on K8s\nExperience with GitOps tooling (ArgoCD, Flux, or similar)\nSystems Fundamentals\nStrong networking and DNS fundamentals: TCP/IP, HTTP, load balancing, DNS resolution, TLS, and debugging connectivity issues\nSolid Linux/OS fundamentals: process management, filesystem, memory, systemd, and comfortable debugging with tools like strace, tcpdump, and netstat\nData and Messaging Infrastructure\nRelational databases: experience with PostgreSQL, MySQL, or similar; indexing, query optimization, replication, and backup/restore procedures\nNoSQL databases: familiarity with MongoDB, DynamoDB, Redis, or similar for document/key-value workloads\nCaching: experience with Redis, Memcached, or similar for application and infrastructure-level caching\nMessage queues and streaming: hands-on with Kafka, SQS, RabbitMQ, or similar for event-driven architectures\nStrong SQL skills for debugging and operational queries\nInfrastructure and Observability\nComfortable with the CNCF ecosystem: Helm, Kustomize, cert-manager, Ingress controllers, CNI/CSI interfaces\nHands-on with at least one observability stack (Grafana/Prometheus/Loki, New Relic, Datadog, or similar)\nFamiliarity with GCP and/or AWS: managed Kubernetes (GKE/EKS), networking, IAM, storage, and cloud-native services (SES, SQS, S3, etc.)\nExperience with CDN/edge platforms (Cloudflare, CloudFront, or similar)\nGood to Have:\nExperience building Kubernetes Operators (kubebuilder, operator-sdk, or controller-runtime)\nExperience tuning Kubernetes core components (API server, kubelet, scheduler)\nFamiliarity with AI/LLM infrastructure: token management, cost tracking, agent orchestration\nExperience with CI/CD pipelines (GitHub Actions, automated testing, deployment pipelines)\nInfrastructure as Code experience (Terraform, Pulumi, or similar)\nPrevious work on large-scale distributed systems or platform-as-a-service\nStartup experience, you thrive in fast-paced, ambiguous environments\nWho You Are:\nA generalist who can context-switch between debugging a K8s deployment, setting up a Grafana alert, and configuring CDN rules, all in the same day\nYou enjoy solving complex infrastructure challenges and automating away toil\nYou dig deep, when something breaks, you find the root cause, not just the workaround\nYou communicate clearly and can collaborate effectively in a fast-moving, distributed team\nTech Stack:\nWe don't require previous experience with our entire stack, but enthusiasm for learning is key: Go, Python, Kubernetes, ArgoCD, Helm, GCP, AWS, Cloudflare, Grafana, Prometheus, Loki, New Relic, PagerDuty, PostgreSQL, MongoDB, Redis, Kafka, and GitHub.\n\nBenefits and Perks:\n401(k)\nHealth, dental, and vision insurance\nUnlimited Paid Time Off: take the time you need to recharge and come back refreshed\nFlexible Working Hours: work arrangements that fit your life and commitments\nLet's build the future of software together.","company":"Emergent Labs","rawCompany":"emergent labs","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-04T17:04:29.076Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer - Infrastructure","description":"Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at global scale and are used to build millions of real applications.\nSince our public launch, we've crossed $100M in ARR and grown to over 10M users across 190+ countries. We're backed by Khosla Ventures, SoftBank, Google, Lightspeed, Prosus, Together, and Y Combinator.\nWe're solving the hard part of AI-driven software creation: correctness, reliability, security, and scale in real production systems. The team is built by repeat founders, Olympiad medalists, IIT & IIM alumni, and leaders from Google, Amazon, and Dropbox.\nWe're hiring builders who want ownership, speed, and impact at global scale.\nWhat You'll Do:\nPlatform and Infrastructure\nMaintain stability of our platform consisting of distributed microservices closely interacting with Kubernetes and cloud providers (GCP, AWS)\nManage Kubernetes workloads with ArgoCD (GitOps), deploy, monitor, and troubleshoot application syncs, resource trees, and rollouts\nDebug and resolve complex Kubernetes issues across clusters\nManage CDN and edge infrastructure (Cloudflare) for performance, caching, and traffic management\nAutomate infrastructure lifecycle operations and workflows\nObservability and Incident Response\nOwn the observability stack: Grafana (dashboards, Loki logs, Prometheus metrics) and New Relic (APM, golden metrics, transaction analysis)\nEnhance monitoring, alerting, and distributed tracing across services\nParticipate in on-call rotation via PagerDuty, handle incident response, and perform root cause analysis\nProactively identify reliability risks before they become incidents\nAI Agent Infrastructure\nSupport the platform that runs AI agent workloads including job scheduling, trajectory tracking, environment provisioning, deployments, and cost attribution\nDevelop Kubernetes controllers and operators to extend platform capabilities for agent orchestration\nCollaboration and Internal Tooling\nWork closely with product and backend teams to ensure platform scalability and reliability\nBuild internal tools, automate workflows, and integrate systems to improve team productivity\nStay current with Kubernetes releases, CNCF ecosystem updates, and cloud-native best practices\nWhat We're Looking For:\nCore Requirements\n3+ years of software/platform engineering experience with production systems\nStrong proficiency in Go or Python, you write production code in at least one daily\nHands-on experience building and deploying services on Kubernetes, not just YAML, you've developed something that runs on K8s\nExperience with GitOps tooling (ArgoCD, Flux, or similar)\nSystems Fundamentals\nStrong networking and DNS fundamentals: TCP/IP, HTTP, load balancing, DNS resolution, TLS, and debugging connectivity issues\nSolid Linux/OS fundamentals: process management, filesystem, memory, systemd, and comfortable debugging with tools like strace, tcpdump, and netstat\nData and Messaging Infrastructure\nRelational databases: experience with PostgreSQL, MySQL, or similar; indexing, query optimization, replication, and backup/restore procedures\nNoSQL databases: familiarity with MongoDB, DynamoDB, Redis, or similar for document/key-value workloads\nCaching: experience with Redis, Memcached, or similar for application and infrastructure-level caching\nMessage queues and streaming: hands-on with Kafka, SQS, RabbitMQ, or similar for event-driven architectures\nStrong SQL skills for debugging and operational queries\nInfrastructure and Observability\nComfortable with the CNCF ecosystem: Helm, Kustomize, cert-manager, Ingress controllers, CNI/CSI interfaces\nHands-on with at least one observability stack (Grafana/Prometheus/Loki, New Relic, Datadog, or similar)\nFamiliarity with GCP and/or AWS: managed Kubernetes (GKE/EKS), networking, IAM, storage, and cloud-native services (SES, SQS, S3, etc.)\nExperience with CDN/edge platforms (Cloudflare, CloudFront, or similar)\nGood to Have:\nExperience building Kubernetes Operators (kubebuilder, operator-sdk, or controller-runtime)\nExperience tuning Kubernetes core components (API server, kubelet, scheduler)\nFamiliarity with AI/LLM infrastructure: token management, cost tracking, agent orchestration\nExperience with CI/CD pipelines (GitHub Actions, automated testing, deployment pipelines)\nInfrastructure as Code experience (Terraform, Pulumi, or similar)\nPrevious work on large-scale distributed systems or platform-as-a-service\nStartup experience, you thrive in fast-paced, ambiguous environments\nWho You Are:\nA generalist who can context-switch between debugging a K8s deployment, setting up a Grafana alert, and configuring CDN rules, all in the same day\nYou enjoy solving complex infrastructure challenges and automating away toil\nYou dig deep, when something breaks, you find the root cause, not just the workaround\nYou communicate clearly and can collaborate effectively in a fast-moving, distributed team\nTech Stack:\nWe don't require previous experience with our entire stack, but enthusiasm for learning is key: Go, Python, Kubernetes, ArgoCD, Helm, GCP, AWS, Cloudflare, Grafana, Prometheus, Loki, New Relic, PagerDuty, PostgreSQL, MongoDB, Redis, Kafka, and GitHub.\n\nBenefits and Perks:\n401(k)\nHealth, dental, and vision insurance\nUnlimited Paid Time Off: take the time you need to recharge and come back refreshed\nFlexible Working Hours: work arrangements that fit your life and commitments\nLet's build the future of software together.","datePosted":"2026-08-04T17:04:29.076Z","dateModified":"2026-08-04T17:04:29.076Z","hiringOrganization":{"@type":"Organization","name":"Emergent Labs","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"4dd31a05c5a16bd8b7d8726b"},"url":"https://jobsearcher.com/jobs/4dd31a05c5a16bd8b7d8726b"}}