DevOps Engineer
DevOps EngineerWe're hiring a DevOps Engineer to own and improve the infrastructure that Dili's product and engineering teams depend on. You'll work across AWS, Terraform, GitHub Actions, observability, security, and developer tooling. Your job will be to make it easier and safer for engineers to ship changes, operate production systems, investigate problems, and scale the platform as Dili grows. This is a hands-on engineering role at an early-stage company. You'll build infrastructure, automate operational work, respond to production issues, improve system visibility, and help establish practical standards for reliability and security. You should be comfortable moving between planned infrastructure projects and immediate production needs. You'll work closely with application engineers and the CTO, but you'll be expected to develop your own view of where the platform needs improvement and take ownership of getting that work done.What You'll DoOwn and improve Dili's production infrastructure across AWS, including ECS Fargate, RDS, S3, load balancing, networking, IAM, and related services.Manage infrastructure as code using Terraform and ensure infrastructure changes are reviewable, repeatable, and safe.Build and maintain CI/CD pipelines in GitHub Actions for backend, frontend, infrastructure, migrations, and automated testing.Improve observability across services using logs, metrics, traces, dashboards, and alerts.Investigate production issues, identify root causes, and implement lasting fixes rather than relying on manual intervention.Improve deployment safety through automated checks, progressive rollout practices, rollback procedures, and clear operational visibility.Partner with engineers to improve local development, testing, release workflows, and overall developer experience.Monitor system capacity, reliability, and cost, and make practical improvements as usage grows.Help define incident response, disaster recovery, backup, and business continuity practices.Participate in a shared on-call rotation and help improve the systems and processes that make on-call manageable.What Success Looks LikeFirst 30 DaysLearn Dili's architecture, deployment process, cloud infrastructure, security controls, and operational workflows.Deploy changes through the existing pipeline, review recent incidents and recurring operational pain points, and begin contributing improvements to infrastructure or developer tooling.By 60 DaysTake ownership of a meaningful infrastructure area, such as CI/CD, observability, deployment reliability, cloud security, or production improvements that reduce manual work, improve visibility, or make releases safer and faster.By 90 DaysComplete a significant platform improvement with measurable results, such as faster deployments, fewer production issues, improved recovery time, stronger security controls, or reduced infrastructure cost.Establish a prioritized infrastructure roadmap and become the person the engineering team relies on for production reliability and platform decisions.QualificationsExperience building and operating production infrastructure for a software product.Strong experience with AWS and a practical understanding of cloud networking, compute, storage, databases, and identity management.Experience managing infrastructure through Terraform or another infrastructure-as-code tool.Experience building and maintaining CI/CD pipelines.Comfortable working with containers and containerized production workloads.Strong understanding of production observability, incident investigation, and root-cause analysis.Able to automate operational work using Python, shell scripting, or another programming language.Comfortable working directly with application engineers and understanding how infrastructure decisions affect product development.Strong security judgmentAble to turn ambiguous operational problems into practical, maintainable solutions.Clear, concise written and verbal and Working StyleCompetitive salary and equity.Meaningful ownership over Dili's infrastructure and engineering access to founders and company decision-making.Medical, dental, and vision coverage.Company-funded laptop and home-office setup.Flexible PTO.A hybrid culture based in our Union Square office, with four in-office days per week.A small engineering team with direct communication, lightweight process, and a strong focus on maintainability.