Lead Platform & Reliability Engineer
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
Overview
This is a high-impact player-coach position. You will own the entire Azure platform, ensuring production reliability, and lead the strategy to make our engineers 3–5 times more productive through AI-augmented workflows. There is potential to transition into an Engineering Manager as we scale the team and the need arises.
Responsibilities
Own the Azure cloud platform end-to-end (App Services, Azure SQL, Redis, Front Door, etc.)
Own, build, and evolve the entire Infrastructure as Code (IaC) pipeline using Bicep (primary) and/or Terraform — including AKS clusters, networking, Azure SQL/Redis configurations, and migration of existing App Services to containerized deployments
Ruthlessly drive reliability through AI-powered observability, anomaly detection, and self-healing systems
Build and maintain AI-assisted tooling for log analysis, root-cause investigation, automated remediation, and “smart on-call” triage
Establish chaos engineering, SLOs/SLIs, and incident response practices — with heavy use of AI for faster MTTD/MTTR
Lead blameless post-mortems, write clear incident reports, and communicate reliability status/updates to technical and non-technical stakeholders (including leadership)
Stay deeply hands-on: write and review production C#/.NET 8 and React/Next.js + TypeScript code daily
Lead and participate in on-call rotation (compensated with on-call stipend when implemented)
Qualifications
7+ years software/platform engineering experience
4+ years running proaC (Bicep/Terraform) and CI/CD (Azure DevOps or GitHub Actions)
Proven track record of identifying and eliminating operational toil through automation, Infrastructure as Code (IaC), AI-assisted tooling, and system redesign
Excellent verbal and written communication skills, with proven ability to explain complex technical concepts to diverse audiences and mentor engineers effectively.
Previous people management or tech lead experience (to grow into a manager role)
Proven experience designing and operating event-driven systems at scale (Azure Service Bus, Event Grid, Kafka, or equivalent) in production — this is a core part of our 2026 initiative
Track record of moving monolithic or request/response workloads to asynchronous, resilient event-driven patterns with idempotency, replay, and schema governance
Strong coding ability in C#/.NET and modern React/Next.js + TypeScript
Security hardening and compliance experience (SOC 2, HIPAA, or state education data privacy)
Strongly Preferred
Built or integrated tools that use Azure OpenAI for log summarization, runbook automation, or anomaly detection
What We Offer
Competitive salary based on experience.
Standard benefits, including health, dental, and vision insurance, 401(k) with match, and paid time off.
A small family-owned company culture with a collaborative and innovative work environment.
Opportunities for professional growth and leadership development.
Join us in shaping the future of our software solutions while leading a talented team dedicated to excellence!
Job Type: Full-time
Base Pay: From $130,000.00 per year
Benefits:
401(k)
401(k) 4% Match
Bereavement leave
Dental insurance
Employee assistance program
Flexible spending account
Free parking
Health insurance
Health savings account
Life insurance
Paid time off
Retirement plan
Vision insurance
Experience:
Software/platform engineering: 7 years (Required)
Running production workloads on Microsoft Azure at scale: 4 years (Required)
Ability to Commute:
Georgetown, TX 78628 (Required)
Work Location: In person