Site Reliability Engineer - Observability
Site Reliability Engineer (Observability SME)Location: Hybrid (2–3 days onsite/week) in Atlanta, GAPosition Type: Full-Time, Direct-HireTarget Compensation: $135,000 / yearPosition OverviewWe are seeking a hungry, execution-focused Site Reliability Engineer (SRE) with a strong background in software development and deep expertise in Dynatrace and Azure cloud observability.In this role, you will serve as the primary Subject Matter Expert (SME) in Observability and Monitoring, taking full ownership of our enterprise platform reliability strategy. You won't just look at dashboards—you will dive into telemetry (logs, metrics, traces), perform data mining in Dynatrace and Azure App Insights, identify the exact root cause in serverless .NET code, and write/implement the fixes.If you are an energetic SRE or a Cloud Developer who uses advanced observability tools to diagnose and patch code at scale, this is your opportunity to drive modern SRE practices, AI-assisted workflows, and automation across high-volume digital platforms.What Makes This Role UniqueActive Code Fixes: This is a hybrid SRE/Dev role. You will use observability to answer how something failed, read/analyze .NET code, and execute the actual fix in the codebase. AI & Automation Focus: Extensive opportunity to leverage AI tooling and Dynatrace Davis AI to automate root-cause analysis, build custom workflows, and drive touchless operations. High Impact & Autonomy: You will lead the SRE strategy, build out the backlog, and partner directly with cloud and engineering teams without rigid, legacy playbooks. Key ResponsibilitiesDynatrace & Observability Ownership: Serve as the hands-on SME for Dynatrace and Dynatrace Cloud. Build advanced queries (DQL), custom dashboards, alerts, and automated workflows. Go beyond basic agent installs to drive deep telemetry and data mining. Azure Cloud & Serverless Reliability: Manage, monitor, and troubleshoot serverless components in Azure, specifically Azure Function Apps, App Services, APIM, and backend APIs using Application Insights, Azure Monitor, and KQL. Code-Level Root Cause Analysis: Read and analyze .NET application code to identify bugs, latency bottlenecks, and unhandled exceptions, writing code updates and deploying fixes. Mobile & API Observability: Extend telemetry and performance monitoring to mobile applications (iOS and Android) and end-to-end API transaction flows. Agile Execution & Backlog Delivery: Own feature and story backlog items, driving sprint commitments with high accountability, speed, and creative technical solutions. Core RequirementsQUALIFICATIONS & CRITICAL SKILLSProven Dynatrace Expertise: Deep, hands-on experience using Dynatrace beyond basic agent deployment—including writing DQL/queries, data mining, tuning Davis AI, and building proactive alert models. Software Development Background: Prior development experience in .NET with the ability to read, troubleshoot, and patch code independently. Azure Cloud & Serverless Experience: Practical experience supporting Azure Function Apps, Azure Monitor, Application Insights, Log Analytics, and KQL. Engineering Mindset: Highly proactive, curious, and hungry to solve complex architectural issues—someone who thrives on digging into complex systems rather than coasting. Nice-to-HavesExperience combining Dynatrace with Azure Application Insights or OpenTelemetry. Familiarity with mobile telemetry (iOS/Android) and API Management (APIM). Hands-on experience incorporating AI/LLM tools into everyday dev/SRE workflows.#IT123