Software Development Engineer 5
Software Development Engineer 5 – Agentic ML Infrastructure AMD is building agentic software to improve how teams develop, deploy, and operate machine learning workloads on AMD GPUs.This senior contract role focuses on designing and delivering production-grade intelligent agents across the AMD software stack, including:Training and inference frameworksCluster toolingPerformance analysisOperational workflows for large-scale GPU deploymentsThe role will lead implementation of Python systems that combine LLM orchestration, tool integration, and rigorous evaluation to automate high-value engineering work, including:Investigating distributed job failuresInterpreting telemetry and profilersAccelerating performance analysisReducing manual toil in ML operationsThis is a senior individual contributor engagement. The engineer is expected to own significant components end-to-end, make sound technical decisions within program direction, and deliver maintainable code with minimal supervision.Key Responsibilities Agent Framework & LLM Engineering Architect and implement Python agent frameworks, including:Tool adaptersOrchestration loopsStructured outputsSession handlingCLI or service interfaces suitable for production useBuild LLM-powered workflows using LiteLLM, DSPy, or equivalent.Implement bounded prompts, allowlisted tools, budget caps, structured JSON outputs, and evidence-backed reporting.ML & GPU Infrastructure Integrate agents with the AMD GPU and ML software ecosystem, including:ROCm toolingPyTorchDistributed training and inference stacksProfilersCluster metricsLog pipelinesDebug utilitiesSolve ML operations problems on training and inference clusters, including:Multi-node failure analysisPerformance regression investigationConfiguration and launch issuesOperational automationQuality, Testing & Evaluation Own agent quality and evaluation infrastructure.Develop fixture-driven tests and regression suites.Implement JSON Schema validation and accuracy metrics.Integrate testing and evaluation into CI before rollout.Deployment & Documentation Deliver deployment-ready artifacts, including:Docker packagingKubernetes packagingRunbooksTechnical documentationApply security-conscious defaults for customer-controlled environments.Collaboration & Delivery Partner with framework, performance, and field teams.Incorporate code review feedback and iterate based on production and pilot learnings.Own significant technical components independently and deliver production-quality solutions with minimal supervision.Required Qualifications Education & Experience Bachelor's degree or higher in:Computer ScienceComputer EngineeringElectrical EngineeringRelated fieldEquivalent practical experience accepted.10+ years of professional software engineering experience.5+ years of experience developing production Python systems at scale.Distributed ML & GPU Systems Deep experience with distributed ML training or inference on GPU clusters.Experience with:Multi-node jobsCollective communication failuresLog and metric correlation across ranks and nodesStrong hands-on PyTorch experience.Production familiarity with large-scale ML stacks such as:Megatron-LMDeepSpeedTorchTitanvLLMEquivalent technologiesML Debugging & Performance Engineering Proven ML debugging and performance engineering experience across multiple failure modes, including:Throughput regressionMemory errorsNumerical instabilityMisconfigurationProfiler or trace analysisLLM Agents & Software Engineering Demonstrated experience delivering LLM agent or orchestration systems in production or near-production environments.Experience with:Tool routingStructured outputsReliabilityTestabilityTrack record of delivering complex software on schedule.Strong capabilities in:Clean, maintainable codeAutomated testingCode reviewTechnical communicationAbility to work independently, prioritize ambiguous requirements, and align weekly with a technical lead.Preferred Qualifications Open-source ML development, including:Upstream framework repositoriesCommunity CI patternsOSS tooling integrationML performance optimization, including:ParallelismThroughput/MFU tuningRoofline analysisProfiler-driven investigationROCm and GPU cluster operations, including:MetricsProfilersHealth monitoringSlurmKubernetes job environmentsSecurity-aware agent deployment, including:Secret handlingEgress controlCustomer VPC constraintsOn-premises deployment constraintsWorking Model Duration: 12 monthsSchedule: Full-time contractCadence: Weekly sync with technical leadDelivery Model: Milestone-drivenScope: Implementation across AMD agent and ML infrastructure programsPriorities: Allocation may shift based on program priorities at the direction of the technical leadWhat This Role Is Not This position is not focused on:Graphics driver developmentKernel developmentLow-level GPU driver developmentResearch-only work without production deliverablesJunior or mid-level implementationSDE 5 expectations: Senior ownership, strong technical depth, and demonstrated expertise across ML systems and agent engineering.