Software Development Engineer
Key ResponsibilitiesArchitect and implement Python agent frameworks: tool adapters, orchestration loops, structured outputs, session handling, and CLI or service interfaces suitable for production use. Build LLM-powered workflows (LiteLLM, DSPy, or equivalent) with bounded prompts, allowlisted tools, budget caps, structured JSON outputs, and evidence-backed reporting. Integrate agents with the AMD GPU and ML software ecosystem: ROCm tooling,PyTorch, distributed training and inference stacks, profilers, cluster metrics, log pipelines, and debug utilities. Own quality and eval infrastructure: fixture-driven tests, JSON Schema validation, regression suites, accuracy metrics, and CI integration before rollout. Solve ML operations problems on training and inference clusters: multi-node failure analysis, performance regression investigation, configuration and launch issues, and operational automation. Deliver deployment-ready artifacts: Docker / Kubernetes packaging, runbooks, documentation, and security-conscious defaults for customer-controlled environments. Partner with framework, performance, and field teams; incorporate code review feedback and iterate based on production and pilot learnings. Required qualificationsBachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field; equivalent practical experience accepted. 10+ years of professional software engineering experience, including 5+ years on production Python systems at scale. Deep experience with distributed ML training or inference on GPU clusters: multi-node jobs, collective communication failures, log and metric correlation across ranks and nodes. Strong hands-on PyTorch background and production familiarity with large-scale stacks (Megatron-LM, DeepSpeed, TorchTitan, vLLM, or equivalent). Proven ML debugging and performance engineering across multiple failure modes: throughput regression, memory errors, numerical instability, misconfiguration, and profiler or trace analysis. Demonstrated delivery of LLM agent or orchestration systems in production or nearproduction settings: tool routing, structured outputs, reliability, and testability. Track record of shipping complex software on schedule: clean code, automated tests, code review discipline, and clear technical communication. Ability to work independently, prioritize across ambiguous requirements, and align weekly with a technical lead. Preferred QualificationsOpen-source ML development: upstream framework repos, community CI patterns, and integration with OSS tooling. ML performance optimization: parallelism, throughput/MFU tuning, roofline analysis, and profiler-driven investigation. ROCm and GPU cluster operations: metrics, profilers, health monitoring, Slurm or Kubernetes job environments. Security-aware agent deployment: secret handling, egress control, and customer VPC or on-prem constraints.