{"schemaVersion":"jobsearcher.job.v1","id":"9aaac4eb0d66f59f5eb1481e","url":"https://jobsearcher.com/jobs/9aaac4eb0d66f59f5eb1481e","canonicalUrl":"https://jobsearcher.com/jobs/9aaac4eb0d66f59f5eb1481e","title":"Software Development Engineer","description":"Key ResponsibilitiesArchitect and implement Python agent frameworks: tool adapters, orchestration loops, structured outputs, session handling, and CLI or service interfaces suitable for production use. Build LLM-powered workflows (LiteLLM, DSPy, or equivalent) with bounded prompts, allowlisted tools, budget caps, structured JSON outputs, and evidence-backed reporting. Integrate agents with the AMD GPU and ML software ecosystem: ROCm tooling,PyTorch, distributed training and inference stacks, profilers, cluster metrics, log pipelines, and debug utilities. Own quality and eval infrastructure: fixture-driven tests, JSON Schema validation, regression suites, accuracy metrics, and CI integration before rollout. Solve ML operations problems on training and inference clusters: multi-node failure analysis, performance regression investigation, configuration and launch issues, and operational automation. Deliver deployment-ready artifacts: Docker / Kubernetes packaging, runbooks, documentation, and security-conscious defaults for customer-controlled environments. Partner with framework, performance, and field teams; incorporate code review feedback and iterate based on production and pilot learnings. Required qualificationsBachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field; equivalent practical experience accepted. 10+ years of professional software engineering experience, including 5+ years on production Python systems at scale. Deep experience with distributed ML training or inference on GPU clusters: multi-node jobs, collective communication failures, log and metric correlation across ranks and nodes. Strong hands-on PyTorch background and production familiarity with large-scale stacks (Megatron-LM, DeepSpeed, TorchTitan, vLLM, or equivalent). Proven ML debugging and performance engineering across multiple failure modes: throughput regression, memory errors, numerical instability, misconfiguration, and profiler or trace analysis. Demonstrated delivery of LLM agent or orchestration systems in production or nearproduction settings: tool routing, structured outputs, reliability, and testability. Track record of shipping complex software on schedule: clean code, automated tests, code review discipline, and clear technical communication. Ability to work independently, prioritize across ambiguous requirements, and align weekly with a technical lead. Preferred QualificationsOpen-source ML development: upstream framework repos, community CI patterns, and integration with OSS tooling. ML performance optimization: parallelism, throughput/MFU tuning, roofline analysis, and profiler-driven investigation. ROCm and GPU cluster operations: metrics, profilers, health monitoring, Slurm or Kubernetes job environments. Security-aware agent deployment: secret handling, egress control, and customer VPC or on-prem constraints.","company":"Saicon","rawCompany":"saicon","city":"California","state":"MO","isRemote":false,"isActive":false,"createdAt":"2026-08-29T08:43:21.197Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1251.00","title":"Computer Programmers","slug":"computer-programmers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Development Engineer","description":"Key ResponsibilitiesArchitect and implement Python agent frameworks: tool adapters, orchestration loops, structured outputs, session handling, and CLI or service interfaces suitable for production use. Build LLM-powered workflows (LiteLLM, DSPy, or equivalent) with bounded prompts, allowlisted tools, budget caps, structured JSON outputs, and evidence-backed reporting. Integrate agents with the AMD GPU and ML software ecosystem: ROCm tooling,PyTorch, distributed training and inference stacks, profilers, cluster metrics, log pipelines, and debug utilities. Own quality and eval infrastructure: fixture-driven tests, JSON Schema validation, regression suites, accuracy metrics, and CI integration before rollout. Solve ML operations problems on training and inference clusters: multi-node failure analysis, performance regression investigation, configuration and launch issues, and operational automation. Deliver deployment-ready artifacts: Docker / Kubernetes packaging, runbooks, documentation, and security-conscious defaults for customer-controlled environments. Partner with framework, performance, and field teams; incorporate code review feedback and iterate based on production and pilot learnings. Required qualificationsBachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field; equivalent practical experience accepted. 10+ years of professional software engineering experience, including 5+ years on production Python systems at scale. Deep experience with distributed ML training or inference on GPU clusters: multi-node jobs, collective communication failures, log and metric correlation across ranks and nodes. Strong hands-on PyTorch background and production familiarity with large-scale stacks (Megatron-LM, DeepSpeed, TorchTitan, vLLM, or equivalent). Proven ML debugging and performance engineering across multiple failure modes: throughput regression, memory errors, numerical instability, misconfiguration, and profiler or trace analysis. Demonstrated delivery of LLM agent or orchestration systems in production or nearproduction settings: tool routing, structured outputs, reliability, and testability. Track record of shipping complex software on schedule: clean code, automated tests, code review discipline, and clear technical communication. Ability to work independently, prioritize across ambiguous requirements, and align weekly with a technical lead. Preferred QualificationsOpen-source ML development: upstream framework repos, community CI patterns, and integration with OSS tooling. ML performance optimization: parallelism, throughput/MFU tuning, roofline analysis, and profiler-driven investigation. ROCm and GPU cluster operations: metrics, profilers, health monitoring, Slurm or Kubernetes job environments. Security-aware agent deployment: secret handling, egress control, and customer VPC or on-prem constraints.","datePosted":"2026-08-29T08:43:21.197Z","dateModified":"2026-08-29T08:43:21.197Z","hiringOrganization":{"@type":"Organization","name":"Saicon","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"California","addressRegion":"MO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"9aaac4eb0d66f59f5eb1481e"},"url":"https://jobsearcher.com/jobs/9aaac4eb0d66f59f5eb1481e"}}