Software Development Engineer 5
Overview: Tek Wissen is a global workforce management provider headquartered in Ann Arbor, Michigan that offers strategic talent solutions to our clients worldwide. This Client is an American multinational semiconductor company based in Austin, that develops computer processors and related technologies for business and consumer markets. Global company that specializes in manufacturing semiconductor devices used in computer processing. The company also produces flash memories, graphics processors, motherboard chip sets, and a variety of components used in consumer electronics goods..Job Title: Software Development Engineer 5 - Agentic ML InfrastructureDuration: 6 Months to HireWork Location: Santa Clara, CAJob Type: Temporary AssignmentWork Type: OnsiteAbout the roleClient is building agentic software to improve how teams develop, deploy, and operate machine learning workloads on Client GPUs.This senior contract role focuses on designing and delivering production-grade intelligent agents across the CLIENT software stack: training and inference frameworks, cluster tooling, performance analysis, and operational workflows for large-scale GPU deployments.You will lead implementation of Python systems that combine LLM orchestration, tool integration, and rigorous evaluation to automate high-value engineering work: investigating distributed job failures, interpreting telemetry and profilers, accelerating performance analysis, and reducing manual toil in ML operations.This is a senior individual contributor engagement: you are expected to own significantcomponents end to end, make sound technical decisions within program direction, and deliver maintainable code with minimal supervision.Key responsibilities Architect and implement Python agent frameworks: tool adapters, orchestration loops, structured outputs, session handling, and CLI or service interfaces suitable for production use.Build LLM-powered workflows (LiteLLM, DSPy, or equivalent) with bounded prompts, allowlisted tools, budget caps, structured JSON outputs, and evidence-backed reporting.Integrate agents with the CLIENT GPU and ML software ecosystem: ROCm tooling, PyTorch, distributed training and inference stacks, profilers, cluster metrics, log pipelines, and debug utilities.Own quality and eval infrastructure: fixture-driven tests, JSON Schema validation, regression suites, accuracy metrics, and CI integration before rollout.Solve ML operations problems on training and inference clusters: multi-node failure analysis, performance regression investigation, configuration and launch issues, and operational automation.Deliver deployment-ready artifacts: Docker / Kubernetes packaging, runbooks, documentation, and security-conscious defaults for customer-controlled environments.Partner with framework, performance, and field teams; incorporate code review feedbackand iterate based on production and pilot learnings.Required qualifications Bachelor's degree or higher in Computer Science, Computer Engineering, ElectricalEngineering, or a related field; equivalent practical experience accepted.10+ years of professional software engineering experience, including 5+ years on production Python systems at scale.Deep experience with distributed ML training or inference on GPU clusters: multi-node jobs, collective communication failures, log and metric correlation across ranks and nodes.Strong hands-on PyTorch background and production familiarity with large-scale stacks (Megatron-LM, DeepSpeed, TorchTitan, vLLM, or equivalent).Proven ML debugging and performance engineering across multiple failure modes: throughput regression, memory errors, numerical instability, misconfiguration, and profiler or trace analysis.Demonstrated delivery of LLM agent or orchestration systems in production or near production settings: tool routing, structured outputs, reliability, and testability.Track record of shipping complex software on schedule: clean code, automated tests,code review discipline, and clear technical communication.Ability to work independently, prioritize across ambiguous requirements, and align weekly with a technical lead.Preferred qualifications Open-source ML development: upstream framework repos, community CI patterns, and integration with OSS tooling.ML performance optimization: parallelism, throughput/MFU tuning, roofline analysis,and profiler-driven investigation.ROCm and GPU cluster operations: metrics, profilers, health monitoring, Slurm orKubernetes job environments.Security-aware agent deployment: secret handling, egress control, and customer VPC or on-prem constraints.Working model Cadence: weekly sync with technical lead; milestone-driven deliveryScope: implementation across CLIENT agent and ML infrastructure programs; allocation may shift with program priorities at lead directionWhat this role is notGraphics driver, kernel, or low-level GPU driver development.Research-only role without production deliverables.Junior or mid-level implementation; SDE 5 expects senior ownership and technical depth across ML systems and agent engineering.TekWissen Group is an equal opportunity employer supporting workforce diversity.