{"schemaVersion":"jobsearcher.job.v1","id":"56de436366cff4cdf8c2d7a2","url":"https://jobsearcher.com/jobs/56de436366cff4cdf8c2d7a2","canonicalUrl":"https://jobsearcher.com/jobs/56de436366cff4cdf8c2d7a2","title":"Systems Development Engineer","description":"About Us\nAt Union, we are solving one of the hardest challenges in AI infrastructure today: enabling high-velocity iteration while maintaining seamless production-readiness for AI workloads at scale.\n\nFlyte, the open-source project we steward, has emerged as the modern standard for data and AI orchestration, and is trusted by leading technology organizations including LinkedIn, Stripe, and Wayve to run millions of mission-critical workflows on the platform. These workflows comprise data preparation, model training, and scaled inference spanning thousands of GPUs, all major clouds, and on-premise infrastructure.\n\nWe have a technical founding team who created Flyte while at Lyft, a deep bench of infrastructure experts from top companies, and have raised from top investors like NEA and Nava Ventures.\n\nAbout the Role\nWe are hiring a Systems Development Engineer to improve the reliability, operability, and customer experience of our production platform. This is a 50/50 operations and engineering role. Part of your time will be spent investigating customer-impacting production issues, and part will be spent building the tools, automation, system design, and engineering practices that prevent those issues from recurring.\n\nFor this role, production is the customer. You will work from real operational signals: customer issues, incidents, on-call pages, recurring support patterns, and gaps in observability or automation. You will sit in engineering and partner with customer-facing teams to turn those signals into durable platform improvements.\n\nThis is not a traditional support role. It is a systems engineering role for someone who can debug deeply, communicate clearly, and write software that reduces operational load. The full engineering team backs you on on-call.\n\nThis role is hybrid, based out of our Seattle office.\n\nWhat You'll Do\n\nInvestigate and resolve customer-impacting production issues across cloud infrastructure, workflow execution, access control, storage, networking, deployment systems, and observability.\n\nIdentify patterns in customer issues and convert them into automation, product improvements, runbooks, tests, or design changes.\n\nBuild internal tools and diagnostics that make production issues easier to detect, understand, and resolve.\n\nImprove platform observability, including logs, metrics, dashboards, alerts, and customer-visible debugging information.\n\nParticipate in design and development so systems are easier to operate, debug, and support, and guide engineering teams toward durable fixes.\n\nDefine and uphold operational engineering practices: production readiness, alert quality, runbook discipline, observability standards, regression prevention, and code quality.\n\nDrive measurable reductions in on-call pages, recurring customer issues, manual operational work, and time-to-resolution.\n\nWhat We're Looking For\n\nStrong software engineering skills in Python, Go, Java, or a similar language.\n\nExperience debugging production systems across multiple layers of the stack.\n\nPractical knowledge of Kubernetes, Linux, cloud infrastructure, distributed systems, networking, storage, and IAM.\n\nExperience with infrastructure-as-code, deployment systems, CI/CD, observability, and operational automation.\n\nAbility to move from ambiguous customer symptoms to clear technical diagnosis and durable remediation.\n\nStrong judgment about when to fix directly, automate, elevate, redesign, or build a broader platform improvement.\n\nClear written and verbal communication, especially around root cause analysis, technical recommendations, runbooks, and design feedback.\n\nA bias toward reducing toil through engineering rather than repeatedly solving the same issue by hand.\n\nPreferred Experience\n\nOperating customer-facing SaaS, cloud infrastructure, self-hosted or on-prem deployments, or workflow orchestration systems.\n\nBatch workloads, autoscaling, capacity management, identity and access systems, storage systems, or platform observability.\n\nImproving on-call health, reducing ticket volume, or building production diagnostics.\n\nWorking across support, customer success, product, and engineering teams.\n\nSuccess Looks Like\n\nCustomer issues are diagnosed faster and recur less often.\n\nEngineering teams receive actionable feedback from production and customer pain.\n\nCommon operational problems become automated workflows, better diagnostics, clearer runbooks, or product fixes.\n\nOn-call pages trend toward roughly one per month.\n\nThe platform becomes easier to operate, easier to debug, and safer to change.\n\nCustomers experience fewer production surprises and faster resolution when issues do happen.\n\nBenefits & Belonging\n\nAt Union.ai we know that employees who feel their best can build amazing things and we are proud to offer best in class benefits that will continually evolve and grow as the needs of our employees do. Benefits may vary based on country.\n\nExcellent medical - We pay 100% of your premiums and 90% for your dependents\n\nGenerous dental and vision plans- We pay 90% of the premiums for you and your dependents\n\nMeaningful equity in the form of options – all employees are owners here\n\nUnlimited time off + 12 company holidays\n\n401K match - Union.ai matches 100% of contributions up to the first 3%, and 50% up to 5%\n\n12 weeks paid parental leave for primary and secondary caregivers\n\nFlexible work schedule (some restrictions apply)\n\nFor in office employees: Lunch provided onsite and well stocked kitchen with snacks and drinks.\n\nWe believe that our differences are what bring us together to achieve truly special outcomes. We strive to be inclusive and focus on building teams that embody that quality too. Union.ai is an equal-opportunity employer and we encourage you to apply, even if your experience doesn’t align exactly with our job description.\n\n#J-18808-Ljbffr","company":"Linuxconfig","rawCompany":"linuxconfig","city":"Seattle","state":"WA","isRemote":false,"isActive":false,"createdAt":"2026-09-04T03:39:16.464Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Systems Development Engineer","description":"About Us\nAt Union, we are solving one of the hardest challenges in AI infrastructure today: enabling high-velocity iteration while maintaining seamless production-readiness for AI workloads at scale.\n\nFlyte, the open-source project we steward, has emerged as the modern standard for data and AI orchestration, and is trusted by leading technology organizations including LinkedIn, Stripe, and Wayve to run millions of mission-critical workflows on the platform. These workflows comprise data preparation, model training, and scaled inference spanning thousands of GPUs, all major clouds, and on-premise infrastructure.\n\nWe have a technical founding team who created Flyte while at Lyft, a deep bench of infrastructure experts from top companies, and have raised from top investors like NEA and Nava Ventures.\n\nAbout the Role\nWe are hiring a Systems Development Engineer to improve the reliability, operability, and customer experience of our production platform. This is a 50/50 operations and engineering role. Part of your time will be spent investigating customer-impacting production issues, and part will be spent building the tools, automation, system design, and engineering practices that prevent those issues from recurring.\n\nFor this role, production is the customer. You will work from real operational signals: customer issues, incidents, on-call pages, recurring support patterns, and gaps in observability or automation. You will sit in engineering and partner with customer-facing teams to turn those signals into durable platform improvements.\n\nThis is not a traditional support role. It is a systems engineering role for someone who can debug deeply, communicate clearly, and write software that reduces operational load. The full engineering team backs you on on-call.\n\nThis role is hybrid, based out of our Seattle office.\n\nWhat You'll Do\n\nInvestigate and resolve customer-impacting production issues across cloud infrastructure, workflow execution, access control, storage, networking, deployment systems, and observability.\n\nIdentify patterns in customer issues and convert them into automation, product improvements, runbooks, tests, or design changes.\n\nBuild internal tools and diagnostics that make production issues easier to detect, understand, and resolve.\n\nImprove platform observability, including logs, metrics, dashboards, alerts, and customer-visible debugging information.\n\nParticipate in design and development so systems are easier to operate, debug, and support, and guide engineering teams toward durable fixes.\n\nDefine and uphold operational engineering practices: production readiness, alert quality, runbook discipline, observability standards, regression prevention, and code quality.\n\nDrive measurable reductions in on-call pages, recurring customer issues, manual operational work, and time-to-resolution.\n\nWhat We're Looking For\n\nStrong software engineering skills in Python, Go, Java, or a similar language.\n\nExperience debugging production systems across multiple layers of the stack.\n\nPractical knowledge of Kubernetes, Linux, cloud infrastructure, distributed systems, networking, storage, and IAM.\n\nExperience with infrastructure-as-code, deployment systems, CI/CD, observability, and operational automation.\n\nAbility to move from ambiguous customer symptoms to clear technical diagnosis and durable remediation.\n\nStrong judgment about when to fix directly, automate, elevate, redesign, or build a broader platform improvement.\n\nClear written and verbal communication, especially around root cause analysis, technical recommendations, runbooks, and design feedback.\n\nA bias toward reducing toil through engineering rather than repeatedly solving the same issue by hand.\n\nPreferred Experience\n\nOperating customer-facing SaaS, cloud infrastructure, self-hosted or on-prem deployments, or workflow orchestration systems.\n\nBatch workloads, autoscaling, capacity management, identity and access systems, storage systems, or platform observability.\n\nImproving on-call health, reducing ticket volume, or building production diagnostics.\n\nWorking across support, customer success, product, and engineering teams.\n\nSuccess Looks Like\n\nCustomer issues are diagnosed faster and recur less often.\n\nEngineering teams receive actionable feedback from production and customer pain.\n\nCommon operational problems become automated workflows, better diagnostics, clearer runbooks, or product fixes.\n\nOn-call pages trend toward roughly one per month.\n\nThe platform becomes easier to operate, easier to debug, and safer to change.\n\nCustomers experience fewer production surprises and faster resolution when issues do happen.\n\nBenefits & Belonging\n\nAt Union.ai we know that employees who feel their best can build amazing things and we are proud to offer best in class benefits that will continually evolve and grow as the needs of our employees do. Benefits may vary based on country.\n\nExcellent medical - We pay 100% of your premiums and 90% for your dependents\n\nGenerous dental and vision plans- We pay 90% of the premiums for you and your dependents\n\nMeaningful equity in the form of options – all employees are owners here\n\nUnlimited time off + 12 company holidays\n\n401K match - Union.ai matches 100% of contributions up to the first 3%, and 50% up to 5%\n\n12 weeks paid parental leave for primary and secondary caregivers\n\nFlexible work schedule (some restrictions apply)\n\nFor in office employees: Lunch provided onsite and well stocked kitchen with snacks and drinks.\n\nWe believe that our differences are what bring us together to achieve truly special outcomes. We strive to be inclusive and focus on building teams that embody that quality too. Union.ai is an equal-opportunity employer and we encourage you to apply, even if your experience doesn’t align exactly with our job description.\n\n#J-18808-Ljbffr","datePosted":"2026-09-04T03:39:16.464Z","dateModified":"2026-09-04T03:39:16.464Z","hiringOrganization":{"@type":"Organization","name":"Linuxconfig","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Seattle","addressRegion":"WA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"56de436366cff4cdf8c2d7a2"},"url":"https://jobsearcher.com/jobs/56de436366cff4cdf8c2d7a2"}}