{"schemaVersion":"jobsearcher.job.v1","id":"03a2b2d5490f1bebe86ebbad","url":"https://jobsearcher.com/jobs/03a2b2d5490f1bebe86ebbad","canonicalUrl":"https://jobsearcher.com/jobs/03a2b2d5490f1bebe86ebbad","title":"Network Automation & Reliability Engineer","description":"Overview\nOntrac Solutions is seeking a Network Automation & Reliability Engineer to operate and automate hyperscale data center, backbone, and out-of-band (OOB) networks. This is a Python-first role: you will spend more of your week writing code that operates the network than typing on a CLI. You will build the tooling, telemetry pipelines, and self-healing automation that keep a global fleet of data centers and POP sites running, and you will carry the on-call pager for the systems you build.\nWe are looking for an engineer who is genuinely production-grade on both sides of the job — deep routing and switching fundamentals and real, sustained Python software work. Candidates who are strong on one side and thin on the other will not clear screening.\nWhat your application must clearly show\nWe screen against the requirements below exactly as written — your resume should make these easy to find:\n\nPython you actually wrote, described as software, not as a skills keyword.Name the project, what it did, roughly how large it was, and who used it. \"Python (scripting)\" in a skills list will not clear this bar.A GitHub, GitLab, or public repo link is strongly preferred— we look at code.\nThespecific Python network librariesyou have used in production — for example Netmiko, NAPALM, Nornir, pyATS/Genie, Scrapli, ncclient, or a vendor SDK — and what you built with each.\nBGP in production at scale.Name the policy work: local preference, MED, communities, import/export policy, route reflection, ECMP, or multihoming across carriers.\nHands-on withat least twoof Arista EOS, Juniper Junos (QFX / SRX / PTX / MX), or Cisco IOS-XR / NX-OS — named by platform, not just by vendor.\nTelemetry and observabilityyou implemented: gNMI/gRPC, OpenConfig, streaming telemetry, flow telemetry, or SNMP-to-TSDB pipelines. Say what you subscribed to and what you did with the data.\nConfig-as-code:Jinja2 templating, NETCONF/YANG or REST-API-driven provisioning, golden configs, ZTP, and the Git/CI workflow you shipped changes through.\nYournetworking certifications, named, with dates and credential IDs or verification links— we verify certifications.\nWhether you have carriedproduction on-call, and at what scale (sites, devices, or POPs).\nShortly after you apply you will receive a short role-specific questionnaire — completing it promptly is the fastest way to move into screening.\nRequired Qualifications:\n\nPython — primary requirement:Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts.\nRouting & switching depth:Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays.\nMulti-vendor hardware:Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS.\nAutomation frameworks:Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management.\nTelemetry & monitoring:gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them.\nLinux:Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting.\nVersion control and CI:Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions).\nProduction on-call:Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis.\nExperience level:Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply.\nLocation & work authorization:Must be located in the United States and authorized to work in the US.\nPreferred Qualifications:\n\nOut-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity.\nData center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing.\nMACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC.\nOptical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation.\nBuilding or contributing to a networksource of truth(NetBox or in-house) aggregating BGP, link-state, and drain-state data.\nApplying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization.\nA master's degree in Network Engineering, Telecommunications, or Computer Science.\nKey Responsibilities — Network Automation & Tooling (primary focus):\n\nDesign, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet.\nBuild model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate configuration drift.\nDevelop automated remediation for recurring failure classes — route flaps, hardware faults, optical degradation — triggered by syslog and telemetry.\nReduce operational toil: identify manual runbook steps that recur, and replace them with tested, reviewed code.\nKey Responsibilities — Backbone, Data Center & Out-of-Band Operations:\n\nOperate and scale data center fabrics, backbone links, and out-of-band management networks across a global footprint of data centers and POP sites.\nDesign and tune BGP policy — local preference, MED, communities, import/export policy, default-route propagation — to control path selection and eliminate single-carrier points of failure.\nSupport capacity expansion and site turn-ups: high-level design, port channelization, optics and cabling standards, and hardware qualification.\nExecute zero-downtime change management in production, including hitless migrations and staged rollbacks.\nKey Responsibilities — Telemetry, Observability & Reliability:\n\nOperationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and tune subscription jobs for signal quality and management-plane efficiency.\nBuild and maintain observability that supports hop-by-hop path tracing, multi-layer fault isolation, and fast root cause analysis.\nContribute to a real-time network source of truth aggregating BGP, link-state, and drain-state data.\nParticipate in a 24x7 on-call rotation; lead incident response, write RCAs, and drive the follow-up automation that prevents recurrence.\nKey Responsibilities — Security & Change Governance:\n\nMaintain and optimize firewall and ACL policy across multi-vendor platforms, keeping rule sets scoped and performant.\nSupport SIRT/PSIRT CVE remediation through automated regression testing and config-as-code pipelines.\nAuthor and maintain technical documentation — designs, runbooks, and API contracts — for the tooling and networks you own.","company":"Ontrac Solutions","rawCompany":"ontrac solutions","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-10T15:11:07.016Z","occupations":[{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1241.00","title":"Computer Network Architects","slug":"computer-network-architects"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Network Automation & Reliability Engineer","description":"Overview\nOntrac Solutions is seeking a Network Automation & Reliability Engineer to operate and automate hyperscale data center, backbone, and out-of-band (OOB) networks. This is a Python-first role: you will spend more of your week writing code that operates the network than typing on a CLI. You will build the tooling, telemetry pipelines, and self-healing automation that keep a global fleet of data centers and POP sites running, and you will carry the on-call pager for the systems you build.\nWe are looking for an engineer who is genuinely production-grade on both sides of the job — deep routing and switching fundamentals and real, sustained Python software work. Candidates who are strong on one side and thin on the other will not clear screening.\nWhat your application must clearly show\nWe screen against the requirements below exactly as written — your resume should make these easy to find:\n\nPython you actually wrote, described as software, not as a skills keyword.Name the project, what it did, roughly how large it was, and who used it. \"Python (scripting)\" in a skills list will not clear this bar.A GitHub, GitLab, or public repo link is strongly preferred— we look at code.\nThespecific Python network librariesyou have used in production — for example Netmiko, NAPALM, Nornir, pyATS/Genie, Scrapli, ncclient, or a vendor SDK — and what you built with each.\nBGP in production at scale.Name the policy work: local preference, MED, communities, import/export policy, route reflection, ECMP, or multihoming across carriers.\nHands-on withat least twoof Arista EOS, Juniper Junos (QFX / SRX / PTX / MX), or Cisco IOS-XR / NX-OS — named by platform, not just by vendor.\nTelemetry and observabilityyou implemented: gNMI/gRPC, OpenConfig, streaming telemetry, flow telemetry, or SNMP-to-TSDB pipelines. Say what you subscribed to and what you did with the data.\nConfig-as-code:Jinja2 templating, NETCONF/YANG or REST-API-driven provisioning, golden configs, ZTP, and the Git/CI workflow you shipped changes through.\nYournetworking certifications, named, with dates and credential IDs or verification links— we verify certifications.\nWhether you have carriedproduction on-call, and at what scale (sites, devices, or POPs).\nShortly after you apply you will receive a short role-specific questionnaire — completing it promptly is the fastest way to move into screening.\nRequired Qualifications:\n\nPython — primary requirement:Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts.\nRouting & switching depth:Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays.\nMulti-vendor hardware:Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS.\nAutomation frameworks:Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management.\nTelemetry & monitoring:gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them.\nLinux:Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting.\nVersion control and CI:Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions).\nProduction on-call:Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis.\nExperience level:Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply.\nLocation & work authorization:Must be located in the United States and authorized to work in the US.\nPreferred Qualifications:\n\nOut-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity.\nData center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing.\nMACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC.\nOptical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation.\nBuilding or contributing to a networksource of truth(NetBox or in-house) aggregating BGP, link-state, and drain-state data.\nApplying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization.\nA master's degree in Network Engineering, Telecommunications, or Computer Science.\nKey Responsibilities — Network Automation & Tooling (primary focus):\n\nDesign, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet.\nBuild model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate configuration drift.\nDevelop automated remediation for recurring failure classes — route flaps, hardware faults, optical degradation — triggered by syslog and telemetry.\nReduce operational toil: identify manual runbook steps that recur, and replace them with tested, reviewed code.\nKey Responsibilities — Backbone, Data Center & Out-of-Band Operations:\n\nOperate and scale data center fabrics, backbone links, and out-of-band management networks across a global footprint of data centers and POP sites.\nDesign and tune BGP policy — local preference, MED, communities, import/export policy, default-route propagation — to control path selection and eliminate single-carrier points of failure.\nSupport capacity expansion and site turn-ups: high-level design, port channelization, optics and cabling standards, and hardware qualification.\nExecute zero-downtime change management in production, including hitless migrations and staged rollbacks.\nKey Responsibilities — Telemetry, Observability & Reliability:\n\nOperationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and tune subscription jobs for signal quality and management-plane efficiency.\nBuild and maintain observability that supports hop-by-hop path tracing, multi-layer fault isolation, and fast root cause analysis.\nContribute to a real-time network source of truth aggregating BGP, link-state, and drain-state data.\nParticipate in a 24x7 on-call rotation; lead incident response, write RCAs, and drive the follow-up automation that prevents recurrence.\nKey Responsibilities — Security & Change Governance:\n\nMaintain and optimize firewall and ACL policy across multi-vendor platforms, keeping rule sets scoped and performant.\nSupport SIRT/PSIRT CVE remediation through automated regression testing and config-as-code pipelines.\nAuthor and maintain technical documentation — designs, runbooks, and API contracts — for the tooling and networks you own.","datePosted":"2026-08-10T15:11:07.016Z","dateModified":"2026-08-10T15:11:07.016Z","hiringOrganization":{"@type":"Organization","name":"Ontrac Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"03a2b2d5490f1bebe86ebbad"},"url":"https://jobsearcher.com/jobs/03a2b2d5490f1bebe86ebbad"}}