{"schemaVersion":"jobsearcher.job.v1","id":"dfd43d524dc879b3381e5cca","url":"https://jobsearcher.com/jobs/dfd43d524dc879b3381e5cca","canonicalUrl":"https://jobsearcher.com/jobs/dfd43d524dc879b3381e5cca","title":"Site Reliability Engineer","description":"Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment.\nOur customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast‑growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security.\nBacked by world‑class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people‑centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI‑ready operations.\nAbout the Role This is not a \"keep the lights on\" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.\nIf you thrive in environments where you’re handed a problem rather than a playbook this role is for you.\nWhat You'll Own Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use\nDrive the reliability and operational excellence of the Forward SaaS platform\nBuild and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers do\nLead incident response: on‑call rotations, runbooks, post‑mortems, and the follow‑through to make sure the same incident doesn't happen twice\nPartner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviews\nHelp define and build the SRE team as the company scales — this is a foundational hire with a path to leadership\nWhat We're Looking For 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment\nProven experience building or significantly maturing an SRE function — not just operating within one someone else built\nStrong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus\nHands‑on experience with Kubernetes and container orchestration in production environments\nDeep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similar\nStrong scripting and automation skills in Python, Bash, or similar\nExperience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)\nTrack record of owning and improving incident response processes including blameless post‑mortems and SLO‑driven reliability improvements\nAbility to communicate clearly with both engineering teams and non‑technical stakeholders — you can explain an outage to a customer‑facing team without jargon and explain an SLO to an executive without losing them\nNice to Have Experience supporting enterprise or federal government customers with high availability requirements\nExperience in a foundational or early SRE hire capacity at a growth stage company\nWhat This Role Is Not A pure ops or NOC role — you are building and engineering, not just monitoring\nA siloed function — you will be deeply embedded with product and engineering teams\nA ticket‑taker — you will be proactively identifying and solving reliability problems before they become incidents\nThe base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location.\n\n#J-18808-Ljbffr","company":"Forward","rawCompany":"forward","city":"Santa Clara","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-19T03:18:34.843Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1231.00","title":"Computer Network Support Specialists","slug":"computer-network-support-specialists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer","description":"Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment.\nOur customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast‑growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security.\nBacked by world‑class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people‑centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI‑ready operations.\nAbout the Role This is not a \"keep the lights on\" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.\nIf you thrive in environments where you’re handed a problem rather than a playbook this role is for you.\nWhat You'll Own Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use\nDrive the reliability and operational excellence of the Forward SaaS platform\nBuild and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers do\nLead incident response: on‑call rotations, runbooks, post‑mortems, and the follow‑through to make sure the same incident doesn't happen twice\nPartner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviews\nHelp define and build the SRE team as the company scales — this is a foundational hire with a path to leadership\nWhat We're Looking For 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment\nProven experience building or significantly maturing an SRE function — not just operating within one someone else built\nStrong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus\nHands‑on experience with Kubernetes and container orchestration in production environments\nDeep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similar\nStrong scripting and automation skills in Python, Bash, or similar\nExperience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)\nTrack record of owning and improving incident response processes including blameless post‑mortems and SLO‑driven reliability improvements\nAbility to communicate clearly with both engineering teams and non‑technical stakeholders — you can explain an outage to a customer‑facing team without jargon and explain an SLO to an executive without losing them\nNice to Have Experience supporting enterprise or federal government customers with high availability requirements\nExperience in a foundational or early SRE hire capacity at a growth stage company\nWhat This Role Is Not A pure ops or NOC role — you are building and engineering, not just monitoring\nA siloed function — you will be deeply embedded with product and engineering teams\nA ticket‑taker — you will be proactively identifying and solving reliability problems before they become incidents\nThe base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location.\n\n#J-18808-Ljbffr","datePosted":"2026-07-19T03:18:34.843Z","dateModified":"2026-07-19T03:18:34.843Z","hiringOrganization":{"@type":"Organization","name":"Forward","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Santa Clara","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"dfd43d524dc879b3381e5cca"},"url":"https://jobsearcher.com/jobs/dfd43d524dc879b3381e5cca"}}