{"schemaVersion":"jobsearcher.job.v1","id":"de3ba98f8c1d87bc8475c30b","url":"https://jobsearcher.com/jobs/de3ba98f8c1d87bc8475c30b","canonicalUrl":"https://jobsearcher.com/jobs/de3ba98f8c1d87bc8475c30b","title":"Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US","description":"About us:\n\nWorking at Tech Holding isn't just a job, it's an opportunity to be a part of something bigger. We are a full-service consulting firm that was founded on the premise of delivering predictable outcomes and high-quality solutions to our clients. Our founders and team members have industry experience and have held senior positions in a wide variety of companies – from emerging startups to large Fortune 50 firms – and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability.\n\nThe Role:\n\nWe are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform.This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives.\nThis is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation.\nYou will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling.\n\nKey Responsibilities:\n\nEstablish performance, throughput, latency, and capacity baselines for critical customer and platform workflows\nDefine and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds\nInstrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services\nIdentify system bottlenecks and lead cross-functional remediation efforts with engineering teams\nBuild capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost\nLead load, stress, soak, spike, failure, and recovery testing in representative environments\nDevelop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events\nDrive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements\nPartner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates\nOwn technical readiness assessments for major pilots, partnerships, and production launches\nCreate operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures\nLead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work\nMake infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership\nRecommend capacity and reliability investments before they become production constraints\n\nRequired Skills:\n\nSignificant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline\nExperience supporting production systems with meaningful scale, traffic, latency, or availability requirements\nDeep understanding of observability, performance analysis, capacity planning, and reliability engineering\nStrong hands-on experience with cloud infrastructure and production distributed systems\nDeep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes\nExperience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics\nHands-on experience performing load, stress, soak, scalability, and resilience testing\nAbility to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements\nExperience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios\nStrong incident management and root-cause analysis experience\nAbility to translate technical performance and reliability risks into clear business implications for senior leadership\nStrong judgment around when systems genuinely require optimization versus when additional complexity is premature\n\nNice to have:\n\nExperience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms\nExperience creating capacity-cost models and forecasting infrastructure requirements\nExperience building performance and reliability gates into CI/CD pipelines\nExperience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnerships\nExperience leading reliability or performance initiatives that span multiple engineering teams\nWhat Success Looks Like\nWithin your first several months, you will have:\nEstablished measurable throughput, latency, and capacity baselines for critical platform journeys\nDefined initial SLOs, error budgets, dashboards, alerts, and performance thresholds\nIdentified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap\nValidated representative high-scale scenarios through load, soak, stress, failure, and recovery testing\nDeveloped a capacity and cost model showing how the platform can support significant increases in demand\nEstablished production-readiness criteria and clear go/no-go evidence for major launches and partnerships\nCreated repeatable scale-up, incident, rollback, and dependency-failure runbooks\nGiven leadership a clear, evidence-based understanding of the platform's current capacity envelope and future scaling requirements\nEmployment type:\nContract\n\n*Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time\n\n#LI-Remote\n\nTech Holding is proud to be an Equal Opportunity Employer and is committed to fostering a diverse and inclusive workplace. We welcome applicants from all backgrounds and experiences, and we consider qualified applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic. If you require accommodation in the application process, please contact our HR","company":"Tech Holding","rawCompany":"tech holding","isRemote":true,"isActive":false,"createdAt":"2026-08-22T15:56:52.499Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US","description":"About us:\n\nWorking at Tech Holding isn't just a job, it's an opportunity to be a part of something bigger. We are a full-service consulting firm that was founded on the premise of delivering predictable outcomes and high-quality solutions to our clients. Our founders and team members have industry experience and have held senior positions in a wide variety of companies – from emerging startups to large Fortune 50 firms – and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability.\n\nThe Role:\n\nWe are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform.This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives.\nThis is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation.\nYou will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling.\n\nKey Responsibilities:\n\nEstablish performance, throughput, latency, and capacity baselines for critical customer and platform workflows\nDefine and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds\nInstrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services\nIdentify system bottlenecks and lead cross-functional remediation efforts with engineering teams\nBuild capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost\nLead load, stress, soak, spike, failure, and recovery testing in representative environments\nDevelop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events\nDrive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements\nPartner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates\nOwn technical readiness assessments for major pilots, partnerships, and production launches\nCreate operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures\nLead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work\nMake infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership\nRecommend capacity and reliability investments before they become production constraints\n\nRequired Skills:\n\nSignificant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline\nExperience supporting production systems with meaningful scale, traffic, latency, or availability requirements\nDeep understanding of observability, performance analysis, capacity planning, and reliability engineering\nStrong hands-on experience with cloud infrastructure and production distributed systems\nDeep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes\nExperience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics\nHands-on experience performing load, stress, soak, scalability, and resilience testing\nAbility to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements\nExperience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios\nStrong incident management and root-cause analysis experience\nAbility to translate technical performance and reliability risks into clear business implications for senior leadership\nStrong judgment around when systems genuinely require optimization versus when additional complexity is premature\n\nNice to have:\n\nExperience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms\nExperience creating capacity-cost models and forecasting infrastructure requirements\nExperience building performance and reliability gates into CI/CD pipelines\nExperience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnerships\nExperience leading reliability or performance initiatives that span multiple engineering teams\nWhat Success Looks Like\nWithin your first several months, you will have:\nEstablished measurable throughput, latency, and capacity baselines for critical platform journeys\nDefined initial SLOs, error budgets, dashboards, alerts, and performance thresholds\nIdentified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap\nValidated representative high-scale scenarios through load, soak, stress, failure, and recovery testing\nDeveloped a capacity and cost model showing how the platform can support significant increases in demand\nEstablished production-readiness criteria and clear go/no-go evidence for major launches and partnerships\nCreated repeatable scale-up, incident, rollback, and dependency-failure runbooks\nGiven leadership a clear, evidence-based understanding of the platform's current capacity envelope and future scaling requirements\nEmployment type:\nContract\n\n*Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time\n\n#LI-Remote\n\nTech Holding is proud to be an Equal Opportunity Employer and is committed to fostering a diverse and inclusive workplace. We welcome applicants from all backgrounds and experiences, and we consider qualified applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic. If you require accommodation in the application process, please contact our HR","datePosted":"2026-08-22T15:56:52.499Z","dateModified":"2026-08-22T15:56:52.499Z","hiringOrganization":{"@type":"Organization","name":"Tech Holding","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"de3ba98f8c1d87bc8475c30b"},"url":"https://jobsearcher.com/jobs/de3ba98f8c1d87bc8475c30b"}}