{"schemaVersion":"jobsearcher.job.v1","id":"3d83eab7fb9113e5eef18dc4","url":"https://jobsearcher.com/jobs/3d83eab7fb9113e5eef18dc4","canonicalUrl":"https://jobsearcher.com/jobs/3d83eab7fb9113e5eef18dc4","title":"Systems Observability Specialist","description":"Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.\n\nThis is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.\n\nJob Title: Systems Observability Specialist\nLocation: 100% Remote (U.S.)\nPosition Type: Full-time, Direct W2\nSalary Range: $135,000–$155,000 Annually\nExperience Required: 6+ years\n\nSponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.\n\nJob Summary:\nWe are looking for an Systems Observability Specialist to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run. The role spans the full observability stack — from collection agents and pipelines to long-term storage, dashboards, and alerting workflows — with a strong focus on usability, signal quality, and operational ROI. The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders.\n\nKey Responsibilities\nDesign and operate enterprise-grade observability platforms covering metrics, logs, traces, events, and synthetic monitoring.\nArchitect Prometheus / Thanos / Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments for high availability and scale.\nDevelop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging conventions.\nDefine and enforce SLOs, SLIs, and error budgets, and build the dashboards and alerts that operationalize them.\nBuild alerting strategies that minimize noise, surface actionable signals, and integrate cleanly with on-call workflows in PagerDuty, Opsgenie, or similar tools.\nOperate large-scale time-series and log storage platforms, balancing retention, query performance, and cost.\nDesign distributed tracing pipelines and help teams use traces to diagnose latency and reliability issues.\nDevelop self-service tooling, paved-road libraries, and templates that make adoption of observability standards easy for product teams.\nDrive cost management and label-cardinality discipline across the observability estate.\nLead incident response readiness improvements through better dashboards, alerting hygiene, and post-incident analysis tooling.\nPartner with SRE and platform teams to integrate observability into deployment pipelines, canary analysis, and progressive delivery workflows.\nEvaluate and recommend observability vendors and open-source tools based on cost, capability, and operational maturity.\nMentor engineering teams on observability fundamentals, debugging techniques, and SLO-driven operations.\nMaintain documentation, onboarding guides, and runbooks for the observability platform.\n\nRequired Qualifications\nBachelor’s degree in Computer Science or a related field.\nFive or more years of experience in SRE, platform engineering, or observability roles.\nDeep hands-on experience with Prometheus, Grafana, and at least one major commercial observability platform such as Datadog, New Relic, or Splunk.\nStrong understanding of OpenTelemetry, distributed tracing, and structured logging.\nProficiency in at least one general-purpose language such as Go, Python, or Java.\nExperience operating high-cardinality, high-throughput metrics and log pipelines.\nStrong understanding of SLOs, error budgets, and SRE principles.\nExperience integrating observability with CI/CD and incident management tooling.\nSolid grasp of Linux internals, networking, and container platforms.\nExcellent communication and collaboration skills.\n\nPreferred Qualifications\nExperience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.\nContributions to OpenTelemetry or observability open-source projects.\nFamiliarity with eBPF-based observability tooling.\nExperience driving observability cost optimization initiatives.\nExposure to regulated environments with audit-grade logging requirements.\n\nHow to Apply\nWould you like to know more about this opportunity? For immediate consideration, please send your resume to Harry@bvteck.com or contact us at (908)676-4399. Learn more about Bright Vision Technologies at www.bvteck.com.\n\nBright Vision Technologies is an Equal Opportunity Employer.\nEqual Employment Opportunity (EEO) Statement\nBright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.\nBV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.\nlVrJ7Ge2gF","company":"Brightvisiontechnologies","rawCompany":"brightvisiontechnologies","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-03T22:52:28.210Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Systems Observability Specialist","description":"Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.\n\nThis is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.\n\nJob Title: Systems Observability Specialist\nLocation: 100% Remote (U.S.)\nPosition Type: Full-time, Direct W2\nSalary Range: $135,000–$155,000 Annually\nExperience Required: 6+ years\n\nSponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.\n\nJob Summary:\nWe are looking for an Systems Observability Specialist to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run. The role spans the full observability stack — from collection agents and pipelines to long-term storage, dashboards, and alerting workflows — with a strong focus on usability, signal quality, and operational ROI. The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders.\n\nKey Responsibilities\nDesign and operate enterprise-grade observability platforms covering metrics, logs, traces, events, and synthetic monitoring.\nArchitect Prometheus / Thanos / Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments for high availability and scale.\nDevelop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging conventions.\nDefine and enforce SLOs, SLIs, and error budgets, and build the dashboards and alerts that operationalize them.\nBuild alerting strategies that minimize noise, surface actionable signals, and integrate cleanly with on-call workflows in PagerDuty, Opsgenie, or similar tools.\nOperate large-scale time-series and log storage platforms, balancing retention, query performance, and cost.\nDesign distributed tracing pipelines and help teams use traces to diagnose latency and reliability issues.\nDevelop self-service tooling, paved-road libraries, and templates that make adoption of observability standards easy for product teams.\nDrive cost management and label-cardinality discipline across the observability estate.\nLead incident response readiness improvements through better dashboards, alerting hygiene, and post-incident analysis tooling.\nPartner with SRE and platform teams to integrate observability into deployment pipelines, canary analysis, and progressive delivery workflows.\nEvaluate and recommend observability vendors and open-source tools based on cost, capability, and operational maturity.\nMentor engineering teams on observability fundamentals, debugging techniques, and SLO-driven operations.\nMaintain documentation, onboarding guides, and runbooks for the observability platform.\n\nRequired Qualifications\nBachelor’s degree in Computer Science or a related field.\nFive or more years of experience in SRE, platform engineering, or observability roles.\nDeep hands-on experience with Prometheus, Grafana, and at least one major commercial observability platform such as Datadog, New Relic, or Splunk.\nStrong understanding of OpenTelemetry, distributed tracing, and structured logging.\nProficiency in at least one general-purpose language such as Go, Python, or Java.\nExperience operating high-cardinality, high-throughput metrics and log pipelines.\nStrong understanding of SLOs, error budgets, and SRE principles.\nExperience integrating observability with CI/CD and incident management tooling.\nSolid grasp of Linux internals, networking, and container platforms.\nExcellent communication and collaboration skills.\n\nPreferred Qualifications\nExperience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.\nContributions to OpenTelemetry or observability open-source projects.\nFamiliarity with eBPF-based observability tooling.\nExperience driving observability cost optimization initiatives.\nExposure to regulated environments with audit-grade logging requirements.\n\nHow to Apply\nWould you like to know more about this opportunity? For immediate consideration, please send your resume to Harry@bvteck.com or contact us at (908)676-4399. Learn more about Bright Vision Technologies at www.bvteck.com.\n\nBright Vision Technologies is an Equal Opportunity Employer.\nEqual Employment Opportunity (EEO) Statement\nBright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.\nBV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.\nlVrJ7Ge2gF","datePosted":"2026-08-03T22:52:28.210Z","dateModified":"2026-08-03T22:52:28.210Z","hiringOrganization":{"@type":"Organization","name":"Brightvisiontechnologies","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"3d83eab7fb9113e5eef18dc4"},"url":"https://jobsearcher.com/jobs/3d83eab7fb9113e5eef18dc4"}}