Bright Vision Technologies Logo

Bright Vision Technologies

Site Observability Engineer

Posted 2 Days Ago
Be an Early Applicant
In-Office
Bellevue, WA, USA
100K-150K Annually
Senior level
In-Office
Bellevue, WA, USA
100K-150K Annually
Senior level
Designs and operates enterprise observability platforms for metrics, logs, traces, events, synthetic monitoring, dashboards, and alerting. Establishes instrumentation standards, SLOs, SLIs, error budgets, and actionable on-call workflows. Manages high-scale telemetry storage, tracing pipelines, platform costs, and reliability tooling while partnering with SRE and platform teams. Builds self-service tools, integrates observability into CI/CD and progressive delivery, evaluates vendors, improves incident readiness, mentors engineers, and maintains documentation and runbooks.
The summary above was generated by AI
Site Observability Engineer- Remote 
 
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. 
 
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. 

 
Job Title:
Site Observability Engineer

Location: 100% Remote (U.S.) 
Position Type: Full-time, Direct W2 
Salary Range: $100,000–$150,000 Annually 
Experience Required: 6+ years 
 
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. 
 

Job Summary 
We are looking for a Site Observability Engineer to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run. The role spans the full observability stack — from collection agents and pipelines to long-term storage, dashboards, and alerting workflows — with a strong focus on usability, signal quality, and operational ROI. The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders.
Key Responsibilities
  • Design and operate enterprise-grade observability platforms covering metrics, logs, traces, events, and synthetic monitoring.
  • Architect Prometheus / Thanos / Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments for high availability and scale.
  • Develop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging conventions.
  • Define and enforce SLOs, SLIs, and error budgets, and build the dashboards and alerts that operationalize them.
  • Build alerting strategies that minimize noise, surface actionable signals, and integrate cleanly with on-call workflows in PagerDuty, Opsgenie, or similar tools.
  • Operate large-scale time-series and log storage platforms, balancing retention, query performance, and cost.
  • Design distributed tracing pipelines and help teams use traces to diagnose latency and reliability issues.
  • Develop self-service tooling, paved-road libraries, and templates that make adoption of observability standards easy for product teams.
  • Drive cost management and label-cardinality discipline across the observability estate.
  • Lead incident response readiness improvements through better dashboards, alerting hygiene, and post-incident analysis tooling.
  • Partner with SRE and platform teams to integrate observability into deployment pipelines, canary analysis, and progressive delivery workflows.
  • Evaluate and recommend observability vendors and open-source tools based on cost, capability, and operational maturity.
  • Mentor engineering teams on observability fundamentals, debugging techniques, and SLO-driven operations.
  • Maintain documentation, onboarding guides, and runbooks for the observability platform.
Required Qualifications
  • Bachelor’s degree in Computer Science or a related field.
  • Five or more years of experience in SRE, platform engineering, or observability roles.
  • Deep hands-on experience with Prometheus, Grafana, and at least one major commercial observability platform such as Datadog, New Relic, or Splunk.
  • Strong understanding of OpenTelemetry, distributed tracing, and structured logging.
  • Proficiency in at least one general-purpose language such as Go, Python, or Java.
  • Experience operating high-cardinality, high-throughput metrics and log pipelines.
  • Strong understanding of SLOs, error budgets, and SRE principles.
  • Experience integrating observability with CI/CD and incident management tooling.
  • Solid grasp of Linux internals, networking, and container platforms.
  • Excellent communication and collaboration skills.
Preferred Qualifications
  • Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.
  • Contributions to OpenTelemetry or observability open-source projects.
  • Familiarity with eBPF-based observability tooling.
  • Experience driving observability cost optimization initiatives.
  • Exposure to regulated environments with audit-grade logging requirements.

How to Apply 
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com
 
Bright Vision Technologies is an Equal Opportunity Employer. 
 

Similar Jobs

33 Minutes Ago
Hybrid
Seattle, WA, USA
136K-234K Annually
Senior level
136K-234K Annually
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Design, build, and operate low-latency, high-throughput ad delivery systems for bidding, ranking, auctions, and budget enforcement. Integrate ML signals, collaborate with product and data teams, improve observability and reliability, lead design discussions, and mentor junior engineers.
Top Skills: A/B TestingApache FlinkSparkCloud EnvironmentsGrpcJavaKotlinMachine LearningMicroservices
An Hour Ago
Hybrid
Seattle, WA, USA
50K-94K Annually
Mid level
50K-94K Annually
Mid level
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Provides underwriting and processing support for Subcontract Default Insurance, including new business, renewals, cancellations, endorsements, broker setup, licensing, reconciliations, regulatory requirements, and reporting. Responds to internal and external customer inquiries, maintains files, supports technical issue resolution and user testing, and assists with billing discrepancies. Mentors colleagues, serves as a workflow subject-matter expert, and performs office support and special projects.
Top Skills: Bond Processing SystemsCorporate SystemsSdi Systems
2 Hours Ago
In-Office or Remote
US
166K-269K Annually
Expert/Leader
166K-269K Annually
Expert/Leader
Consumer Web • eCommerce • Machine Learning • Software • Sports • Analytics
Own an AI portfolio across a business domain, identify and prioritize opportunities, and personally build production agentic solutions end to end. Establish reusable patterns, evaluations, human-in-the-loop controls, and measurable business outcomes. Partner with functional leaders, earn stakeholder trust, guide platform choices, and improve team capability through mentoring, design and code reviews, and interviewing. Build reliable, secure, scalable applications around AI systems across frontend and backend layers.
Top Skills: Ai AgentsBackend SystemsClaudeFrontend ApisOpenaiOrchestration Platforms

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account