Top Remote Site Reliability Engineer Jobs in Seattle, WA

28 Days AgoSaved
Remote
United States
Senior level
Senior level
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills: AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
28 Days AgoSaved
Remote
United States
115K-175K Annually
Entry level
115K-175K Annually
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills: AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Reposted 6 Days AgoSaved
In-Office or Remote
United States
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills: AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
29 Days AgoSaved
In-Office or Remote
United States
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills: AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
6 Days AgoSaved
Remote
United States
150K-170K Annually
Senior level
150K-170K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Own reliability for Kong’s Volcano internal developer platform by defining SLOs, incident practices, and observability. Design multi-region Kubernetes infrastructure, GitOps deployment automation, preview environments, managed PostgreSQL, Redis, and object storage. Lead reliability and compliance initiatives with engineering, security, and OCTO leadership while evaluating emerging edge, serverless, vector database, and AI infrastructure technologies.
Top Skills: ArgocdAutoscalingCi/CdCniDatadogGitopsGrafanaHelmIngressKubernetesObject StoragePostgresPrometheusRedisService MeshTerraformTerragrunt
One Month AgoSaved
In-Office or Remote
United States
76K-136K Annually
Entry level
76K-136K Annually
Entry level
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills: AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
One Month AgoSaved
In-Office or Remote
United States
135K-160K Annually
Senior level
135K-160K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills: Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Reposted One Month AgoSaved
Remote or Hybrid
Seattle, WA, USA
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills: ArgocdAWSDatadogGitGoKubernetesPythonTerraform
7 Days AgoSaved
In-Office or Remote
United States
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead site reliability engineering for Akamai’s compute infrastructure and services. Develop automation, define reliability requirements and standards, establish SLOs and KPIs, troubleshoot complex distributed-system and hardware issues, manage incidents and postmortems, and participate in on-call rotations. Coordinate engineering teams, support strategic initiatives, mentor engineers, and guide restoration of service-impacting issues.
Top Skills: Cloud InfrastructureDistributed SystemsInfrastructure AutomationLinux
8 Days AgoSaved
Remote
USA
160K-208K Annually
Senior level
160K-208K Annually
Senior level
Artificial Intelligence • Healthtech • Software
Build and maintain scalable infrastructure platforms supporting domestic and international workloads. Develop CI/CD, declarative lifecycle management, Kubernetes clusters, monitoring, and automation systems. Troubleshoot infrastructure issues, respond to alerts, minimize downtime, streamline delivery pipelines and database changes, and help guide SRE team direction. Collaborate with engineers, data scientists, and technology professionals while promoting a high-performance, cross-functional culture.
Top Skills: AWSAzureContainerdContinuous DeploymentContinuous IntegrationDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
8 Days AgoSaved
Remote
United States
120K-170K Annually
Senior level
120K-170K Annually
Senior level
Software
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
Top Skills: Amazon Web ServicesCheckovClaude CodeDatadogDockerDynatraceGithub ActionsGitlab CiGoJavaScriptKubernetesAzureNew RelicPrisma CloudPythonTerraformWiz
Reposted 2 Months AgoSaved
Easy Apply
Remote or Hybrid
United States
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills: AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 9 Days AgoSaved
Remote
USA
134K-184K Annually
Senior level
134K-184K Annually
Senior level
Healthtech
Lead the migration from legacy Azure services to a Kubernetes-based, containerized microservices platform. Design, build, and scale infrastructure, implement observability (monitoring/alerting/logging), drive incident response and SLOs, automate with IaC and CI/CD, optimize cost and networking, mentor teams, and document systems to ensure reliable, scalable healthcare platform operations.
Top Skills: .NetAWSAzureAzure Entra IdBashC#DatadogGCPGithub ActionsGitlab Ci/CdGrafanaHelmKubernetesPrometheusPythonTerraform
9 Days AgoSaved
Remote
United States
156K-199K Annually
Senior level
156K-199K Annually
Senior level
Cybersecurity
Owns the FedRAMP-authorized AWS GovCloud environment, ensuring reliability, security, compliance, patching, vulnerability remediation, continuous monitoring, deployments, hardening, audits, incident response, and on-call support. The role requires hands-on Kubernetes, AWS, infrastructure-as-code, CI/CD, vulnerability management, automation, and regulated-environment experience, while collaborating with SRE and security teams.
Top Skills: AnsibleAws GovcloudCloudFormationFips-Validated CryptographyGithub ActionsGoJenkinsKubernetesNessusOpenscapPythonQualysRubySIEMTenableTerraform
One Month AgoSaved
Remote
United States
90K-100K Annually
Mid level
90K-100K Annually
Mid level
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills: Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
One Month AgoSaved
Remote
United States
125K-150K Annually
Mid level
125K-150K Annually
Mid level
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills: ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
One Month AgoSaved
Remote or Hybrid
USA
136K-181K Annually
Entry level
136K-181K Annually
Entry level
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills: Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
12 Days AgoSaved
Remote
USA
152K-205K Annually
Senior level
152K-205K Annually
Senior level
Information Technology • Security • Software • Cybersecurity
Owns production reliability for high-throughput, low-latency systems, including observability, SLOs, incident response, capacity planning, infrastructure as code, progressive delivery, chaos testing, and operational tooling. The role requires hands-on software and infrastructure engineering, production Redis/ElastiCache expertise, cloud infrastructure knowledge, security awareness, mentoring, and participation in on-call operations.
Top Skills: AWSDatadogEksElasticacheGoGrafanaKubernetesOpentelemetryPrometheusPythonRedisTerraform
12 Days AgoSaved
Remote
United States
210K-275K Annually
Senior level
210K-275K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Design and maintain reliable, scalable infrastructure for Replit’s global platform. Responsibilities include building observability and alerting systems, automating infrastructure with infrastructure-as-code, managing CI/CD pipelines, defining SLOs and SLIs, leading incident response and postmortems, maintaining runbooks, optimizing performance, and improving capacity, availability, and recovery times.
Top Skills: AnsibleCi/CdDatadogGoGoogle Cloud PlatformGrafanaKubernetesPrometheusPulumiPythonTerraform
13 Days AgoSaved
Remote
United States
85K-141K Annually
Senior level
85K-141K Annually
Senior level
Cloud • Security • Cybersecurity
Owns operational capabilities for FedRAMP-regulated cloud environments, including monitoring, alerting, backup and recovery, continuous compliance evidence, incident response, and automation. The role leads major incident command, improves runbooks and operational standards, supports client-facing service delivery, partners on transitions to managed operations, and mentors engineers. Requires deep cloud operations, observability, infrastructure-as-code, resilience engineering, security-control frameworks, and senior escalation experience.
Top Skills: AnsibleAWSAzureCi/CdGCPGoInfrastructure As CodeJSONOscalPolicy As CodePythonSIEMTerraform
14 Days AgoSaved
Remote
United States
Senior level
Senior level
Software
Own reliability, performance, scalability, and operational standards across on-premises, private-cloud, and AWS environments. Build infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling; define SLOs, SLIs, error budgets, and readiness standards. Lead incident response, root-cause analysis, disaster-recovery readiness, release coordination, and preventive automation. Mentor engineers, coach teams on operational practices, coordinate on-call coverage, and ensure infrastructure meets security and compliance requirements.
Top Skills: AnsibleAuto ScalingAWSBashCi/CdClaudeCloudwatchCrowdstrikeDatadogDistributed SystemsDnsDockerEc2EcsGitGithub CopilotGitlabGitopsIamKubernetesLinuxLoad BalancingOracle LinuxPythonQualysRapid7RhelS3Tcp/IpTerraformVpcWireshark
Reposted 14 Days AgoSaved
Remote
USA
160K-208K Annually
Senior level
160K-208K Annually
Senior level
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills: AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Reposted One Month AgoSaved
Remote
United States
115K-135K Annually
Mid level
115K-135K Annually
Mid level
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills: ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
One Month AgoSaved
Remote
USA
Entry level
Entry level
Cybersecurity
Assist with monitoring system performance, uptime, and reliability; support incident response, troubleshooting, and root cause analysis; monitor SRE alerts in Slack and report issues; learn cloud platforms; and collaborate with development and DevOps teams.
Top Skills: AWSAzureDevOpsDevsecopsGCPGrafanaKubernetesPrometheusSlack
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account