Top SRE Engineer Jobs in Seattle, WA

Reposted 7 Days AgoSaved
In-Office
Seattle, WA
102K-261K Annually
Junior
102K-261K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve Microsoft Defender services in US government cloud environments: provide 24x7 on-call support, run live-site incident response and postmortems, automate deployments and tooling, validate security/compliance during onboarding, collaborate with engineering teams, and apply software engineering best practices to increase reliability and observability at scale.
Top Skills: C#Ci/CdCloudDistributed SystemsGoJavaMicrosoft DefenderObservabilityPython
Reposted 7 Days AgoSaved
In-Office
Seattle, WA
Senior level
Senior level
Cloud • Software • Database
The Site Reliability Engineer will optimize and scale managed services across cloud providers, automate infrastructure, enhance monitoring, and ensure system reliability.
Top Skills: AWSAzureBashGCPGrafanaKubernetesLokiMimirPrometheusPython
10 Days AgoSaved
In-Office or Remote
Seattle, WA
120K-261K Annually
Mid level
120K-261K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills: AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
10 Days AgoSaved
In-Office or Remote
Seattle, WA
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Develop and maintain code and automation for highly scalable M365 sovereign cloud services. Operate live sites, participate in on-call rotations, troubleshoot incidents, improve observability, design automation for deployments, and collaborate with product and security stakeholders to ensure reliability and compliance.
Top Skills: AzureCC#C++CopilotExchange Online ProtectionExchange TransportGenerative AiJavaJavaScriptMicrosoft 365Microsoft Defender For OfficeOffice 365OnedrivePurviewPythonSharepointTeams
10 Days AgoSaved
In-Office
Seattle, WA
102K-219K Annually
Mid level
102K-219K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Own reliability and operational health for Substrate services in regulated environments. Respond to incidents as on-call engineer, diagnose and fix production issues, implement automation, develop monitoring and telemetry for SLOs, lead post-incident reviews, and collaborate with engineering teams to embed reliability and security.
Top Skills: Department Of Defense EnvironmentsExchange OnlineGcc High (Gcch)Gcc Moderate (Gccm)M365 CopilotMicrosoft CloudMicrosoft SubstrateMonitoringOn-Call AutomationSlosTelemetry
10 Days AgoSaved
In-Office
Seattle, WA
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead reliability strategy and SRE best practices for Substrate services in regulated environments. Serve as a senior on-call engineer, lead incident response and post-incident remediation, architect large-scale automation, observability, and self-healing systems, influence cross-organizational design for reliability, security, and compliance, and mentor senior engineers while representing SRE perspectives to leadership.
Top Skills: AutomationDod)Exchange OnlineGcc HighMicrosoft 365Microsoft Government Cloud (Gcc ModerateMicrosoft SubstrateObservabilitySelf-HealingSlo Frameworks
10 Days AgoSaved
In-Office or Remote
Seattle, WA
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead qualification, performance validation, and production readiness for new Azure Storage hardware and firmware. Drive test planning, automation frameworks, large-scale telemetry and benchmark analysis, root-cause investigations across software/hardware/firmware, and partner with engineering and vendors to resolve reliability and performance issues.
Top Skills: Automation FrameworksAzureAzure StorageBenchmarkingFirmwareNetworkingSsdsTelemetry
Reposted 3 Hours AgoSaved
Remote
Seattle, WA
164K-220K Annually
Senior level
164K-220K Annually
Senior level
Robotics • Software
Own reliability across vehicle and cloud stacks for AUV operations: onboard Jetson/ROS2 compute, topside systems, cloud ingestion/processing and customer platform. Build automation, observability, runbooks, and self-recovery to reduce on-call toil; manage AWS infrastructure, IaC, container orchestration, and reliability targets. Participate in shared 12-hour on-call shifts and field deployments, mentor team on operational excellence.
Top Skills: AWSBashContainerizationDockerGoGrafanaIamJetsonKubernetesLinuxPrometheusPythonRosRos 2Terraform
Reposted YesterdaySaved
In-Office or Remote
Seattle, WA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills: AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Reposted YesterdaySaved
In-Office or Remote
Seattle, WA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills: AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Reposted YesterdaySaved
In-Office or Remote
Seattle, WA
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills: AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Reposted YesterdaySaved
In-Office or Remote
Seattle, WA
76K-136K Annually
Mid level
76K-136K Annually
Mid level
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills: Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted YesterdaySaved
Remote
Seattle, WA
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills: AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Reposted YesterdaySaved
In-Office or Remote
Seattle, WA
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability, automation, and observability for high-density AI hardware infrastructure. Build Python-based IaC tooling, telemetry pipelines, Prometheus/Grafana dashboards, and AI-assisted tooling. Run 24x7 incident response, coordinate vendors and field technicians, define operational readiness, and drive post-mortems to improve uptime and performance.
Top Skills: Bare-MetalBgpGrafanaIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonSlackTimeseries EnginesVirtualized Environments
Reposted YesterdaySaved
Remote
Seattle, WA
Senior level
Senior level
Software
Drive SRE practices for VA enterprise healthcare platforms: automate infrastructure and CI/CD, define SLIs/SLOs, improve observability and reliability, support incident response, and ensure cloud-native, secure, compliant operations in AWS and containerized environments.
Top Skills: AnsibleAWSBashCi/CdCloudwatchDockerEcsEksElkGoGrafanaInfrastructure As CodeKubernetesLinuxOpentelemetryPowershellPrometheusPythonSplunkTerraform
Reposted YesterdaySaved
Remote
Seattle, WA
180K-224K Annually
Senior level
180K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills: Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
2 Days AgoSaved
Remote
Seattle, WA
165K-195K Annually
Senior level
165K-195K Annually
Senior level
Fintech • Real Estate • Software
Lead reliability and observability efforts across the org: design and maintain Kubernetes and AWS infrastructure, build CI/CD pipelines, drive IaC standards (Terraform/Crossplane), partner with 16+ teams to roll out tools and processes, participate in on-call rotation and incident response, and use AI tools to accelerate work.
Top Skills: Ai ToolsArgoAurora PostgresAWSCi/Cd PipelinesCrossplaneDatadogDocumentdb (Mongo)EcsEksGithub ActionsHelmKubernetesMongoDBPostgresRdsTerraform
11 Days AgoSaved
Hybrid
Seattle, WA
170K-220K Annually
Senior level
170K-220K Annually
Senior level
Artificial Intelligence • Legal Tech • Software • Generative AI
Lead and own the release and deployment process, manage GitHub workflows and Actions, build and maintain AWS infrastructure and observability, automate deployments and internal tooling, respond to incidents and be on-call, contribute code to reliability tooling, and support global/offshore teams across time zones.
Top Skills: AWSBashChatgptCi/CdClaudeEc2GitGithub ActionsIamLambdaMetricsObservability (LogsPostgres SqlPythonRdsTraces)TypescriptVpc
Reposted 11 Days AgoSaved
In-Office
Seattle, WA
180K-240K Annually
Senior level
180K-240K Annually
Senior level
Artificial Intelligence • Software • Generative AI
Lead reliability, scalability, and operational health of a production platform. Evolve Kubernetes, CI/CD, IaC, and observability. Build tooling and automation, improve monitoring/incident response, partner with engineering to identify and mitigate scaling risks, and influence platform direction across reliability, security, performance, and cost.
Top Skills: Ci/CdCloud-Native ArchitectureContainer OrchestrationGitopsGpu ProvisioningIncident ResponseInfrastructure As CodeKubernetesLoggingMetricsMulti-CloudObservabilityPythonTracingTypescript
Reposted 2 Days AgoSaved
Remote
Seattle, WA
150K-195K Annually
Senior level
150K-195K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills: DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Reposted 12 Days AgoSaved
In-Office
Seattle, WA
165K-230K Annually
Senior level
165K-230K Annually
Senior level
Aerospace • Other
Design, deploy, and automate infrastructure for on‑prem and cloud compute. Manage core services (databases, monitoring, storage), collaborate with software teams to build scalable, operable systems, and own the service lifecycle from design through deployment, operation, and refinement to ensure secure, reliable, and autonomous satellite software services.
Top Skills: AnsibleBashBazelCloudDatabasesKubernetesLinuxMakefilesMonitoringPythonTcp/IpTerraform
Reposted 3 Days AgoSaved
In-Office or Remote
Seattle, WA
Senior level
Senior level
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills: BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
Reposted 3 Days AgoSaved
Remote
Seattle, WA
117K-181K Annually
Senior level
117K-181K Annually
Senior level
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills: AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
4 Days AgoSaved
Remote
Seattle, WA
200K-270K Annually
Expert/Leader
200K-270K Annually
Expert/Leader
Social Media • Software
Design, implement, and operate infrastructure for a federated social network. Own reliability, availability, observability, incident response, deployments, capacity planning, and cost management. Build automation and tooling, scale bare-metal and cloud systems for millions of users, lead incident reviews, mentor engineers, and manage vendor relationships to ensure operational excellence.
Top Skills: Bare-MetalCapacity PlanningCloud ServicesColocationDatabasesDeployment And Rollback SystemsGoIncident ResponseKubernetesLinuxMonitoringNetworkingObservability SystemsProduction AutomationStorage
Reposted 13 Days AgoSaved
In-Office
Seattle, WA
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
The role involves building and managing observability infrastructure in GCP, automating deployments, and optimizing data processes for high reliability.
Top Skills: GkeGoGCPGrafanaKubernetesOpentelemetryPythonRubySplunkTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account