Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Seattle, WA
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve Microsoft Defender services in US government cloud environments: provide 24x7 on-call support, run live-site incident response and postmortems, automate deployments and tooling, validate security/compliance during onboarding, collaborate with engineering teams, and apply software engineering best practices to increase reliability and observability at scale.
Top Skills:
C#Ci/CdCloudDistributed SystemsGoJavaMicrosoft DefenderObservabilityPython
Cloud • Software • Database
The Site Reliability Engineer will optimize and scale managed services across cloud providers, automate infrastructure, enhance monitoring, and ensure system reliability.
Top Skills:
AWSAzureBashGCPGrafanaKubernetesLokiMimirPrometheusPython
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills:
AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Develop and maintain code and automation for highly scalable M365 sovereign cloud services. Operate live sites, participate in on-call rotations, troubleshoot incidents, improve observability, design automation for deployments, and collaborate with product and security stakeholders to ensure reliability and compliance.
Top Skills:
AzureCC#C++CopilotExchange Online ProtectionExchange TransportGenerative AiJavaJavaScriptMicrosoft 365Microsoft Defender For OfficeOffice 365OnedrivePurviewPythonSharepointTeams
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Own reliability and operational health for Substrate services in regulated environments. Respond to incidents as on-call engineer, diagnose and fix production issues, implement automation, develop monitoring and telemetry for SLOs, lead post-incident reviews, and collaborate with engineering teams to embed reliability and security.
Top Skills:
Department Of Defense EnvironmentsExchange OnlineGcc High (Gcch)Gcc Moderate (Gccm)M365 CopilotMicrosoft CloudMicrosoft SubstrateMonitoringOn-Call AutomationSlosTelemetry
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead reliability strategy and SRE best practices for Substrate services in regulated environments. Serve as a senior on-call engineer, lead incident response and post-incident remediation, architect large-scale automation, observability, and self-healing systems, influence cross-organizational design for reliability, security, and compliance, and mentor senior engineers while representing SRE perspectives to leadership.
Top Skills:
AutomationDod)Exchange OnlineGcc HighMicrosoft 365Microsoft Government Cloud (Gcc ModerateMicrosoft SubstrateObservabilitySelf-HealingSlo Frameworks
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead qualification, performance validation, and production readiness for new Azure Storage hardware and firmware. Drive test planning, automation frameworks, large-scale telemetry and benchmark analysis, root-cause investigations across software/hardware/firmware, and partner with engineering and vendors to resolve reliability and performance issues.
Top Skills:
Automation FrameworksAzureAzure StorageBenchmarkingFirmwareNetworkingSsdsTelemetry
Reposted 3 Hours AgoSaved
Robotics • Software
Own reliability across vehicle and cloud stacks for AUV operations: onboard Jetson/ROS2 compute, topside systems, cloud ingestion/processing and customer platform. Build automation, observability, runbooks, and self-recovery to reduce on-call toil; manage AWS infrastructure, IaC, container orchestration, and reliability targets. Participate in shared 12-hour on-call shifts and field deployments, mentor team on operational excellence.
Top Skills:
AWSBashContainerizationDockerGoGrafanaIamJetsonKubernetesLinuxPrometheusPythonRosRos 2Terraform
Reposted YesterdaySaved
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills:
AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills:
AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills:
Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills:
AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Cloud • Security • Software • Cybersecurity
Lead reliability, automation, and observability for high-density AI hardware infrastructure. Build Python-based IaC tooling, telemetry pipelines, Prometheus/Grafana dashboards, and AI-assisted tooling. Run 24x7 incident response, coordinate vendors and field technicians, define operational readiness, and drive post-mortems to improve uptime and performance.
Top Skills:
Bare-MetalBgpGrafanaIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonSlackTimeseries EnginesVirtualized Environments
Software
Drive SRE practices for VA enterprise healthcare platforms: automate infrastructure and CI/CD, define SLIs/SLOs, improve observability and reliability, support incident response, and ensure cloud-native, secure, compliant operations in AWS and containerized environments.
Top Skills:
AnsibleAWSBashCi/CdCloudwatchDockerEcsEksElkGoGrafanaInfrastructure As CodeKubernetesLinuxOpentelemetryPowershellPrometheusPythonSplunkTerraform
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills:
Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Fintech • Real Estate • Software
Lead reliability and observability efforts across the org: design and maintain Kubernetes and AWS infrastructure, build CI/CD pipelines, drive IaC standards (Terraform/Crossplane), partner with 16+ teams to roll out tools and processes, participate in on-call rotation and incident response, and use AI tools to accelerate work.
Top Skills:
Ai ToolsArgoAurora PostgresAWSCi/Cd PipelinesCrossplaneDatadogDocumentdb (Mongo)EcsEksGithub ActionsHelmKubernetesMongoDBPostgresRdsTerraform
Artificial Intelligence • Legal Tech • Software • Generative AI
Lead and own the release and deployment process, manage GitHub workflows and Actions, build and maintain AWS infrastructure and observability, automate deployments and internal tooling, respond to incidents and be on-call, contribute code to reliability tooling, and support global/offshore teams across time zones.
Top Skills:
AWSBashChatgptCi/CdClaudeEc2GitGithub ActionsIamLambdaMetricsObservability (LogsPostgres SqlPythonRdsTraces)TypescriptVpc
Artificial Intelligence • Software • Generative AI
Lead reliability, scalability, and operational health of a production platform. Evolve Kubernetes, CI/CD, IaC, and observability. Build tooling and automation, improve monitoring/incident response, partner with engineering to identify and mitigate scaling risks, and influence platform direction across reliability, security, performance, and cost.
Top Skills:
Ci/CdCloud-Native ArchitectureContainer OrchestrationGitopsGpu ProvisioningIncident ResponseInfrastructure As CodeKubernetesLoggingMetricsMulti-CloudObservabilityPythonTracingTypescript
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills:
DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Aerospace • Other
Design, deploy, and automate infrastructure for on‑prem and cloud compute. Manage core services (databases, monitoring, storage), collaborate with software teams to build scalable, operable systems, and own the service lifecycle from design through deployment, operation, and refinement to ensure secure, reliable, and autonomous satellite software services.
Top Skills:
AnsibleBashBazelCloudDatabasesKubernetesLinuxMakefilesMonitoringPythonTcp/IpTerraform
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills:
BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
Reposted 3 Days AgoSaved
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills:
AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Social Media • Software
Design, implement, and operate infrastructure for a federated social network. Own reliability, availability, observability, incident response, deployments, capacity planning, and cost management. Build automation and tooling, scale bare-metal and cloud systems for millions of users, lead incident reviews, mentor engineers, and manage vendor relationships to ensure operational excellence.
Top Skills:
Bare-MetalCapacity PlanningCloud ServicesColocationDatabasesDeployment And Rollback SystemsGoIncident ResponseKubernetesLinuxMonitoringNetworkingObservability SystemsProduction AutomationStorage
Cloud
The role involves building and managing observability infrastructure in GCP, automating deployments, and optimizing data processes for high reliability.
Top Skills:
GkeGoGCPGrafanaKubernetesOpentelemetryPythonRubySplunkTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Seattle, WA Companies Hiring SRE Engineers
See AllPopular Seattle, WA Engineering Job Searches
Engineering Jobs in Seattle, WA
Software Engineer Jobs in Seattle, WA
Android Developer Jobs in Seattle, WA
C# Jobs in Seattle, WA
C++ Jobs in Seattle, WA
DevOps Jobs in Seattle, WA
Front End Developer Jobs in Seattle, WA
Golang Jobs in Seattle, WA
Hardware Engineer Jobs in Seattle, WA
iOS Developer Jobs in Seattle, WA
Java Developer Jobs in Seattle, WA
Javascript Jobs in Seattle, WA
Linux Jobs in Seattle, WA
Engineering Manager Jobs in Seattle, WA
.NET Developer Jobs in Seattle, WA
PHP Developer Jobs in Seattle, WA
Python Jobs in Seattle, WA
QA Jobs in Seattle, WA
Ruby Jobs in Seattle, WA
Salesforce Developer Jobs in Seattle, WA
Scala Jobs in Seattle, WA
Automation Engineer Jobs in Seattle, WA
AWS Engineer Jobs in Seattle, WA
Backend Engineer Jobs in Seattle, WA
Cloud Engineer Jobs in Seattle, WA
Controls Engineer Jobs in Seattle, WA
CTO Jobs in Seattle, WA
Design Engineer Jobs in Seattle, WA
DevOps Engineer Jobs in Seattle, WA
Director of Engineering Jobs in Seattle, WA
Electrical Engineering Jobs in Seattle, WA
Embedded Software Engineer Jobs in Seattle, WA
Field Engineer Jobs in Seattle, WA
Full-Stack Engineer Jobs in Seattle, WA
Infrastructure Engineer Jobs in Seattle, WA
Manufacturing Engineer Jobs in Seattle, WA
Mechanical Design Engineer Jobs in Seattle, WA
Mechanical Engineering Jobs in Seattle, WA
Network Engineer Jobs in Seattle, WA
Platform Engineer Jobs in Seattle, WA
Principal Engineer Jobs in Seattle, WA
Process Engineer Jobs in Seattle, WA
Product Engineer Jobs in Seattle, WA
Project Engineer Jobs in Seattle, WA
QA Engineer Jobs in Seattle, WA
Robotics Engineer Jobs in Seattle, WA
Security Engineer Jobs in Seattle, WA
Software Architect Jobs in Seattle, WA
Software Development Manager Jobs in Seattle, WA
Solutions Architect Jobs in Seattle, WA
Solutions Engineer Jobs in Seattle, WA
SRE Engineer Jobs in Seattle, WA
Staff Engineer Jobs in Seattle, WA
Staff Software Engineer Jobs in Seattle, WA
Systems Engineer Jobs in Seattle, WA
Web Developer Jobs in Seattle, WA
All Filters
Total selected ()
No Results
No Results
























