Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Seattle, WA
Reposted YesterdaySaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Financial Services
Design, deploy, and operate secure, highly available cloud infrastructure and container platforms. Build and maintain IaC (Terraform), CI/CD pipelines, monitoring and logging, security integrations, disaster recovery, and on-call support. Collaborate with developers to automate deployments and improve reliability, scalability, and operational efficiency.
Top Skills:
Aqua SecurityAWSAws CloudwatchAws CodepipelineAzureBashCircleCICloud FoundryDynatraceEc2EcsEksElastic StackElasticsearchElk StackGitlab CiGrafanaIamJavaJenkinsKibanaKubernetesLogstashNode.jsPrometheusPythonRdsS3ShellSnykSonarqubeSpinnakerSplunkSpring BootTerraformTrivyVpc
Insurance
As a Site Reliability Engineer II, you will build, test, and maintain the technology infrastructure for Openly's insurance platform, focusing on automation, monitoring, incident response, and operational decisions.
Top Skills:
Aiven DebeziumArcgisBigQueryCircleCICloud FunctionsCloud RunCloudsqlComposer/AirflowDatadogFivetranGcp GcsGitGoGCPJupyter NotebooksKafkaKubernetesNuxtPostgresPub/SubPythonRSQLTailwindTerraformVuejsWebpack
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Reposted 7 Days AgoSaved
Software • Defense
Own reliability, scalability, and security for on-prem and AWS deployments. Build observability (Prometheus/Loki/Grafana/ELK), define SLOs/SLIs, lead incident response and postmortems, automate infrastructure (Terraform/Ansible), operate Kubernetes clusters, embed security/compliance controls, eliminate operational toil, and mentor teams.
Top Skills:
AlloyAnsibleAWSAws GovcloudBashCloudFormationDatadogElkGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubernetesLokiPrometheusPythonRmfStigsTerraform
Reposted 7 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills:
Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 16 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 17 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
As a Cloud Cost Utilization SRE at GitLab, you'll manage cloud spending, improve tracking and optimization of cloud usage, and collaborate with finance and engineering teams to enhance cost efficiency across AWS and GCP.
Top Skills:
AnsibleAWSElkGCPGrafanaLokiMimirPrometheusTempoTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native, production-grade Kubernetes platforms and tooling to improve reliability, operability, and developer experience. Develop IaC and CI/CD automation, use observability to debug distributed systems, document self-service workflows, and influence engineering teams on scalable architectural patterns. Scope and deliver platform projects with focus on security, cost efficiency, and service stability.
Top Skills:
AksApmAWSAzureC#Ci/CdContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and develop SRE practices for secure, highly regulated Microsoft CISO services. Architect and automate hybrid/cloud infrastructure, manage petabyte-scale data platforms and pipelines, implement IaC and disaster recovery, deliver telemetry and automation, participate in on-call rotations, and mentor engineers to improve reliability, diagnosability, security, and compliance.
Top Skills:
AciAksArm TemplatesAWSAzureAzure BicepAzure Container AppsAzure Event HubsAzure Key VaultAzure SynapseAzure VmsBashBicepC#DockerGCPHadoopIaasJavaKubernetesMicrosoft 365 (ExchangePowershellPythonSharepointSkypeSparkTeams)Terraform
Big Data • Cloud
Design and automate scale and resilience tests for Qumulo's hybrid cloud storage platform. Build test frameworks and schedules, automate manual tests, troubleshoot failures across VMs and hardware, analyze cluster/C error logs, implement monitoring and alerts, participate in on-call, and help set release quality standards.
Top Skills:
AnsibleArgoAWSAzureCContainersGCPGrafanaInfluxdbJenkinsKubernetesNfsOpenmetricsPrometheusPythonS3SmbTerraformUbuntuVms
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
The Senior Site Reliability Engineer I will enhance Axon's observability platform, work on distributed tracing, log aggregation, and metrics infrastructure, and develop internal tools while collaborating with engineering teams.
Top Skills:
ArgocdCdkCortexGoGrafanaHelmJaegerJavaLokiOpentelemetryPrometheusPythonTerraform
Cloud • Information Technology • Internet of Things • Professional Services • Software
Design, build, and maintain automation and deployment tooling to improve reliability and scalability across global cloud environments. Troubleshoot distributed systems, develop CI/CD and testing frameworks, automate cluster and environment provisioning, and collaborate with engineering and product teams to reduce operational overhead and support large-scale platform growth.
Top Skills:
AnsibleGitlab CiLinuxRspecRuby
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Lead design and implementation of a Kubernetes-based self-service platform, driving infrastructure-as-code, GitOps practices, datastore provisioning, observability, and architectural standards. Partner with application teams to resolve performance bottlenecks and build scalable, automated platform solutions for large-scale distributed systems.
Top Skills:
AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesMicroservicesMySQLNew RelicPostgresPulumiTerraform
Artificial Intelligence • Software • Generative AI
Design, operate, and improve highly available, cloud-native systems for AI workloads. Own production reliability, incident response, on-call practices, observability, SLOs/SLIs, postmortems, and reliability automation across multi-region platforms.
Top Skills:
AWSAzureDistributed SystemsGCPGoGpu Cluster ManagementJavaKubernetesLogsMetricsOciPythonSlo/Sla FrameworksTracing
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Build and improve automation and tooling to detect, analyze, and mitigate live-site incidents for Azure Cosmos DB. Collaborate with engineering and customers to enhance telemetry, observability, and proactive alerting, perform automated root-cause analysis, and influence product architecture to meet strict SLOs.
Top Skills:
Azure Cosmos DbAzure Data FactoryAzure Event GridAzure PostgresqlAzure Service BusAzure Sql DbAzure Synapse AnalyticsJupyter NotebooksLogic AppsMeltMicrosoft FabricObservabilityPower BI
Healthtech • Software • Analytics • Business Intelligence
Lead and own reliability for critical backend and distributed systems: design, launch, on-call, incident leadership, SLO/SLI/error budget definition, automation to remove toil, observability improvement, resilience testing, mentoring, and cross-team reliability initiatives for production healthcare workflows.
Top Skills:
AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills:
AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
Reposted 23 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Fintech • Payments
The Senior Staff SRE leads reliability engineering initiatives, drives operational excellence, mentors staff, and influences architecture to enhance system reliability and performance.
Top Skills:
Ai/MlAWSAzureDockerElk StackGCPGrafanaKubernetesMySQLNoSQLPostgresSplunk
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills:
Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Seattle, WA Engineering Job Searches
Engineering Jobs in Seattle, WA
Software Engineer Jobs in Seattle, WA
Android Developer Jobs in Seattle, WA
C# Jobs in Seattle, WA
C++ Jobs in Seattle, WA
DevOps Jobs in Seattle, WA
Front End Developer Jobs in Seattle, WA
Golang Jobs in Seattle, WA
Hardware Engineer Jobs in Seattle, WA
iOS Developer Jobs in Seattle, WA
Java Developer Jobs in Seattle, WA
Javascript Jobs in Seattle, WA
Linux Jobs in Seattle, WA
Engineering Manager Jobs in Seattle, WA
.NET Developer Jobs in Seattle, WA
PHP Developer Jobs in Seattle, WA
Python Jobs in Seattle, WA
QA Jobs in Seattle, WA
Ruby Jobs in Seattle, WA
Salesforce Developer Jobs in Seattle, WA
Scala Jobs in Seattle, WA
Automation Engineer Jobs in Seattle, WA
AWS Engineer Jobs in Seattle, WA
Backend Engineer Jobs in Seattle, WA
Cloud Engineer Jobs in Seattle, WA
Controls Engineer Jobs in Seattle, WA
CTO Jobs in Seattle, WA
Design Engineer Jobs in Seattle, WA
DevOps Engineer Jobs in Seattle, WA
Director of Engineering Jobs in Seattle, WA
Electrical Engineering Jobs in Seattle, WA
Embedded Software Engineer Jobs in Seattle, WA
Field Engineer Jobs in Seattle, WA
Full-Stack Engineer Jobs in Seattle, WA
Infrastructure Engineer Jobs in Seattle, WA
Manufacturing Engineer Jobs in Seattle, WA
Mechanical Design Engineer Jobs in Seattle, WA
Mechanical Engineering Jobs in Seattle, WA
Network Engineer Jobs in Seattle, WA
Platform Engineer Jobs in Seattle, WA
Principal Engineer Jobs in Seattle, WA
Process Engineer Jobs in Seattle, WA
Product Engineer Jobs in Seattle, WA
Project Engineer Jobs in Seattle, WA
QA Engineer Jobs in Seattle, WA
Robotics Engineer Jobs in Seattle, WA
Security Engineer Jobs in Seattle, WA
Software Architect Jobs in Seattle, WA
Software Development Manager Jobs in Seattle, WA
Solutions Architect Jobs in Seattle, WA
Solutions Engineer Jobs in Seattle, WA
SRE Engineer Jobs in Seattle, WA
Staff Engineer Jobs in Seattle, WA
Staff Software Engineer Jobs in Seattle, WA
Systems Engineer Jobs in Seattle, WA
Web Developer Jobs in Seattle, WA
All Filters
Total selected ()
No Results
No Results





.png)

























