Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Seattle, WA
Edtech • HR Tech • Software
Lead reliability engineering for critical SaaS services by defining SLOs, managing major incidents, improving observability, architecting scalable infrastructure, advancing deployment safety, and mentoring engineers. The role partners with engineering and product leaders to balance feature delivery with system resilience, while driving reliability standards, automation, and recruiting across the SRE organization.
Top Skills:
AWSContainer OrchestrationInfrastructure As Code
Artificial Intelligence • Software • Conversational AI • Automation
Build and improve Replicant’s AI-native platform, including site reliability, CI/CD, developer tooling, observability, incident management, cloud infrastructure, and autonomous-agent harnesses. Own production infrastructure reliability at scale, reduce operational toil, improve deployment workflows, participate in on-call rotation, and shape platform engineering patterns. The role uses TypeScript/Node.js, Python, Terraform, Kubernetes, Helm, GCP, and modern monitoring tools.
Top Skills:
ClaudeCursorDatadogFreeswitchGCPGitlab CiGrafanaHelmKubernetesLlmsNode.jsPrometheusPythonSipTerraformTypescript
Other
Build and improve Replicant’s AI-native platform infrastructure, including site reliability, CI/CD, developer tooling, observability, incident management, and agent harness systems. Own production infrastructure patterns, deployment workflows, and developer self-service across Kubernetes and cloud environments. Participate in on-call and incident response while improving scalability, availability, and operational efficiency for real-time conversational AI traffic.
Top Skills:
ClaudeCursorDatadogFreeswitchGCPGitlab CiGrafanaHelmKubernetesLlmsNode.jsPrometheusPythonSipTerraformTypescript
Artificial Intelligence • Big Data • Information Technology • Other • Software • Database • Biotech
Provides frontline support for high-availability cloud systems, monitors infrastructure, manages incidents, coordinates outage bridges, documents technical issues, and develops automation and AI agentic workflows. The role partners with development teams to improve detection and resolution, supports AWS technologies, databases, networks, CI/CD tools, and observability platforms. This is a Monday–Thursday overnight shift within a 24/7 command center and includes some holiday coverage.
Top Skills:
Ai Agentic WorkflowsAWSAws Application SignalsBashCassandraChefCi/CdCouchdbFirewallsHarnessIntrusion Detection SystemsJavaLan/WanLoad BalancersMssqlMySQLNew RelicPowershellPrompt EngineeringProxy ServersPuppetPythonTcp/IpVirtualization
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills:
AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Other
Design and operate cloud platforms supporting backend telecom services. Automate deployments, scaling, recovery, and infrastructure provisioning; monitor production systems; maintain observability, alerting, and dashboards; support incident response and on-call operations; manage CI/CD pipelines; and enable engineering, telecom, and data teams through reliable tools and infrastructure.
Top Skills:
AnsibleAWSAzureBashCassandraCircleCICloudFormationDatadogDnsDockerElasticsearchElk StackGitlab CiGoGCPGrafanaHttp/HttpsIamJaegerJenkinsKafkaKubernetesKvmLinuxNoSQLOpentelemetryPerlPrometheusPythonRubySaltstackSplunkSQLTcp/IpTerraformUnixVMware
Agency • Cloud • Professional Services • Software
Improve AWS production infrastructure reliability, observability, performance, and operational maturity. Build Terraform infrastructure, enhance CI/CD, automate operational work, manage incident response and on-call operations, lead postmortems, improve application resilience, support capacity planning and database reliability, and collaborate on security hardening and compliance. Mentor engineers and promote reliability practices across the organization.
Top Skills:
AWSCi/CdCircleCIDatadogGithub ActionsGitlab CiLinuxNew RelicPostgresRubyRuby On RailsSlisSlosTerraform
Artificial Intelligence • HR Tech • Professional Services • Software
Evaluate AI-generated documents, spreadsheets, and presentations for accuracy, rigor, domain quality, and presentation quality in incident management, reliability, and SRE. Apply specialized rubrics, identify factual and aesthetic errors, and provide structured written feedback. This is a fully remote, flexible independent-contractor engagement requiring at least five years of relevant professional experience, English fluency, and proficiency with Microsoft Office and Google Workspace.
Top Skills:
Google SlidesGoogle WorkspaceMS OfficePowerPoint
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills:
AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
Cloud
Build and operate reliable, scalable, secure infrastructure for SaaS security and Snowflake data systems. Automate infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and related tools. Ensure security and compliance, participate in on-call rotations, lead incident response and root-cause analysis, and collaborate with development, data science, and security teams on architectural decisions and service implementation.
Top Skills:
Artificial IntelligenceCi/CdFlywayInfrastructure As CodeKubernetesMachine LearningSnowflakeSpinnakerTerraform
Cloud
Designs, builds, and operates reliable, scalable infrastructure for security SaaS and Snowflake data systems. Automates infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and Flyway. Partners with security, development, and data science teams on secure, compliant architecture. Participates in on-call rotations, leads critical incident response and root-cause analysis, and implements preventative improvements. In-person onboarding and travel to the Toronto office are required during the first employment week.
Top Skills:
Ci/CdContainerizationFlywayInfrastructure As CodeKubernetesSnowflakeSpinnakerTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills:
AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills:
AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills:
Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills:
DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills:
AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Blockchain • Energy • Cryptocurrency
Hands-on role to assess, implement, test, and document backup, restore, failover, and recovery capabilities. Inventory critical systems, design and automate backup and restoration, run recovery exercises, produce runbooks, validate recoverability, measure RTO/RPO, and train system owners. Collaborate with Security, SRE, DevOps, QA, and application teams to harden shared recovery capabilities and transfer operational ownership.
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills:
.NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills:
Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Healthtech • Software • Analytics • Business Intelligence
Senior SRE responsible for designing, building, and operating reliable, scalable distributed systems; owning production reliability (SLOs/SLIs, incident response, MTTR reduction); automating toil with software and platform tooling; driving observability, capacity planning, and cross-team reliability improvements; mentoring engineers and running blameless postmortems.
Top Skills:
AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Aerospace • Other
Design, deploy, and scale on-prem Kubernetes clusters and core infrastructure for Starlink. Build automation, manage databases, monitoring, and distributed storage. Collaborate with engineers to improve service lifecycle, availability, and performance; troubleshoot across the Starlink stack and drive reliability improvements.
Top Skills:
AnsibleBashBazelC++GoKubernetesLinuxMakefilesOci ContainersPythonTcp/IpTerraform
Aerospace • Other
Design, deploy, and scale on‑premise compute and core infrastructure for Starlink. Develop automation, manage databases/monitoring/distributed storage, collaborate with software teams, troubleshoot end-to-end, and improve deployment and developer velocity.
Top Skills:
AnsibleBashCC++DatabasesDistributed StorageDockerGoHypervisor TechnologiesKubernetesLinuxMonitoringPythonTcp/IpTerraformVirtualization
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Seattle, WA Engineering Job Searches
Engineering Jobs in Seattle, WA
Software Engineer Jobs in Seattle, WA
Android Developer Jobs in Seattle, WA
C# Jobs in Seattle, WA
C++ Jobs in Seattle, WA
DevOps Jobs in Seattle, WA
Front End Developer Jobs in Seattle, WA
Golang Jobs in Seattle, WA
Hardware Engineer Jobs in Seattle, WA
iOS Developer Jobs in Seattle, WA
Java Developer Jobs in Seattle, WA
Javascript Jobs in Seattle, WA
Linux Jobs in Seattle, WA
Engineering Manager Jobs in Seattle, WA
.NET Developer Jobs in Seattle, WA
PHP Developer Jobs in Seattle, WA
Python Jobs in Seattle, WA
QA Jobs in Seattle, WA
Ruby Jobs in Seattle, WA
Salesforce Developer Jobs in Seattle, WA
Scala Jobs in Seattle, WA
Automation Engineer Jobs in Seattle, WA
AWS Engineer Jobs in Seattle, WA
Backend Engineer Jobs in Seattle, WA
Cloud Engineer Jobs in Seattle, WA
Controls Engineer Jobs in Seattle, WA
CTO Jobs in Seattle, WA
Design Engineer Jobs in Seattle, WA
DevOps Engineer Jobs in Seattle, WA
Director of Engineering Jobs in Seattle, WA
Electrical Engineering Jobs in Seattle, WA
Embedded Software Engineer Jobs in Seattle, WA
Field Engineer Jobs in Seattle, WA
Full-Stack Engineer Jobs in Seattle, WA
Infrastructure Engineer Jobs in Seattle, WA
Manufacturing Engineer Jobs in Seattle, WA
Mechanical Design Engineer Jobs in Seattle, WA
Mechanical Engineering Jobs in Seattle, WA
Network Engineer Jobs in Seattle, WA
Platform Engineer Jobs in Seattle, WA
Principal Engineer Jobs in Seattle, WA
Process Engineer Jobs in Seattle, WA
Product Engineer Jobs in Seattle, WA
Project Engineer Jobs in Seattle, WA
QA Engineer Jobs in Seattle, WA
Robotics Engineer Jobs in Seattle, WA
Security Engineer Jobs in Seattle, WA
Software Architect Jobs in Seattle, WA
Software Development Manager Jobs in Seattle, WA
Solutions Architect Jobs in Seattle, WA
Solutions Engineer Jobs in Seattle, WA
SRE Engineer Jobs in Seattle, WA
Staff Engineer Jobs in Seattle, WA
Staff Software Engineer Jobs in Seattle, WA
Systems Engineer Jobs in Seattle, WA
Web Developer Jobs in Seattle, WA
All Filters
Total selected ()
No Results
No Results



.png)

.png)























