Onebrief Logo

Onebrief

Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided

Reposted 20 Days Ago
Remote
Hiring Remotely in United States
180K-220K Annually
Senior level
Remote
Hiring Remotely in United States
180K-220K Annually
Senior level
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
The summary above was generated by AI
Consequential Work. Dedicated People.
About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.

Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Security Clearance, Location, and Onsite Notice:

This is a hybrid role, requiring regular work on-site at customer locations in Arlington, VA - about 50/50 on-site vs remote.

If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).

Active Secret Clearance required.

About The Role

We're hiring a Site Reliability Engineer to join our Infrastructure & Security team. You'll work closely with product engineers, fellow SREs, security, and customer success.

This is an SRE role for someone who's comfortable in application code. Much of the reliability and performance work happens in the codebase (primarily TypeScript), so you'll fix problems at the source rather than working around them in the infrastructure. You'll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the field will feed directly back into the product.

You'll ship code that makes Onebrief more stable, faster, and easier to deploy and operate. The work sits at the seam between engineering and operations, and it's weighted toward engineering.

About You

You treat reliability as a feature, not an afterthought, and you'd rather fix a problem in the code than route around it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into each step.

You're comfortable reading and writing application code, and you're just as happy dropping into a kubectl shell to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and clear runbooks are part of building software, not extra credit.

You mentor others and push a culture of blameless postmortems. You work naturally with product and platform teams, helping them move fast without breaking things by giving them the tools, tests, and observability that make quick recovery real.

What You'll Do

You'll help make our production application reliable, scalable, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like:

  • Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. You'll partner with product engineers on design decisions, review code with reliability and security in mind, and treat "make the app better" as a first-class part of the job rather than something you hand off.

  • Building observability that developers actually use: Design and run our monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). The goal is alerts and dashboards tied to real application behavior, so teams catch issues before users do.

  • Owning reliability targets: Define and measure SLIs and SLOs, wire up alerting that feeds them, and be the person who can say what "reliable" means for our systems and prove it with data.

  • Leading incident response: Act as incident responder, and incident commander when needed. Run blameless post-mortems (AARs) that find the actual root cause and turn it into a code or process fix so it doesn't happen again.

  • Automating away toil: Spot the repetitive operational work and write software to kill it. Share what works with other teams, including those running in air-gapped environments, and help them get production-ready.

What We Look For
  • An active Secret clearance

  • 5+ years in software engineering, SRE, or a related role, with real time spent writing and shipping application code

  • Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript)

  • Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage

  • Experience with incident response, root cause analysis, and turning findings into lasting fixes

  • A collaborator who works well across product, platform, and DevOps teams and shares context openly

Technical expertise
  • Application development in TypeScript (Node and/or a modern front-end framework)

  • CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins)

  • Testing and quality practices as part of the delivery process

  • Comfort with at least one of Python, Go, or Bash for tooling and automation

  • Working knowledge of containers and Kubernetes (enough to debug and deploy, not necessarily to stand up clusters from scratch)

  • Networking fundamentals and secure configuration basics

Bonus points (nice to have)
  • Observability: Grafana stack, ELK, or Datadog

  • Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud)

  • Kubernetes cluster design and operations

  • Designing meaningful SLIs/SLOs with error budgets for distributed systems

  • GitOps practices and toolchains

  • DoD environments and compliance frameworks (RMF, STIGs, ICD 503)

  • Service mesh (Istio, Linkerd)

  • On-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V)

  • Relevant certs (AWS DevOps Engineer, CKA/CKAD)


Notice to Third Party Recruitment Agencies

Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.

Similar Jobs at Onebrief

2 Days Ago
Remote
United States
126K-154K Annually
Mid level
126K-154K Annually
Mid level
Software • Defense
The Procurement Specialist manages corporate indirect procurement operations, including vendor contract cleanup, record maintenance, renewals, onboarding, purchase requests, approvals, purchase orders, and vendor documentation. They operate Zip, maintain data flows with NetSuite, track compliance and renewals, support negotiations, train stakeholders on procurement policies, and optimize workflows. The role also supports spend tracking, budgeting, forecasting, and cross-functional coordination with Finance, IT, GRC, Legal, and business units.
Top Skills: AirbaseCoupaNetSuiteProcurifyZip
12 Days Ago
Remote
United States
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Software • Defense
Own and improve Onebrief’s corporate security stack across endpoints, identity, SaaS, browsers, Zero Trust, EDR, SIEM, and MDM. Establish secure configuration baselines, detect and remediate drift, integrate telemetry, and build API-driven or infrastructure-as-code automation. Translate CMMC 2.0 and NIST-aligned requirements into continuously validated technical controls and reliable audit evidence. Partner with Corporate IT, Security Operations, GRC, and application owners to strengthen security posture and reduce manual compliance work.
Top Skills: Configuration ManagementCrowdstrikeEdrGithub ActionsInfrastructure-As-CodeMdmMfaOktaOkta WorkflowsRest ApisSIEMSplunkSsoWebhooksWorkspace OneZero TrustZscaler
16 Days Ago
Remote
United States
134K-180K Annually
Entry level
134K-180K Annually
Entry level
Software • Defense
Own the credibility of AtomEngine’s military simulation entity catalog. Lead catalog strategy, research standards, quantitative performance modeling, AI-directed content production, verification and validation, automated testing, competitive wargames, documentation, and analyst mentoring. Partner with customers, engineers, and subject matter experts to prioritize catalog development, defend modeling assumptions, identify inaccuracies, and continuously improve simulation data quality.
Top Skills: Agentic Ai WorkflowsAtomengineAutomated TestingBug-Tracking SystemsGame Modding ToolsSimulation Authoring EnvironmentsVersion Control Systems

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account