SpaceX Logo

SpaceX

Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

Posted An Hour Ago
Be an Early Applicant
In-Office
Redmond, WA, USA
165K-270K Annually
Senior level
In-Office
Redmond, WA, USA
165K-270K Annually
Senior level
Design, deploy, operate, and scale GPU and CPU infrastructure supporting classified national security missions. Manage Kubernetes and AI clusters, Linux systems, databases, monitoring, distributed storage, virtualization, and GPU-as-a-service platforms. Build automation for large-scale on-premise environments, improve reliability and performance, collaborate with AI engineers and customers, and mentor junior engineers. Lead technical decisions and service lifecycles while supporting high availability across Top Secret data centers.
The summary above was generated by AI

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

SR. SITE RELIABILITY ENGINEER (STARSHIELD) 

At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation. Starshield is the world’s largest US government satellite constellation and is tasked with providing immediate access to critical intelligence and national security data for the US government anywhere on the globe. We design, build, test, and operate all parts of the system – receivers that allow users to connect within minutes, and the software that brings it all together. We’ve only begun to scratch the surface of Starshield's global impact and are looking for best-in-class engineers to help us further our ambitious goals.

As an engineer focused on Starshield's software and GPU infrastructure, you will design, operate and scale the infrastructure which supports critical national security missions. These positions cover a variety of areas ranging from Site Reliability Engineering, Developer Operations, and GPU platforms. You will develop automation to deploy and manage on-premise compute resources, create highly scalable and maintainable software products, and directly collaborate with engineering across the board.

RESPONSIBILITIES: 

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes\AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
  • Monitoring and alerting supporting systems to have high availability
  • Identify areas for improvement and create innovative solutions that enable high system availability
  • Mentor and train junior engineers
  • As a senior engineer you must lead the team to technical excellence – your decisions guide the team

BASIC QUALIFICATIONS:

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux operating systems; OR 7+ years of professional experience in software, DevOps, or site reliability engineering in lieu of a degree
  • 5+ year of experience with Kubernetes
  • 5+ year of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

PREFERRED SKILLS AND EXPERIENCE:

  • 5+ years of experience with Python and Python-based development frameworks
  • Experience managing Kubernetes clusters, not just using them
  • Knowledge of Linux boot process and systems configuration
  • Deep understanding of testing, continuous integration, build, deployment & continuous monitoring
  • Understanding of relevant build technologies, such as Bazel and Makefiles
  • Focus on performance bottlenecks and performance improvement techniques
  • Understanding of distributed databases and data modeling
  • Experience with automatically managing thousands of servers (eg: Terraform or Ansible)
  • Strong networking knowledge of TCP/IP
  • Cloud virtualization experience (as the cloud provider)
  • Experience using NVIDIA GPU deployment stacks (Blackwell/Rubin)
  • Excellent communications skills with the ability to communicate with customers, peers, management etc. in both formal and informal situations
  • Active Top Secret, Top Secret SCI, or DOE Level Q clearance

ADDITIONAL REQUIREMENTS:

  • Must be willing to work extended hours and weekends as needed
  • Must be willing to travel domestically and globally in the future when needed
  • This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment. While the clearance may not be immediately necessary upon hire, we encourage you to initiate the application process promptly upon accepting this offer. Your ability to secure the necessary clearance is essential for fulfilling key responsibilities of the role. Should you be unable to obtain it, SpaceX reserves the right to modify or terminate your employment to align with operational needs.

COMPENSATION AND BENEFITS:
Pay Range:
Level 3: $165,000.00 - $270,000.00

Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience.

Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees in Washington State accrue paid sick time in compliance with state and federal law. Company shuttles are offered to employees for roundtrip travel from select Seattle locations to the SpaceX Redmond office Monday to Friday.

Those with an active clearance will receive a 10% differential, up to an additional $20,000 annually, once officially briefed into a classified program.

ITAR REQUIREMENTS:

  • To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.  

SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.

Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to [email protected]

SpaceX Redmond, Washington, USA Office

22630 NE Marketplace Dr, Redmond, Washington, United States, 98053

SpaceX Seattle, Washington, USA Office

Seattle, United States

Similar Jobs

54 Minutes Ago
Hybrid
Seattle, WA, USA
150K-438K Annually
Senior level
150K-438K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads strategic tax planning and compliance for asset and wealth management clients. Oversees complex tax returns, analyzes financial data, develops optimization strategies, manages executive client relationships, drives business development, mentors tax professionals, promotes technology adoption, and provides thought leadership while ensuring regulatory compliance and minimizing risk.
2 Hours Ago
Hybrid
Kirkland, WA, USA
191K-334K Annually
Senior level
191K-334K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads strategy, roadmap, and execution for ServiceNow’s global cloud platform across commercial, sovereign, partner-hosted, and customer-operated environments. Aligns engineering, operations, security, legal, compliance, infrastructure, go-to-market, and customer teams. Evaluates market, regulatory, customer, and hyperscaler developments to guide investment decisions, supports strategic customer opportunities, and owns cross-functional initiatives through ambiguity. The role requires deep cloud and enterprise SaaS expertise, strong executive communication, and the ability to influence without authority.
Top Skills: Ai-Powered ToolsAWSEnterprise SaasGoogle Cloud PlatformAzureServicenow Cloud Platform
3 Hours Ago
Hybrid
Seattle, WA, USA
50-50 Hourly
Internship
50-50 Hourly
Internship
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
Software Engineering Interns will work on front-end, back-end, full-stack, security, or mobile projects. Responsibilities include scoping, designing, implementing, and shipping end-to-end features, collaborating with mentors and engineering teams, iterating quickly, and gaining experience with APIs, frameworks, coding agents, and programming languages while building products for millions of users.
Top Skills: APIsC#FrameworksGoJavaJavaScriptPythonTypescript

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account