Lead the architecture and hands-on development of distributed AI inference runtime systems. Responsibilities include scheduling, batching, KV-cache and memory management, distributed execution, kernel optimization, profiling, benchmarking, and performance analysis. Enable new model architectures, improve latency and throughput, develop validation and safe-rollout systems, and collaborate with infrastructure, compiler, hardware, cloud, and research teams. Mentor engineers and establish performance-engineering practices for Arm’s production AI platform.
As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models.
You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm's AI platform.
Responsibilities:
Required Skills and Experience :
"Nice To Have" Skills and Experience :
In Return:
You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities to support AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. The success of our AI projects will be directly influenced by your work, crafting the company's AI inference capabilities and defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!
Salary Range:
$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at Arm
At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [email protected] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.
Hybrid Working at Arm
Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.
Equal Opportunities at Arm
Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don't discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm's AI platform.
Responsibilities:
- Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms.
- Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.
- Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.
- Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.
- Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices.
Required Skills and Experience :
- 5+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.
- Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.
- Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, handling of memory resources, and data movement.
- Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.
"Nice To Have" Skills and Experience :
- Experience developing or modifying inference schedulers, cache managers, batching systems, disaggregated or distributed execution paths.
- Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution.
- Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization.
- Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries.
In Return:
You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities to support AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. The success of our AI projects will be directly influenced by your work, crafting the company's AI inference capabilities and defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!
Salary Range:
$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at Arm
At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [email protected] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.
Hybrid Working at Arm
Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.
Equal Opportunities at Arm
Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don't discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Similar Jobs at Arm
Artificial Intelligence • Internet of Things • Semiconductor
Leads project management across AI and Developer Platforms, establishing scalable operating models, planning frameworks, governance, risk management, and delivery standards. Manages and develops a distributed Project Management team, oversees portfolio and execution rhythms, aligns cross-functional stakeholders, resolves trade-offs, and provides leadership visibility into progress, dependencies, risks, resources, and outcomes. Uses data, automation, dashboards, and AI to improve delivery efficiency and organizational decision-making.
Top Skills:
Ai PlatformsDashboardsDelivery Data And MetricsDeveloper ToolingSoftware InfrastructureWorkflow Automation
Artificial Intelligence • Internet of Things • Semiconductor
Build and operate secure, scalable AI compute platform services, including backend APIs, asynchronous workflows, Kubernetes controllers, and multi-tenant workload capabilities. The role focuses on distributed training and inference, platform reliability, observability, developer tooling, and infrastructure simplification. Responsibilities include collaborating across engineering teams, improving scalability and developer productivity, and providing technical leadership for production AI platform capabilities.
Top Skills:
APIsAsynchronous ProcessingContainersGitopsGoGrafanaKafkaKubernetesKubernetes OperatorsOpentelemetryPostgresPrometheusPythonPyTorchRayRelational DatabasesRpcVllm
Artificial Intelligence • Internet of Things • Semiconductor
Design and build secure, scalable platform services for distributed AI training and inference. Develop backend APIs, Kubernetes controllers, orchestration systems, asynchronous workflows, multi-tenant access controls, and developer tools. Improve reliability, observability, scalability, and engineering productivity while partnering with infrastructure and AI teams. The role also provides technical leadership, mentoring, and cross-functional collaboration throughout platform development and production delivery.
Top Skills:
ContainersGitopsGoGrafanaKafkaKubernetesOpentelemetryPostgresPrometheusPythonPyTorchRayRelational DatabasesRpcVllm
What you need to know about the Seattle Tech Scene
Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.
Key Facts About Seattle Tech
- Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Amazon, Microsoft, Meta, Google
- Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Madrona, Fuse, Tola, Maveron
- Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

