A client of BEARCloud is seeking an experienced Senior NPU Architect to help design and optimize next-generation AI acceleration technology. This role will work across architecture, hardware, and software teams to develop high-performance NPU solutions leveraging advanced in-memory computing.
The ideal candidate brings strong expertise in computer architecture and microarchitecture, with experience designing AI accelerators, NPUs, GPUs, SoCs, or other high-performance compute architectures.
What You'll Do
Design and define NPU architecture and microarchitecture, including compute engines, memory subsystems, and on-chip networks.
Develop architectural features based on AI workload requirements and hardware/software performance needs.
Build and enhance architectural and performance models using C++ simulation frameworks.
Analyze performance, identify bottlenecks, and optimize compute, memory movement, and hardware utilization.
Evaluate modern AI/ML workloads including LLMs, diffusion models, and CNNs to guide architectural decisions.
Partner with design and verification teams to develop validation and testing strategies for new architectural features.
Collaborate across architecture, hardware, software, and AI teams on hardware/software co-design.
Research emerging AI accelerator technologies and contribute to future product architecture.
What We're Looking For
BS or MS in Electrical Engineering, Computer Science, Computer Engineering, or a related field with approximately 5–8 years of relevant industry experience.
Ph.D. with 2–4 years of applicable experience is a plus.
Strong knowledge of computer architecture, microarchitecture, and digital design.
Experience with NPUs, AI accelerators, GPUs, custom silicon, SoCs, or high-performance compute architectures.
Understanding of AI/ML algorithms, frameworks, and workloads.
Strong programming experience with C/C++ and Python.
Experience with Verilog or SystemVerilog.
Knowledge of memory architecture, on-chip interconnects/networking, and performance modeling.
Ability to collaborate effectively across architecture, RTL/design, verification, software, and AI teams.
Ideal Background
You'll be especially well suited for this role if you've worked on NPU or AI accelerator architecture, performance modeling, memory systems, on-chip networks, or custom silicon and enjoy solving performance challenges at the intersection of AI, hardware, and software.
Numbers & Facts
Location
San Francisco, California (Remote)
Skills
Algorithmsunmatched
Architectural Servicesunmatched
Artificial Intelligence (AI)unmatched
C Programming Languageunmatched
C++ Programming Languageunmatched
Computer Architectureunmatched
Computer Engineeringunmatched
Computer Programmingunmatched
Computer Scienceunmatched
Design Verificationunmatched
Electrical Engineeringunmatched
Emerging Technologyunmatched
GPU (Graphics Processing Unit)unmatched
Graphic Designunmatched
Hardware Architectureunmatched
Memory Hardwareunmatched
Memory Subsystemunmatched
Network Performance/Analysisunmatched
Performance Analysisunmatched
Performance Modelingunmatched
Python Programming/Scripting Languageunmatched
RTL Designunmatched
Simulationunmatched
Software Designunmatched
SystemVerilogunmatched
Team Playerunmatched
Test Plan/Scheduleunmatched
Test Strategyunmatched
Validation Testingunmatched
Verilog Hardware Description Languageunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.