Own the RAS and debug architecture for our DPUs - error detection, correction, and reporting, FIT budgeting, and the trace, debug, and telemetry infrastructure - from early path-finding through silicon bring-up Set FIT-rate targets from the intended usages and deployment models, and budget them across the design - memories, logic, on-chip interfaces, and links Define the error detection, correction, reporting, and containment architecture -- parity, ECC/SECDED, poisoning and poison propagation, error logging, and how errors are surfaced to firmware, host software, and platform management Define the trace, debug, and performance-monitoring architecture for post-silicon debug and for software and firmware debugging - on-chip trace, event and counter telemetry, crash and state capture, and the JTAG/debug-access model Decide what belongs in hardware versus firmware versus software, and define how the hardware presents itself to them - register and programming models, error and interrupt models, and trace/telemetry interfaces Work closely with design, DV, and PD teams on feature definition and PPA tradeoff; refine architecture to meet the design and PD constraints. Support post-silicon bring-upBachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 8+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs Experience with RAS concepts: FIT-rate estimation and budgeting, failure modes (including silent data corruption), error detection, correction, and containment, and reliability targets for data center deployments.