Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

SOFTWARE · Infrastructure / Hardware

DPU Silicon RAS and Debug Architect

Meta

Senior · 8+ yrsOn-site · Sunnyvale, CA⋅Austin, TX⋅Menlo Park, CAListed 1d ago
Apply now

Backed by

VC portfolio

HQ

🇺🇸 Menlo Park, CA, United States

Open roles

22

How to stand out for DPU Silicon RAS and Debug Architect at Meta

Auto Match agent

Let ORI find you the best jobs.

Set up your Auto Match agent once - your target role, level and where you want to work. It searches every day, scores each opening against your profile and resume, and delivers the ones worth applying to, with a prepared application a click away.

Searches every day Scored against your profile Applications prepared for you
Sign in & set up Auto Match agent

Resume & career call

Get your resume reviewed for this role - 30-minute 1:1 call

Line-by-line resume feedback for this application, how to position your Role Readiness, and a clear plan for what to do next - with a OneRoadmap career coach.

Experience Senior · 6+ yrs (8+ years)

About the role

structured by ORI

Skip to main content # DPU Silicon RAS and Debug Architect Sunnyvale, CA +2 locations Hardware \+ 1 more Apply now Meta's Infrastructure Silicon organization designs custom silicon that powers our data center infrastructure — SmartNICs/IPUs/DPUs, AI accelerators, and networking ASICs. We are looking for a Data…

What you will do

  • Own the RAS and Debug architecture for our DPUs - error detection, correction, and reporting, FIT budgeting, and the trace, debug, and telemetry infrastructure - from early path-finding through silicon bring-up
  • Set FIT-rate targets from the intended usages and deployment models, and budget them across the design - memories, logic, on-chip interfaces, and links
  • Define the error detection, correction, reporting, and containment architecture - parity, ECC/SECDED, poisoning and poison propagation, error logging, and how errors are surfaced to firmware, host software, and platform management
  • Define the trace, debug, and performance-monitoring architecture for post-silicon debug and for software and firmware debugging - on-chip trace, event and counter telemetry, crash and state capture, and the JTAG/debug-access model
  • Decide what belongs in hardware versus firmware versus software, and define how the hardware presents itself to them - register and programming models, error and interrupt models, and trace/telemetry interfaces

What they are looking for

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • Bachelor's degree in Computer Science, Computer Engineering or Electrical Engineering or equivalent work experience
  • 8+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
  • Understanding of RAS concepts: FIT-rate estimation and budgeting, failure modes (including silent data corruption), error detection, correction, and containment, and reliability targets for data center deployments.
  • Familiarity with data center reliability, serviceability, and manageability requirements
  • Understanding of error-protection mechanisms - parity, ECC/SECDED, CRC, data poisoning and poison propagation, and lockstep/redundancy techniques - and their area, latency, and power trade-offs
  • Understanding of memory and interface RAS LPDDR/DDR RAS (ECC, on-die ECC, error scrubbing, post-package repair) and PCIe RAS (Advanced Error Reporting, ECRC, link error detection and recovery)
  • Understanding of error reporting and handling architectures (ARM and/or x86) - error logging and registers, interrupts, machine-check and AER-style reporting, and escalation to firmware, host software, and platform/BMC management
  • Understanding of on-chip debug and trace architectures JTAG and debug access, on-chip trace, breakpoints and watchpoints, and crash and state capture (e.g., ARM CoreSight or comparable)
  • Understanding of performance-monitoring architectures - hardware performance counters and event telemetry - and their use in post-silicon and software/firmware debug
  • Familiarity with DFT concepts (scan, MBIST/LBIST, boundary scan) and how they interact with RAS and debug
  • Familiarity with processor ISA debug mechanisms (ARM and/or x86) and instruction/execution trace mechanisms such as ETM

Nice to have

  • Experience defining FIT budgets and reliability targets for hyperscale data center silicon, including soft-error rate (SER) analysis and mitigation
  • 15+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
  • Experience with hardware description languages (e.g., SystemVerilog, VHDL) and simulation environments used in ASIC development flows
  • Experience developing Python-based (or other scripting) automation pipelines for debug, telemetry collection, and data analysis
  • PhD in Computer Science, Computer Engineering or Electrical Engineering
  • Familiarity with functional-safety standards (e.g., ISO 26262) and reliability qualification methods
  • Experience designing error-reporting and machine-check/AER architectures and the firmware error-handling flows built on them
  • Experience designing on-chip trace and debug subsystems (e.g., ARM CoreSight) and the associated post-silicon debug tooling

Benefits

  • bonus
  • equity
RASDebugASICsPCIeJTAGARMx86ECCARM CoreSightETMSystemVerilogVHDLPython
Full posting text

Skip to main content # DPU Silicon RAS and Debug Architect Sunnyvale, CA +2 locations Hardware \+ 1 more Apply now Meta's Infrastructure Silicon organization designs custom silicon that powers our data center infrastructure — SmartNICs/IPUs/DPUs, AI accelerators, and networking ASICs. We are looking for a Data Processing Unit (DPU) RAS and Debug Architect to define the reliability, availability, and serviceability architecture for our DPU ASICs.

You'll set FIT-rate targets from the intended usages and deployment models, define how errors are detected, corrected, and reported, and define the trace, debug, and performance-monitoring architecture that supports both post-silicon debug and software and firmware debugging. You'll decide what belongs in hardware versus firmware versus software, define how the hardware presents itself to firmware and software, and work hand in hand with the firmware, driver, and RTL teams - carrying your designs from early path-finding through implementation and silicon bring-up.

* * * DPU Silicon RAS and Debug Architect Responsibilities Own the RAS and Debug architecture for our DPUs - error detection, correction, and reporting, FIT budgeting, and the trace, debug, and telemetry infrastructure - from early path-finding through silicon bring-up Set FIT-rate targets from the intended usages and deployment models, and budget them across the design - memories, logic, on-chip interfaces, and links Define the error detection, correction, reporting, and containment architecture - parity, ECC/SECDED, poisoning and poison propagation, error logging, and how errors are surfaced to firmware, host software, and platform management Define the trace, debug, and performance-monitoring architecture for post-silicon debug and for software and firmware debugging - on-chip trace, event and counter telemetry, crash and state capture, and the JTAG/debug-access model Decide what belongs in hardware versus firmware versus software, and define how the hardware presents itself to them - register and programming models, error and interrupt models, and trace/telemetry interfaces Work closely with design, DV, and PD teams on feature definition and PPA tradeoff; refine architecture to meet the design and PD constraints.

Define architecture to maximize DV complexity including defining specific features to help ease DV Support Design, DV, PD and DFT teams to resolve interface and integration issues as they come up. Support post-silicon bring-up Minimum Qualifications Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience Bachelor's degree in Computer Science, Computer Engineering or Electrical Engineering or equivalent work experience 8+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs Understanding of RAS concepts: FIT-rate estimation and budgeting, failure modes (including silent data corruption), error detection, correction, and containment, and reliability targets for data center deployments.

Familiarity with data center reliability, serviceability, and manageability requirements Understanding of error-protection mechanisms - parity, ECC/SECDED, CRC, data poisoning and poison propagation, and lockstep/redundancy techniques - and their area, latency, and power trade-offs Understanding of memory and interface RAS LPDDR/DDR RAS (ECC, on-die ECC, error scrubbing, post-package repair) and PCIe RAS (Advanced Error Reporting, ECRC, link error detection and recovery) Understanding of error reporting and handling architectures (ARM and/or x86) - error logging and registers, interrupts, machine-check and AER-style reporting, and escalation to firmware, host software, and platform/BMC management Understanding of on-chip debug and trace architectures JTAG and debug access, on-chip trace, breakpoints and watchpoints, and crash and state capture (e.g., ARM CoreSight or comparable) Understanding of performance-monitoring architectures - hardware performance counters and event telemetry - and their use in post-silicon and software/firmware debug Familiarity with DFT concepts (scan, MBIST/LBIST, boundary scan) and how they interact with RAS and debug Familiarity with processor ISA debug mechanisms (ARM and/or x86) and instruction/execution trace mechanisms such as ETM Knowledge of relevant industry standards and specifications, such as OCP (server, RAS, and telemetry/manageability specifications), JEDEC (LPDDR/DDR), and PCIe Demonstrated ability to drive analysis independently and influence architectural direction through data Preferred Qualifications Experience defining FIT budgets and reliability targets for hyperscale data center silicon, including soft-error rate (SER) analysis and mitigation 15+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs Experience with hardware description languages (e.g., SystemVerilog, VHDL) and simulation environments used in ASIC development flows Experience developing Python-based (or other scripting) automation pipelines for debug, telemetry collection, and data analysis PhD in Computer Science, Computer Engineering or Electrical Engineering Familiarity with functional-safety standards (e.g., ISO 26262) and reliability qualification methods Experience designing error-reporting and machine-check/AER architectures and the firmware error-handling flows built on them Experience designing on-chip trace and debug subsystems (e.g., ARM CoreSight) and the associated post-silicon debug tooling * * * About Meta Meta builds technologies that help people connect, find communities, and grow businesses.

When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.

For those who live in or expect to work from California if hired for this position, please click here for additional information. United States of America: $178,000/year to $250,000/year + bonus + equity + benefits Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable.

In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta. Equal Employment Opportunity Meta is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics.

You may view our Equal Employment Opportunity notice here. Meta is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, fill out the Accommodations request form. * * * Apply for this job Take the first step toward a rewarding career at Meta.

Apply now APPLY NOW Find your role Explore jobs that match your skills and experience. Search by technology, team or location to find an opening that’s right for you. View jobs Meta AI Recruiters can view your conversations with AI. Using AI is optional and does not impact the outcome of your application process. Learn more about AI usage and settings. Sign up to ask follow-up questions

Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…
Meta

Software Development

Meta's mission is to build the future of human connection and the technology that makes it possible. Our technologies help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive and intelligent experiences like virtual reality, wearables and AI to help build the next evolution in social technology.

To help create a safe and respectful online space, we encourage constructive conversation

Company pageWebsite

More at Meta

MetaInfrastructure / Data Center Design, Engineering, & ConstructionPreferredFunded

On-site · Reston, VA⋅Menlo Park, CA

1d ago

Data Center Infrastructure Management (DCIM) Engineer
View role
MetaData Center OperationsPreferredFunded

On-site · Henrico, VA⋅Mesa, AZ⋅Kansas City, MO⋅Ashburn, VA⋅Temple, TX⋅+21 more

Senior · 6+ yrs · 1d ago

Critical Facility Engineer
electricalHVACmechanical+7 more
View role
MetaAI Infrastructure / Facebook / Infrastructure / Data Center / Engineering / Hardware / Design / Artificial IntelligencePreferredFunded

On-site · Sunnyvale, CA⋅Austin, TX

Fresher · 1d ago

ASIC Engineer, Foundry and IP
View role
MetaArtificial Intelligence / AI Infrastructure / AI Research / Security / Technical Security / EngineeringPreferredFunded

On-site · Menlo Park, CA

Senior · 1d ago

Security Engineering Manager, Applied AI
View role
Apply