Hi everyone,
I’m trying to diagnose a recurring stability issue with a new RX 9070 XT system, and I’d really like to hear from other RX 9000 owners who have experienced something similar.
At this point I’m mainly trying to figure out whether this looks like an RDNA4/driver/Windows scheduling issue, or if I should start suspecting a marginally faulty GPU.
System:
- Ryzen 7 7800X3D
- ASUS TUF Gaming B850-PLUS WIFI
- ASUS Prime RX 9070 XT OC 16 GB
- 32 GB Patriot DDR5-6000 CL28, EXPO enabled
- ASUS TUF Gaming 850G EVO PSU
- ASUS ROG XG27ACMES, 1440p / 255 Hz via DisplayPort
- BIOS 1681, latest stable version
- PCIe set to Auto
- iGPU disabled in BIOS
Drivers tested:
Originally AMD Adrenalin 26.8.1.
After the first major crashes I used AMD Cleanup/DDU and did a clean rollback to 26.7.1.
Unfortunately, the problem continued on 26.7.1, so it doesn’t seem to be a simple 26.8.1 regression.
Chipset driver is 8.08.12.551. I also found and fixed a previously missing AMD PCI device driver (DEV_14DE), now using AMD PCI driver 1.0.0.90.
Crashes so far:
August 29 – 26.8.1
Complete hard freeze / DisplayPort signal loss.
Windows later recorded:
VIDEO_ENGINE_TIMEOUT_DETECTED (141)
amdkmdag.sys
August 31 – long idle
Complete hard freeze again.
LiveKernelEvent:
A1000001
AMD Watchdog-related entries involving:
amdfendr
dxgmms2 scheduler
August 31 – PUBG, now on 26.7.1
Driver timeout occurred, but this time Windows successfully recovered instead of completely freezing.
Another:
LiveKernelEvent 141
amdkmdag
September 1
Full system freeze followed by automatic reboot.
This time I got a real BSOD dump:
DPC_WATCHDOG_VIOLATION (0x133)
Subtype:
DPC_QUEUE_EXECUTION_TIMEOUT_EXCEEDED
WinDbg bucket:
0x133_ISR_amdkmdag!unknown_function
amdkmdag appears multiple times in the watchdog/DPC path.
What makes this strange:
The GPU seems completely happy under heavy synthetic load.
I ran:
- OCCT 3D Adaptive Variable, 30 min: PASS
- OCCT Adaptive Switch 25% ↔ 100%, 30 min: PASS
- OCCT VRAM 80%, 30 min: PASS
No errors, no TDR, no overheating.
I also played Apex Legends DX12 for an extended session at roughly 300 FPS with high GPU utilization and normal temperatures, completely stable.
So this does not seem to be a simple case of:
"high GPU load = crash"
It can survive sustained high load and rapid load changes, while sometimes crashing during gaming or even after long idle periods.
RAM / EXPO testing:
Because EXPO/IMC instability was another possibility, I ran a full MemTest86 test with EXPO actually active:
DDR5-6000 MT/s
CL28
Result:
48/48 tests passed
4/4 passes completed
0 errors
~2.5 hours
So RAM/EXPO instability has moved much lower on my suspect list.
I also had two crashes of OCCT’s AM5 Margin Probe, but both were the application itself crashing with exactly the same:
c0000005 Access Violation
and exactly the same offset:
0x10D42
while OCCT reported no actual memory errors.
Combined with the clean MemTest86 result, I currently suspect that may be an OCCT v17/preset issue rather than unstable RAM.
Things already ruled out / tested:
- iGPU disabled: issue still occurred
- Chrome hardware acceleration disabled: issue still occurred
- clean driver rollback: issue still occurred
- GPU temperature problem: no indication
- VRAM errors: none found
- sustained GPU load: stable
- rapid GPU load switching: stable
- PSU is 850 W with factory cabling and shows no obvious symptoms under heavy GPU load
I’m not claiming the GPU itself is definitely good. A marginal hardware fault could still behave this way.
Next tests:
My next clean A/B test will be:
HAGS OFF
Nothing else will be changed.
If it crashes again with another 141 / A1000001 / 0x133 / hard freeze, my final driver test will probably be:
AMD Cleanup Utility in Safe Mode → internet disconnected → clean 26.3.1 WHQL Driver Only installation.
If it still happens on that driver, I’m planning to stop troubleshooting at home and send the complete PC back to the retailer for component A/B testing, especially testing with another GPU.
What I’d really like to know:
Has anyone with an RX 9070 XT / RX 9000 series seen the same combination of:
- LiveKernelEvent 141
- A1000001
- amdkmdag
- DPC_WATCHDOG_VIOLATION 0x133
- hard freeze / black screen / DisplayPort signal loss
- crashes during idle or relatively light transitions despite heavy stress tests being stable?
If you did:
Did HAGS OFF help?
Did an older driver actually fix it?
Did replacing the GPU fix it?
Or did the exact same problem continue with a replacement card?
I’ve found a few reports that look very similar, but I’m trying to separate general RDNA4 driver issues from genuinely defective/marginal cards.
Any dump analysis, driver experience, or reports from people who went through an RMA would be very useful.
Thanks!