Why do I keep having oops and kernel panics on a pangolin11 laptop?

I have been running a pangolin 11 laptop for a few months now. I enjoyed playing around with Arch Linux but decided I was crashing my computer too often with interrupted pacman -Syu calls.
Two days ago I switched to PopOs as that was what originally shipped with the laptop and is the OS System76 suggests for their users. After switching and starting to work on a project using opencode (with glm-5.3 as the model via Ollama cloud), I got a kernel panic and had to reboot. Here is what happened according to opencode after rebooting:

1. What happened (reconstructed timeline)

Time (PDT) Event Source
10:43:31 Boot into kernel 7.0.11 journal, boot -1
10:50:39 apt upgrade installs kernel 7.1.5 (not yet booted) apt history
12:40:47 Last user command: chown -R user:user ~/Projects (completed fine) journal
12:40:54 Last journal entry — no shutdown sequence follows journal
12:41:16 Kernel oops: NULL pointer write, CPU 0, in the idle task pstore (t=7074.703)
12:41:18 Kernel panic - not syncing: Attempted to kill the idle task! pstore (t=7076.994)
~12:41–12:42 Machine sat on the panic screen until power-cycled —
12:42:09 Reboot into the newly installed 7.1.5 journal, boot 0

The journal simply stops mid-session with no shutdown records — a hard, uncontrolled stop. The actual crash text survived because the EFI persistent store (efi_pstore) captured the kernel log buffer at oops/panic time, and systemd-pstore.service archived it to /var/lib/systemd/pstore/ on the next boot.

2. The actual crash

The pstore archive contains two records: the oops (“Oops#1”, 18 chunks) and the resulting panic (“Panic#2”, 18 chunks — the same buffer plus the panic lines). Reassembled, the essential part:

BUG: kernel NULL pointer dereference, address: 0000000000000000
#PF: supervisor write access in kernel mode
#PF: error_code(0x0002) - not-present page
Oops: Oops: 0002 [#1] SMP NOPTI
CPU: 0 UID: 0 PID: 0 Comm: swapper/0 Tainted: G  OE  7.0.11-76070011-generic PREEMPT(lazy)
Hardware name: System76 Pangolin/Pangolin, BIOS 1.07.27_S3 10/22/2024
RIP: 0010:srso_safe_ret+0x1a/0x20
Code: ... 00 <00> 00 0f 1f 00 e8 db ff ff ff 0f 0b ...
RAX: 0000000000000000  RBX: ffff8b31803ed400  ...
CR2: 0000000000000000
Call Trace:
? tick_nohz_stop_tick+0x63/0x260
? raw_spin_rq_unlock+0x14/0x50
? srso_return_thunk+0x5/0x5f
cpuidle_enter+0x2e/0x50
call_cpuidle+0x21/0x50
cpuidle_idle_call+0x173/0x200
do_idle+0x90/0xf0
...
---[ end trace ]---
Kernel panic - not syncing: Attempted to kill the idle task!

Interpretation. The faulting instruction is two zero bytes (00 00 = add %al,(%rax)) with RAX = 0 → a write to address 0. Those bytes are alignment padding inside the SRSO “safe return” thunk — part of the AMD Speculative Return Stack Overflow mitigation (CVE-2023-20569, “Inception”). No legitimate execution path ever runs there. This is a “wild control-flow transfer”: CPU 0 was entering the idle loop, a RET went somewhere it should never go, and execution landed in mitigation thunk padding. Had RAX not happened to be 0, this would have been silent memory corruption instead of a clean fault.
The panic 2.3 s later is the standard kernel response to an oops in the idle task: the oops handler tries to kill the faulting task, discovers it is PID 0 (swapper/0), which can never be killed, and panics with exactly the message above.

Okay, interesting. What is causing this to happen? I had a log of crashes from my previous Arch set-up on the same machine. Opencode found those logs in a directory and put together some stuff I thought was unrelated. So I guess my major question is are these things really unrelated? Are they due to a software issue related to normal kernel behavior running on an AMD CPU? Or is it a physical memory issue? I did run memtest86+ three times last night and got zero errors. More from opencode:

3. Crash history on this hardware — and how I know

You asked directly whether the crashwhy directory was how I knew about crashes on the other distro. Yes — that was the source, with one important precision: The notes in ~/Projects/crashwhy/ (AMDGPUChanges.md, kdump-setup-instructions.txt) are from a previous Arch Linux install on this laptop: they reference mkinitcpio, pacman hooks, systemd-boot UKIs, a btrfs/zstd root, and linux-lts 6.18 as a kdump capture kernel. They were restored from your backup today — the session journal shows /Documents/Projects being moved in from a mounted USB drive earlier today.

  • What those notes document are two other crash classes, not today’s:
  1. A btrfs/zstd compressor oops (ZSTD_compressBlock_doubleFast NULL deref) on kernel 7.0.14 that went un-captured because no kdump backend existed then;
  2. amdgpu DCN 2.1 modeset regressions in the kernel 7.1 series (your notes prescribed escaping to the LTS kernel, amdgpu gpu_recovery=1, and deep S3 suspend).
  • So this is not “the same issue” across two distros. Today’s crash is a third distinct failure mode. What actually spans the two installs is the same physical machine, and a pattern of kernel-space memory-corruption-style crashes (NULL derefs / wild control flow) on kernel 7.0.x-era builds.

Here is what opencode is suggesting to do to mitigate future crashes. Do these seem reasonable?

4. Recommendations (priority order)

  1. **Run memtest86+ overnight

I did that; three memtest86+ passes over all the RAM revealed zero errors.

  1. Make panics self-recovering so you never have to hard power-cycle again:
sudo tee /etc/sysctl.d/99-panic-reboot.conf <<'EOF'
kernel.panic = 10
EOF
  1. Keep an eye on the GPU side too. Your Arch notes flagged amdgpu DCN 2.1 modeset regressions in the 7.1 series, and cosmic-comp logged recurring Failed to render texture ... import for wrong devices errors both before the crash and on the current boot. If you see hangs/freezes (as opposed to idle-time panics), apply the fix from your own notes — on Pop it’s:
sudo tee /etc/modprobe.d/amdgpu.conf <<'EOF'
options amdgpu gpu_recovery=1
EOF
  1. Optional, only if crashes continue: port the kdump setup you built on Arch (kdump-tools + crashkernel reservation) for full vmcore analysis. For now, efi_pstore proved itself — it captured everything needed to diagnose this crash.
  2. If memtest passes and it recurs on 7.1.5, this is reportable to Pop!_OS (support.system76.com) or upstream kernel bugzilla. The signature to cite:
    NULL supervisor write, RIP srso_safe_ret+0x1a/0x20, CPU 0 idle task
    (swapper/0), AMD Zen 2 (Ryzen 5700U), kernel 7.0.11, BIOS 1.07.27_S3,
    microcode 0x8608108
    — and attach the archived pstore records (noting the OE taint
    from Pop’s system76 DKMS modules).

Any suggestions or advice would be greatly appreciated.