I have been running a pangolin 11 laptop for a few months now. I enjoyed playing around with Arch Linux but decided I was crashing my computer too often with interrupted pacman -Syu calls.
Two days ago I switched to PopOs as that was what originally shipped with the laptop and is the OS System76 suggests for their users. After switching and starting to work on a project using opencode (with glm-5.3 as the model via Ollama cloud), I got a kernel panic and had to reboot. Here is what happened according to opencode after rebooting:
1. What happened (reconstructed timeline)
Time (PDT) Event Source 10:43:31 Boot into kernel 7.0.11 journal, boot -1 10:50:39 apt upgradeinstalls kernel 7.1.5 (not yet booted)apt history 12:40:47 Last user command: chown -R user:user ~/Projects(completed fine)journal 12:40:54 Last journal entry — no shutdown sequence follows journal 12:41:16 Kernel oops: NULL pointer write, CPU 0, in the idle task pstore (t=7074.703) 12:41:18 Kernel panic - not syncing: Attempted to kill the idle task!pstore (t=7076.994) ~12:41–12:42 Machine sat on the panic screen until power-cycled — 12:42:09 Reboot into the newly installed 7.1.5 journal, boot 0
The journal simply stops mid-session with no shutdown records — a hard, uncontrolled stop. The actual crash text survived because the EFI persistent store (
efi_pstore) captured the kernel log buffer at oops/panic time, andsystemd-pstore.servicearchived it to/var/lib/systemd/pstore/on the next boot.
2. The actual crash
The pstore archive contains two records: the oops (“Oops#1”, 18 chunks) and the resulting panic (“Panic#2”, 18 chunks — the same buffer plus the panic lines). Reassembled, the essential part:
BUG: kernel NULL pointer dereference, address: 0000000000000000 #PF: supervisor write access in kernel mode #PF: error_code(0x0002) - not-present page Oops: Oops: 0002 [#1] SMP NOPTI CPU: 0 UID: 0 PID: 0 Comm: swapper/0 Tainted: G OE 7.0.11-76070011-generic PREEMPT(lazy) Hardware name: System76 Pangolin/Pangolin, BIOS 1.07.27_S3 10/22/2024 RIP: 0010:srso_safe_ret+0x1a/0x20 Code: ... 00 <00> 00 0f 1f 00 e8 db ff ff ff 0f 0b ... RAX: 0000000000000000 RBX: ffff8b31803ed400 ... CR2: 0000000000000000 Call Trace: ? tick_nohz_stop_tick+0x63/0x260 ? raw_spin_rq_unlock+0x14/0x50 ? srso_return_thunk+0x5/0x5f cpuidle_enter+0x2e/0x50 call_cpuidle+0x21/0x50 cpuidle_idle_call+0x173/0x200 do_idle+0x90/0xf0 ... ---[ end trace ]--- Kernel panic - not syncing: Attempted to kill the idle task!Interpretation. The faulting instruction is two zero bytes (
00 00=add %al,(%rax)) with RAX = 0 → a write to address 0. Those bytes are alignment padding inside the SRSO “safe return” thunk — part of the AMD Speculative Return Stack Overflow mitigation (CVE-2023-20569, “Inception”). No legitimate execution path ever runs there. This is a “wild control-flow transfer”: CPU 0 was entering the idle loop, aRETwent somewhere it should never go, and execution landed in mitigation thunk padding. Had RAX not happened to be 0, this would have been silent memory corruption instead of a clean fault.
The panic 2.3 s later is the standard kernel response to an oops in the idle task: the oops handler tries to kill the faulting task, discovers it is PID 0 (swapper/0), which can never be killed, and panics with exactly the message above.
Okay, interesting. What is causing this to happen? I had a log of crashes from my previous Arch set-up on the same machine. Opencode found those logs in a directory and put together some stuff I thought was unrelated. So I guess my major question is are these things really unrelated? Are they due to a software issue related to normal kernel behavior running on an AMD CPU? Or is it a physical memory issue? I did run memtest86+ three times last night and got zero errors. More from opencode:
3. Crash history on this hardware — and how I know
You asked directly whether the
crashwhydirectory was how I knew about crashes on the other distro. Yes — that was the source, with one important precision: The notes in~/Projects/crashwhy/(AMDGPUChanges.md,kdump-setup-instructions.txt) are from a previous Arch Linux install on this laptop: they reference mkinitcpio, pacman hooks, systemd-boot UKIs, a btrfs/zstd root, and linux-lts 6.18 as a kdump capture kernel. They were restored from your backup today — the session journal shows/Documents/Projectsbeing moved in from a mounted USB drive earlier today.
- What those notes document are two other crash classes, not today’s:
- A btrfs/zstd compressor oops (
ZSTD_compressBlock_doubleFastNULL deref) on kernel 7.0.14 that went un-captured because no kdump backend existed then;- amdgpu DCN 2.1 modeset regressions in the kernel 7.1 series (your notes prescribed escaping to the LTS kernel,
amdgpu gpu_recovery=1, and deep S3 suspend).
- So this is not “the same issue” across two distros. Today’s crash is a third distinct failure mode. What actually spans the two installs is the same physical machine, and a pattern of kernel-space memory-corruption-style crashes (NULL derefs / wild control flow) on kernel 7.0.x-era builds.
Here is what opencode is suggesting to do to mitigate future crashes. Do these seem reasonable?
4. Recommendations (priority order)
- **Run memtest86+ overnight
I did that; three memtest86+ passes over all the RAM revealed zero errors.
- Make panics self-recovering so you never have to hard power-cycle again:
sudo tee /etc/sysctl.d/99-panic-reboot.conf <<'EOF' kernel.panic = 10 EOF
- Keep an eye on the GPU side too. Your Arch notes flagged amdgpu DCN 2.1 modeset regressions in the 7.1 series, and
cosmic-complogged recurringFailed to render texture ... import for wrong deviceserrors both before the crash and on the current boot. If you see hangs/freezes (as opposed to idle-time panics), apply the fix from your own notes — on Pop it’s:sudo tee /etc/modprobe.d/amdgpu.conf <<'EOF' options amdgpu gpu_recovery=1 EOF
- Optional, only if crashes continue: port the kdump setup you built on Arch (
kdump-tools+ crashkernel reservation) for full vmcore analysis. For now,efi_pstoreproved itself — it captured everything needed to diagnose this crash.- If memtest passes and it recurs on 7.1.5, this is reportable to Pop!_OS (support.system76.com) or upstream kernel bugzilla. The signature to cite:
NULL supervisor write, RIPsrso_safe_ret+0x1a/0x20, CPU 0 idle task
(swapper/0), AMD Zen 2 (Ryzen 5700U), kernel 7.0.11, BIOS 1.07.27_S3,
microcode 0x8608108 — and attach the archived pstore records (noting theOEtaint
from Pop’s system76 DKMS modules).
Any suggestions or advice would be greatly appreciated.