From mboxrd@z Thu Jan 1 00:00:00 1970 From: AL-KERNEL To: kernel-cve@kernelcve.org Subject: [CVE-2026-68237][MODERATE REGULAR] drm/amdgpu/userq: fix indefinite fence wait during GPU reset Date: Mon, 10 Aug 2026 13:45:10 -0400 Message-ID: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-AL-KERNEL-CVE: CVE-2026-68237 X-AL-KERNEL-Priority: MODERATE REGULAR X-AL-KERNEL-Severity: MODERATE REGULAR X-AL-KERNEL-Base-Severity: MODERATE X-AL-KERNEL-KPANIC: NO X-AL-KERNEL-ActionableScore: 4 X-AL-KERNEL-ActionableScore-Lower: 3 X-AL-KERNEL-Commit: 3085ae8695e025b39d208f288c6265edc75abbe8 List-Id: CVE: CVE-2026-68237 Priority: MODERATE REGULAR AL-KERNEL base severity: MODERATE KPANIC flag: NO Patch: drm/amdgpu/userq: fix indefinite fence wait during GPU reset Commit: 3085ae8695e025b39d208f288c6265edc75abbe8 Upstream patch: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id=3085ae8695e025b39d208f288c6265edc75abbe8 Original CVE announcement: https://lore.kernel.org/linux-cve-announce/?q=CVE-2026-68237 Analysis date: Mon, 10 Aug 2026 13:45:10 -0400 ActionableScore: 4 ActionableScore lower bound: 3 Actionable bucket: Borderline / manual review recommended, not an Important candidate Manual review required: YES Summary: AMDGPU user queues can retain an unsignaled fence across GPU reset when the queue is not mapped, allowing a local user with GPU render access to wedge eviction, suspend, or teardown paths and cause a denial of service. ====================================================================== ABOUT THIS REPORT ====================================================================== The original Linux kernel CVE announcement for CVE-2026-68237 is available here: https://lore.kernel.org/linux-cve-announce/?q=CVE-2026-68237 The original announcement does not normally provide a security severity estimate, CVSS assessment, or enough information to determine whether the reported kernel bug represents a practically relevant security issue. This report was generated by AL-KERNEL, an AI-assisted Linux kernel vulnerability analysis system developed by Alexander Larkin. It combines an autonomous classifier with LLM-assisted technical analysis and a separate ActionableScore mechanism. The purpose of this report is to prioritize Linux kernel CVEs before manual review, identify cases that require prompt investigation, and support automatic closure of issues that are unlikely to have meaningful security impact. Published priority for this report: MODERATE REGULAR Manual review required: YES A detailed explanation of the methodology and priority rules is included at the end of this message. ====================================================================== AL-KERNEL CLASSIFICATION RESULT ====================================================================== CVE-2026-68237 MODERATE CHECK WITH IMPACT FROM ORIG NN LOW Maybe valid. Check manually. Hints by AL-KERNEL: The best (paranoid) CVSS is 'AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H';*CWE-833;CWE-667;**CWE-400;Other CVSS 'AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:H';BEST CVSS score: '5.5';DESCR 'AMDGPU user queues can leave a pending fence unsignaled across a GPU reset when the queue is not in the mapped state, for example during eviction. The eviction or suspend worker and process teardown can then wait forever on dma_fence_wait_timeout, which can wedge the machine and cause a local denial of service. For the CVSS the PR:L is selected because a local user or process with access to AMDGPU DRM or render interfaces may be able to create user queues and trigger GPU reset or eviction behavior. The issue is not network reachable and requires local interaction with the GPU driver rather than remote packet input. Impact is availability loss only based on the patch evidence, with no supported UAF, OOB access, information leak, or privilege escalation primitive.';YES REQUIRES MANUAL CHECK; ,and ActionableScore result is Borderline / manual review recommended, not an Important candidate (with actual score 4) SKIP CVE-2026-68237 UNKNOWN SKIP No affected files built, so skip this CVE NO - - unknown MAYBE SIMPLEFIX HARDWARE MEMORY DECREASED_TO_MODERATE_REGULAR_FROM_MODERATE7_BASED_ON_GUESSCVSS_AND_CIANNH - - checked ====================================================================== ACTIONABLESCORE ANALYSIS ====================================================================== ActionableScore=4 ActionableScoreLower=3 ## 1. ActionableScore * Conservative score: 3 * Paranoid score: 4 * Final recommended bucket: **Borderline / manual review recommended, not an Important candidate**:: ## 2. Signal breakdown Triggered signals: * Local unprivileged trigger: +1 AMDGPU DRM render nodes are commonly accessible to local unprivileged users, and user queues are plausibly reachable from local GPU workloads. * Reliable kernel crash / strong DoS: +1 The patch and trace show an indefinite fence wait in eviction or suspend worker paths after GPU reset, causing blocked kernel workers and possible machine wedge. * Availability impact realistic: +1 The impact is not merely a transient warning. The machine can become wedged due to forever-pending fences and blocked teardown or suspend paths. * Broad/common local subsystem exposure: +1 in paranoid score only AMDGPU is common on desktops, workstations, and GPU compute systems, and render-node access is often delegated to normal users. Subtracted signals: * Hard or timing-dependent condition: -1 in conservative score Reliable triggering appears to require a queue in a non-MAPPED state, such as mid-eviction, combined with a GPU reset. This is more specific than a simple direct ioctl crash. Not triggered: * No generic memory corruption signal The patch shows an unsignaled fence and indefinite wait, not UAF, double free, OOB access, type confusion, or attacker-controlled overwrite. * No privilege escalation plausible signal There is no supported primitive for LPE. The issue is availability-only based on the patch evidence. * No network reachability This is a local GPU driver path, not packet-triggered or remotely reachable. ## 3. Reachability analysis A local user or local process with access to AMDGPU DRM or render interfaces may be able to create user queues and interact with workloads that enter eviction or reset-related paths. In many desktop or GPU-compute deployments, render-node access is intentionally granted to non-root users, so PR:L is a reasonable practical interpretation. Namespaces and containers may matter if GPU device nodes are passed through into containers. In that case, a container user with access to the render node could plausibly trigger the affected path, but this is still local device access rather than remote reachability. The path is not globally default-exposed on all Linux systems because it requires AMDGPU hardware and user queue functionality. However, on affected AMDGPU systems, local GPU access is a realistic deployment scenario. Call-site confidence: high. The commit includes the affected function, the relevant wait path, and a concrete blocked-worker trace. ## 4. Severity interpretation This behaves like an actionable local DoS rather than an Important-class memory corruption issue. The realistic impact is a system-wide or GPU-subsystem wedge caused by an indefinitely pending fence after reset. Theoretical memory corruption potential is not supported by the patch. There is no evidence of stale object reuse, write-after-free, callback corruption, OOB write, refcount takeover, or privilege escalation. The main reason not to auto-close is practical local availability impact on systems where unprivileged users can access AMDGPU render nodes. ## 5. One-sentence report phrase AMDGPU user queues can retain an unsignaled fence across GPU reset when the queue is not mapped, allowing a local user with GPU render access to wedge eviction, suspend, or teardown paths and cause a denial of service. ## 6. Manual review recommendation MANUAL CHECK RECOMMENDED Manual review is recommended because the issue is a local GPU-driver DoS with potentially system-wide availability impact and realistic PR:L exposure through render nodes. It is not a strong Important candidate because the patch does not support memory corruption, information disclosure, integrity impact, or privilege escalation. ====================================================================== UPSTREAM PATCH SUMMARY ====================================================================== Patch: drm/amdgpu/userq: fix indefinite fence wait during GPU reset Commit: 3085ae8695e025b39d208f288c6265edc75abbe8 Upstream URL: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id=3085ae8695e025b39d208f288c6265edc75abbe8 Commit description: pre_reset only force-completes fences of MAPPED queues. A queue in any other state (e.g. mid-eviction) keeps its last_fence pending; after a GPU reset that fence never signals, so the eviction/suspend worker and process teardown (amdgpu_evf_mgr_flush_suspend) wait on it forever and wedge the machine: INFO: task kworker/6:28 blocked for more than 120 seconds. Workqueue: events amdgpu_eviction_fence_suspend_worker [amdgpu] Call Trace: dma_fence_wait_timeout+0x7e/0x130 amdgpu_userq_evict+0x67/0x140 [amdgpu] amdgpu_eviction_fence_suspend_worker+0xd8/0x160 [amdgpu] process_scheduled_works+0xa6/0x420 Force-complete every queue's fence regardless of state. The unmap and mark-hung step stays gated on MAPPED, since unmapping a queue that is not mapped is invalid. Fixes: 290f46c ("drm/amdgpu: Implement user queue reset functionality") Reviewed-by: Christian König Signed-off-by: Jesse Zhang Signed-off-by: Alex Deucher (cherry picked from commit 9102b39 ) Cc: stable@vger.kernel.org Signed-off-by: Greg Kroah-Hartman Changed files: drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c Diff excerpt: Not included in this email. See the upstream URL for the full patch. Full patch: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id=3085ae8695e025b39d208f288c6265edc75abbe8 ====================================================================== DETAILED REPORT METHODOLOGY ====================================================================== The original Linux kernel CVE announcement for CVE-2026-68237 can be found here: https://lore.kernel.org/linux-cve-announce/?q=CVE-2026-68237 The original CVE announcement normally does not include a security-level estimate. In particular, it may not contain a CVSS assessment, an impact level, or enough information to determine whether the reported bug is a practically relevant security issue. One purpose of this parallel CVE list is to provide that missing technical and prioritization information. The original goal of the AL-KERNEL project was to prioritize Linux kernel CVE analysis automatically before manual review. The system can also help identify non-security issues that may be suitable for automatic closure. This report was generated by AL-KERNEL, an AI-assisted Linux kernel vulnerability analysis system developed by Alexander Larkin. The first analysis stage combines an autonomous classifier with additional LLM-based analysis. The autonomous classifier runs locally on a CPU and is based on a backpropagation neural network. Together, these mechanisms produce a technical vulnerability description, identify likely weakness types, estimate CVSS severity, and provide input for ActionableScore. Two CVSS estimates are retained because incomplete kernel vulnerability information often permits more than one defensible interpretation: Conservative CVSS vector: AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:H The Best / paranoid CVSS vector: AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H The Best / paranoid CVSS score: 5.5 The conservative vector represents a lower-impact interpretation. The Best/paranoid vector intentionally represents a plausible upper-bound interpretation and should not automatically be treated as demonstrated real-world impact. CVSS may also need to be adjusted for a particular Linux deployment, because actual reachability, privileges, enabled kernel configuration, hardware, namespaces, exposed device nodes, and other environmental conditions can differ significantly between systems. A separate ActionableScore mechanism evaluates practical remediation urgency. Its analysis may include reachability, attack prerequisites, subsystem exposure, memory-corruption characteristics, denial-of-service reliability, and possible confidentiality, integrity, or privilege-escalation impact. Conservative ActionableScore: 3 Paranoid ActionableScore: 4 The final base severity is taken directly from the second tab-separated field of the AL-KERNEL classification result. ActionableScore does not replace or independently override that final AL-KERNEL decision, and if ActionableScore adjusted impact level of ALKERNEL, then you would see self-readable flags above like INCREASED_TO_HIGH_BASED_ON_ACTIONABLESCOREHIGHEREQTHAN7. For an AL-KERNEL result of MODERATE, this report uses the following additional presentation split: ActionableScore below 5 -> MODERATE REGULAR ActionableScore 5 or more -> MODERATE 7.0 The distinction between MODERATE REGULAR and MODERATE 7.0 makes it possible to identify Moderate issues that should receive manual analysis and fixes before lower-priority MODERATE REGULAR issues. In many cases, MODERATE REGULAR fixes may wait for a later rebase or routine update. There is one override in which MODERATE REGULAR becomes MODERATE 7.0 even when the ActionableScore is below 5. When the AL-KERNEL result contains the KPANIC flag, a MODERATE result is always presented as MODERATE 7.0. The KPANIC flag selected with few regexps without usage of AI at all, so it helps to detect cases when Kernel Crash happens and similar (to filter False-Negative results from the LLM usage). KPANIC indicates that a reliable kernel crash, kernel panic, or similarly serious kernel availability impact was identified by the classification workflow. AL-KERNEL base severity for this report: MODERATE KPANIC detected for this report: NO Published priority for this report (same as in Subject): MODERATE REGULAR These results are intended to support engineering triage. They are machine-generated estimates, and cases marked for manual review should be validated by a human security engineer before final disposition. For more info read docs linked from here: https://kernelcve.org/ (and you can submit you own patch there to generate such a report for non-existant CVE-id yet). Note that in many cases this AI tool selects higher severity, than real is (means you can expect Importants instead of Moderate 7.0 or Moderates 7.0 instead of regular Moderates). If you see such cases, please use reply email interface to add additional manual analyses info to this particular CVE. And please, please, let me know when you see Lows instead of Importants or Important instead of Low (because particular for such cases I need to tune this AI tool to make it better for this one and next similar). My contact email for such notifications is alexanjelausa@gmail.com (and both send reply to CVE record itself too and see "reply" button below for howto reply).