Bug 2252130
| Summary: | Davinci Resolve crash at startup with kernel 6.6.2 + AMDGPU | ||
|---|---|---|---|
| Product: | [Fedora] Fedora | Reporter: | Yannick Defais <sevmek> |
| Component: | kernel | Assignee: | Kernel Maintainer List <kernel-maint> |
| Status: | CLOSED EOL | QA Contact: | Fedora Extras Quality Assurance <extras-qa> |
| Severity: | medium | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 39 | CC: | acaringi, adscvr, airlied, alciregi, bskeggs, hdegoede, hpa, jarod, josef, kernel-maint, linville, masami256, mchehab, mj.oppliger, nixuser, ptalbert, steved |
| Target Milestone: | --- | Keywords: | Regression |
| Target Release: | --- | ||
| Hardware: | x86_64 | ||
| OS: | Linux | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | --- | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2024-11-27 22:11:25 UTC | Type: | --- |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Yannick Defais
2023-11-29 18:02:17 UTC
I experience something similar - I have an AMD RX 570 GPU with Fedora 39 + ROCm 5.7.1 + Kernel 6.6.3 and am running GPU compute projects through BOINC. With all 6.6.x kernels tested so far the computation "fails" --> it does not throw obvious errors in the application but it never finishes computing. Booting the same system (without config changes) with kernel 6.5.12 does work just fine. The following errors are logged with journalctl -k: Dez 02 16:09:18 kernel: amdgpu 0000:02:00.0: amdgpu: Disabling VM faults because of PRT request! Dez 02 17:01:53 kernel: amdgpu: Failed to reserve buffers in ttm. Dez 02 17:01:53 kernel: amdgpu: Failed to reserve buffers in ttm. Dez 02 17:01:53 kernel: amdgpu: Failed to reserve buffers in ttm. Dez 02 17:01:53 kernel: ------------[ cut here ]------------ Dez 02 17:01:53 kernel: WARNING: CPU: 0 PID: 10 at drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c:1518 amdgpu_amd> Dez 02 17:01:53 kernel: Modules linked in: uinput snd_seq_dummy snd_hrtimer nf_conntrack_netbios_ns nf_conntrack_br> Dez 02 17:01:53 kernel: drm_suballoc_helper amdxcp polyval_clmulni drm_buddy polyval_generic ghash_clmulni_intel g> Dez 02 17:01:53 kernel: CPU: 0 PID: 10 Comm: kworker/0:1 Not tainted 6.6.3-200.fc39.x86_64 #1 Dez 02 17:01:53 kernel: Hardware name: Hewlett-Packard HP Z440 Workstation/212B, BIOS M60 v02.61 03/23/2023 Dez 02 17:01:53 kernel: Workqueue: events delayed_fput Dez 02 17:01:53 kernel: RIP: 0010:amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 02 17:01:53 kernel: Code: df 5b 5d 41 5c e9 7a ad cc db 5b 5d 41 5c c3 cc cc cc cc e8 fc 5b 46 dc eb cc be 03 0> Dez 02 17:01:53 kernel: RSP: 0000:ffffc900000b7cc0 EFLAGS: 00010206 Dez 02 17:01:53 kernel: RAX: ffff88818b509020 RBX: ffff88818b509000 RCX: ffff88818b509000 Dez 02 17:01:53 kernel: RDX: ffff88832e9a7d48 RSI: ffff888269f1b730 RDI: ffff88818b509040 Dez 02 17:01:53 kernel: RBP: ffff888269f1b000 R08: 0000000000000000 R09: 0000000080200010 Dez 02 17:01:53 kernel: R10: 0000000000000000 R11: 0000000000000000 R12: ffff88818b509040 Dez 02 17:01:53 kernel: R13: ffff88829ff07a00 R14: 0000000000000000 R15: ffff88862f000001 Dez 02 17:01:53 kernel: FS: 0000000000000000(0000) GS:ffff888fefa00000(0000) knlGS:0000000000000000 Dez 02 17:01:53 kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 Dez 02 17:01:53 kernel: CR2: 0000559837911784 CR3: 0000000555518002 CR4: 00000000003706f0 Dez 02 17:01:53 kernel: DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 Dez 02 17:01:53 kernel: DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400 Dez 02 17:01:53 kernel: Call Trace: Dez 02 17:01:53 kernel: <TASK> Dez 02 17:01:53 kernel: ? amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 02 17:01:53 kernel: ? __warn+0x81/0x130 Dez 02 17:01:53 kernel: ? amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 02 17:01:53 kernel: ? report_bug+0x171/0x1a0 Dez 02 17:01:53 kernel: ? handle_bug+0x3c/0x80 Dez 02 17:01:53 kernel: ? exc_invalid_op+0x17/0x70 Dez 02 17:01:53 kernel: ? asm_exc_invalid_op+0x1a/0x20 Dez 02 17:01:53 kernel: ? amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 02 17:01:53 kernel: amdgpu_vm_fini+0x49/0x550 [amdgpu] Dez 02 17:01:53 kernel: amdgpu_driver_postclose_kms+0x191/0x280 [amdgpu] Dez 02 17:01:53 kernel: drm_file_free+0x21c/0x270 Dez 02 17:01:53 kernel: drm_release+0x74/0xf0 Dez 02 17:01:53 kernel: __fput+0xf5/0x290 Dez 02 17:01:53 kernel: delayed_fput+0x23/0x30 Dez 02 17:01:53 kernel: process_one_work+0x174/0x340 Dez 02 17:01:53 kernel: worker_thread+0x27b/0x3a0 Dez 02 17:01:53 kernel: ? __pfx_worker_thread+0x10/0x10 Dez 02 17:01:53 kernel: kthread+0xe8/0x120 Dez 02 17:01:53 kernel: ? __pfx_kthread+0x10/0x10 Dez 02 17:01:53 kernel: ret_from_fork+0x34/0x50 Dez 02 17:01:53 kernel: ? __pfx_kthread+0x10/0x10 Dez 02 17:01:53 kernel: ret_from_fork_asm+0x1b/0x30 Dez 02 17:01:53 kernel: </TASK> Dez 02 17:01:53 kernel: ---[ end trace 0000000000000000 ]--- Dez 02 17:01:53 kernel: amdgpu 0000:02:00.0: amdgpu: still active bo inside vm I'll try to get some more data on this and report back here. This seems to be fixed upstream in the upcoming kernel 6.7 series - so we'll just have to wait for it to be backported. See https://gitlab.freedesktop.org/drm/amd/-/issues/3007#note_2199326 as well as https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/log/?h=v6.7-rc4&qt=grep&q=amdkfd Still broken with 6.6.4 and 6.6.6. OpenCL is now working for me since kernel 6.6.7 - although the traces are still logged: Dez 19 05:30:31 kernel: amdgpu: Failed to reserve buffers in ttm. Dez 19 05:30:31 kernel: amdgpu: Failed to reserve buffers in ttm. Dez 19 05:30:31 kernel: amdgpu: Failed to reserve buffers in ttm. Dez 19 05:30:31 kernel: amdgpu: Failed to reserve buffers in ttm. Dez 19 05:30:31 kernel: ------------[ cut here ]------------ Dez 19 05:30:31 kernel: WARNING: CPU: 3 PID: 11999 at drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c:1518 amdgpu_a> Dez 19 05:30:31 kernel: Modules linked in: uinput snd_seq_dummy snd_hrtimer nf_conntrack_netbios_ns nf_conntrack_bro> Dez 19 05:30:31 kernel: polyval_clmulni drm_suballoc_helper amdxcp polyval_generic drm_buddy ghash_clmulni_intel nv> Dez 19 05:30:31 kernel: CPU: 3 PID: 11999 Comm: kworker/3:2 Tainted: G W 6.6.7-200.fc39.x86_64 #1 Dez 19 05:30:31 kernel: Hardware name: Hewlett-Packard HP Z440 Workstation/212B, BIOS M60 v02.61 03/23/2023 Dez 19 05:30:31 kernel: Workqueue: events delayed_fput Dez 19 05:30:31 kernel: RIP: 0010:amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 19 05:30:31 kernel: Code: df 5b 5d 41 5c e9 ba 7b cd cf 5b 5d 41 5c c3 cc cc cc cc e8 7c 30 47 d0 eb cc be 03 00> Dez 19 05:30:31 kernel: RSP: 0018:ffffc90009043cc0 EFLAGS: 00010287 Dez 19 05:30:31 kernel: RAX: ffff8886accbd020 RBX: ffff8886accbd000 RCX: ffff8886accbd000 Dez 19 05:30:31 kernel: RDX: ffff8883456cce48 RSI: ffff8885bf5a9730 RDI: ffff8886accbd040 Dez 19 05:30:31 kernel: RBP: ffff8885bf5a9000 R08: 0000000000000000 R09: 0000000080200013 Dez 19 05:30:31 kernel: R10: 0000000000000000 R11: 0000000000000000 R12: ffff8886accbd040 Dez 19 05:30:31 kernel: R13: ffff8881fce8f400 R14: 0000000000000000 R15: ffff8886a4000001 Dez 19 05:30:31 kernel: FS: 0000000000000000(0000) GS:ffff888fefac0000(0000) knlGS:0000000000000000 Dez 19 05:30:31 kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 Dez 19 05:30:31 kernel: CR2: 00007fa7eb423000 CR3: 0000000105184005 CR4: 00000000003706e0 Dez 19 05:30:31 kernel: DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 Dez 19 05:30:31 kernel: DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400 Dez 19 05:30:31 kernel: Call Trace: Dez 19 05:30:31 kernel: <TASK> Dez 19 05:30:31 kernel: ? amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 19 05:30:31 kernel: ? __warn+0x81/0x130 Dez 19 05:30:31 kernel: ? amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 19 05:30:31 kernel: ? report_bug+0x171/0x1a0 Dez 19 05:30:31 kernel: ? handle_bug+0x3c/0x80 Dez 19 05:30:31 kernel: ? exc_invalid_op+0x17/0x70 Dez 19 05:30:31 kernel: ? asm_exc_invalid_op+0x1a/0x20 Dez 19 05:30:31 kernel: ? amdgpu_amdkfd_gpuvm_destroy_cb+0x116/0x120 [amdgpu] Dez 19 05:30:31 kernel: amdgpu_vm_fini+0x49/0x550 [amdgpu] Dez 19 05:30:31 kernel: amdgpu_driver_postclose_kms+0x191/0x280 [amdgpu] Dez 19 05:30:31 kernel: drm_file_free+0x21c/0x270 Dez 19 05:30:31 kernel: drm_release+0x74/0xf0 Dez 19 05:30:31 kernel: __fput+0xf5/0x290 Dez 19 05:30:31 kernel: delayed_fput+0x23/0x30 Dez 19 05:30:31 kernel: process_one_work+0x174/0x340 Dez 19 05:30:31 kernel: worker_thread+0x27b/0x3a0 Dez 19 05:30:31 kernel: ? __pfx_worker_thread+0x10/0x10 Dez 19 05:30:31 kernel: kthread+0xe8/0x120 Dez 19 05:30:31 kernel: ? __pfx_kthread+0x10/0x10 Dez 19 05:30:31 kernel: ret_from_fork+0x34/0x50 Dez 19 05:30:31 kernel: ? __pfx_kthread+0x10/0x10 Dez 19 05:30:31 kernel: ret_from_fork_asm+0x1b/0x30 Dez 19 05:30:31 kernel: </TASK> Dez 19 05:30:31 kernel: ---[ end trace 0000000000000000 ]--- Dez 19 05:30:31 kernel: amdgpu 0000:02:00.0: amdgpu: still active bo inside vm This message is a reminder that Fedora Linux 39 is nearing its end of life. Fedora will stop maintaining and issuing updates for Fedora Linux 39 on 2024-11-26. It is Fedora's policy to close all bug reports from releases that are no longer maintained. At that time this bug will be closed as EOL if it remains open with a 'version' of '39'. Package Maintainer: If you wish for this bug to remain open because you plan to fix it in a currently maintained version, change the 'version' to a later Fedora Linux version. Note that the version field may be hidden. Click the "Show advanced fields" button if you do not see it. Thank you for reporting this issue and we are sorry that we were not able to fix it before Fedora Linux 39 is end of life. If you would still like to see this bug fixed and are able to reproduce it against a later version of Fedora Linux, you are encouraged to change the 'version' to a later version prior to this bug being closed. Fedora Linux 39 entered end-of-life (EOL) status on 2024-11-26. Fedora Linux 39 is no longer maintained, which means that it will not receive any further security or bug fix updates. As a result we are closing this bug. If you can reproduce this bug against a currently maintained version of Fedora Linux please feel free to reopen this bug against that version. Note that the version field may be hidden. Click the "Show advanced fields" button if you do not see the version field. If you are unable to reopen this bug, please file a new report against an active release. Thank you for reporting this bug and we are sorry it could not be fixed. |