Conversation
| // overriding --threads. dt_alloc_perthread() pools are sized by | ||
| // dt_get_num_threads(), so a larger team would write past their end. | ||
| omp_set_num_threads(darktable.num_openmp_threads); | ||
| #endif |
There was a problem hiding this comment.
And what about the opposite, if GraphicsMagick do set a threads count below darktable.num_openmp_threads?
There was a problem hiding this comment.
Sorry, had to run (literally).
Deslopping (manually, so I learn something, probably nothing new to anyone else): The original issue was that -t 4 resulted in allocating buffers for 4 threads, then GM reset OpenMP to use 12 threads -> crash.
What was tested now:
OMP_NUM_THREADS sets the default number of threads
-t n overrides that to n (darktable.num_openmp_threads) -- and this number is also used for the allocation
OMP_THREAD_LIMIT m sets a hard maximum.
With OMP_NUM_THREADS = 1 and -t 4, initial number of threads is 1, overridden to 4, buffers are allocated according to that: 4 threads, buffers for 4 threads, OK.
With OMP_THREAD_LIMIT = 1 and -t 4, the hard limit is 1; buffers are still allocated according to -t 4, buffers are allocated for 4 threads, 1 thread is used, wasteful, but no crash.
--- original LLM answer below ---
(Codex)
The same call covers that case. It restores darktable's configured count whether GraphicsMagick raised or lowered it.
By reading the code, Codex checked GraphicsMagick 1.3.46's initialization and darktable's per-thread allocator. Fewer workers cannot overrun buffers sized for the configured count. The reset is unconditional.
A runtime probe with 4 requested threads confirmed that GraphicsMagick selected 1 with OMP_NUM_THREADS=1, then the reset restored 4. With OMP_THREAD_LIMIT=2, the actual team stayed at 2. The build passed, and CPU exports using 0068-rawdenoise-xtrans matched the control pixel for pixel in both cases.
There was a problem hiding this comment.
I don't understand this, if GM asked for 2 threads and we set back to 16 we should have the same issue and then potentially crash in GM this time, no?
Or do you mean that Codex has analyzed all the GM sources to ensure that this is not an issue for GM? That seems very fragile anyway as we depends on GM source code changes.
There was a problem hiding this comment.
It ran the integration tests with those environment variables and thread counts (and it also read the sources). I'll continue tomorrow, it's late.
1e89645 to
d060937
Compare
|
Claude Code says GM can't overrun its buffers when darktable raises the count. In the future, G'MIC 3.x and 4.x. could cause problems -- see end of report. Also, 3 new (unrelated) bugs were found in XTrans-related code. I'll file those separately. GM's init (magick/resource.c:676-689) does two things with threads. It calls omp_set_num_threads() on the calling thread, using OMP_NUM_THREADS or the CPU count, and it stores that number as its ThreadsResource limit. It allocates nothing Every per-thread array in GM is sized at the moment it is allocated, from the calling thread's omp_get_max_threads():
In Pascal's example (GM picks 2, darktable restores 16), the restore runs inside dt_init(), before any GM image exists. Every image GM creates afterwards is sized for 16, and its parallel loops run on the same thread with 16 threads. The stored "2" is only read in two places:
GM does have one real constraint: an image must not be processed by more threads than existed when it was created. darktable creates, reads and destroys each GM image inside one function, with no count change in between. On "fragile" The earlier reply said "Codex read the GM sources", which invites the "we depend on GM internals" objection. The stronger argument doesn't need a GM audit:
Separate problem found while checking (not reproduced) G'MIC 3.x and 4.x call omp_set_num_threads(nb_cpus) every time a gmic instance is constructed, unless OMP_NUM_THREADS is set (gmic-v.4.0.5/src/gmic.cpp:3603, reached from :4183). darktable constructs one in lut3d for compressed LUTs (src/iop/lut3dgmic.cpp:48, :105). With -t below the CPU count, that would reset the count on the calling thread (a pipe worker, or darktable-cli's main thread) in the middle of a session. The reset in dt_init() can't cover that. The installed G'MIC is 2.9.4, which doesn't do this, so I couldn't reproduce it. This is the genuinely fragile part, and adding num_threads(dt_get_num_threads()) to DT_OMP_FOR would fix it. It belongs in its own issue, not this PR. |
|
Also: once #22456 and #22457 are fixed, a next fix can be added to allocate buffers based on the thread count, until that changing the allocation would change output. Over to Claude:
|
|
@kofa73 : So if I understand correctly this works with current GM but will fail with GM > 3.x. That's not a nice future :) What would be your proposal on this? If we still want to merge this PR I would like to get more comments about the future GM 3.x and above issue that we are going to hit. |
|
According to Claude Code's analysis:
The agent's proposal was above: adding num_threads(dt_get_num_threads()) to DT_OMP_FOR would fix it; a more detailed version:
|
GraphicsMagick resets the main thread's OpenMP limit to the CPU count. With --threads set lower, darktable-cli can then run more threads than its per-thread buffers hold. Reapply the requested limit after Magick initialization. Fixes darktable-org#22436
OpenMP never runs more threads than OMP_THREAD_LIMIT, but darktable sized its per-thread buffers for the requested count (--threads, or the CPU count), so the slots above the limit were allocated and never used. Lower darktable.num_openmp_threads to the limit before the first omp_set_num_threads(). OMP_THREAD_LIMIT=N now acts like --threads N wherever darktable uses its thread count: buffer sizes, the number of work slices in algorithms such as the bilateral grid, the number of background worker threads, and whether thumbnails are generated in the background. Exports match those made with --threads N, which can differ from the previous output by rounding.
d060937 to
6778d31
Compare
|
I rebased the branch onto current master. Claude Code added a new commit to cap the thread count and thus allocate memory according to the effective number of threads. We now have 3 commits: the original crash fix, the cap on the thread count, and the release note. Over to the bot: The new commit, With the cap, The cap needs #22475. Before that change, X-Trans raw denoise split its work by thread count, so the cap changed its output. The cap has no separate release note (the issue is about the bug fix in the first commit). Testing (Linux, 12 CPUs, Release build, CPU path):
The cap change has been reviewed by Codex CLI and Antigravity. |
|
@TurboGit : shall I create an issue for the GM >= 3 question? (#22438 (comment)) |
GraphicsMagick resets the main thread's OpenMP limit to the CPU count. With --threads set lower, darktable-cli can then run more threads than its per-thread buffers hold. Reapply the requested limit after Magick initialization.
Summary
darktable-cli -t Nwith N below the CPU count crashed with heap corruption.Fix
Call
omp_set_num_threads(darktable.num_openmp_threads)again after theGraphicsMagick / ImageMagick init.
Referenced issue
Fixes #22436
Checklist
src/tests/integration/where the pixelpipe is touched, ordarktable-clias a headless smoke test._(), new preferences are registered indata/darktableconfig.xml.in.RELEASE_NOTES.mdentry was added (only needed if fixing an issue in a release). Do not reference GitHub issues.Test instructions
Tested locally: Linux, 12 CPUs, Release, GraphicsMagick 1.3.46,
--disable-opencl.xtransIV.rafwith-t 4: aborted before, exports now.-t 1,-t 8andno
-talso succeed.0068-rawdenoise-xtrans,0100-invert-xtransand
0145-lens-metadata-xtransIV-modversion-6, run with-t 4and withoutOMP_THREAD_LIMIT, crashed before and finish now.AI assistance
Fixed by Claude Code, reviewed by Codex CLI