Skip to content

common: restore OpenMP thread limit after GraphicsMagick init - #22438

Open
kofa73 wants to merge 3 commits into
darktable-org:masterfrom
kofa73:22436-threads-option-undone-by-graphicsmagick-overflows-perthread-buffers
Open

kofa73 wants to merge 3 commits into
darktable-org:masterfrom
kofa73:22436-threads-option-undone-by-graphicsmagick-overflows-perthread-buffers

Conversation

@kofa73

@kofa73 kofa73 commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

GraphicsMagick resets the main thread's OpenMP limit to the CPU count. With --threads set lower, darktable-cli can then run more threads than its per-thread buffers hold. Reapply the requested limit after Magick initialization.

Summary

darktable-cli -t N with N below the CPU count crashed with heap corruption.

Fix

Call omp_set_num_threads(darktable.num_openmp_threads) again after the
GraphicsMagick / ImageMagick init.

Referenced issue

Fixes #22436

Checklist

  • I have read CONTRIBUTING.md and the coding style.
  • I have not merged master into the topic branch.
  • The pull request is one logical change, and every commit compiles on its own.
  • I ran the relevant tests: unit tests, src/tests/integration/ where the pixelpipe is touched, or darktable-cli as a headless smoke test.
  • [N/A] New user-visible strings use _(), new preferences are registered in data/darktableconfig.xml.in.
  • A RELEASE_NOTES.md entry was added (only needed if fixing an issue in a release). Do not reference GitHub issues.

Test instructions

Tested locally: Linux, 12 CPUs, Release, GraphicsMagick 1.3.46, --disable-opencl.

  • xtransIV.raf with -t 4: aborted before, exports now. -t 1, -t 8 and
    no -t also succeed.
  • The darktable-cli commands of 0068-rawdenoise-xtrans, 0100-invert-xtrans
    and 0145-lens-metadata-xtransIV-modversion-6, run with -t 4 and without
    OMP_THREAD_LIMIT, crashed before and finish now.

AI assistance

Fixed by Claude Code, reviewed by Codex CLI

Comment thread src/common/darktable.c
// overriding --threads. dt_alloc_perthread() pools are sized by
// dt_get_num_threads(), so a larger team would write past their end.
omp_set_num_threads(darktable.num_openmp_threads);
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And what about the opposite, if GraphicsMagick do set a threads count below darktable.num_openmp_threads?

@kofa73 kofa73 Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, had to run (literally).

Deslopping (manually, so I learn something, probably nothing new to anyone else): The original issue was that -t 4 resulted in allocating buffers for 4 threads, then GM reset OpenMP to use 12 threads -> crash.
What was tested now:
OMP_NUM_THREADS sets the default number of threads
-t n overrides that to n (darktable.num_openmp_threads) -- and this number is also used for the allocation
OMP_THREAD_LIMIT m sets a hard maximum.

With OMP_NUM_THREADS = 1 and -t 4, initial number of threads is 1, overridden to 4, buffers are allocated according to that: 4 threads, buffers for 4 threads, OK.
With OMP_THREAD_LIMIT = 1 and -t 4, the hard limit is 1; buffers are still allocated according to -t 4, buffers are allocated for 4 threads, 1 thread is used, wasteful, but no crash.

--- original LLM answer below ---

(Codex)

The same call covers that case. It restores darktable's configured count whether GraphicsMagick raised or lowered it.

By reading the code, Codex checked GraphicsMagick 1.3.46's initialization and darktable's per-thread allocator. Fewer workers cannot overrun buffers sized for the configured count. The reset is unconditional.

A runtime probe with 4 requested threads confirmed that GraphicsMagick selected 1 with OMP_NUM_THREADS=1, then the reset restored 4. With OMP_THREAD_LIMIT=2, the actual team stayed at 2. The build passed, and CPU exports using 0068-rawdenoise-xtrans matched the control pixel for pixel in both cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't understand this, if GM asked for 2 threads and we set back to 16 we should have the same issue and then potentially crash in GM this time, no?

Or do you mean that Codex has analyzed all the GM sources to ensure that this is not an issue for GM? That seems very fragile anyway as we depends on GM source code changes.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It ran the integration tests with those environment variables and thread counts (and it also read the sources). I'll continue tomorrow, it's late.

@kofa73
kofa73 force-pushed the 22436-threads-option-undone-by-graphicsmagick-overflows-perthread-buffers branch from 1e89645 to d060937 Compare September 30, 2026 16:44
@kofa73 kofa73 self-assigned this Sep 30, 2026
@kofa73
kofa73 marked this pull request as draft September 30, 2026 18:54
@kofa73

kofa73 commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Claude Code says GM can't overrun its buffers when darktable raises the count. In the future, G'MIC 3.x and 4.x. could cause problems -- see end of report.

Also, 3 new (unrelated) bugs were found in XTrans-related code. I'll file those separately.


GM's init (magick/resource.c:676-689) does two things with threads. It calls omp_set_num_threads() on the calling thread, using OMP_NUM_THREADS or the CPU count, and it stores that number as its ThreadsResource limit. It allocates nothing
per-thread there.

Every per-thread array in GM is sized at the moment it is allocated, from the calling thread's omp_get_max_threads():

  • Per-image pixel views: created by AllocateImage (image.c:380, sized at pixel_cache.c:517) and indexed by omp_get_thread_num() (pixel_cache.c:566).
  • Per-operation scratch: omp_data_view.c:93, plus the pixel iterators and the gradient and drawing code.
  • Thread-specific data: the variant sized by OpenMP is only compiled when neither pthreads nor Win32 is available (tsd.c:18-27), so it doesn't apply on our platforms.

In Pascal's example (GM picks 2, darktable restores 16), the restore runs inside dt_init(), before any GM image exists. Every image GM creates afterwards is sized for 16, and its parallel loops run on the same thread with 16 threads.

The stored "2" is only read in two places:

  • the JPEG XL coder (coders/jxl.c:472), which passes it to libjxl's own thread pool; that pool isn't indexed by OpenMP thread number.
  • the gm command-line tool, which darktable doesn't use.

GM does have one real constraint: an image must not be processed by more threads than existed when it was created. darktable creates, reads and destroys each GM image inside one function, with no count change in between.

On "fragile"

The earlier reply said "Codex read the GM sources", which invites the "we depend on GM internals" objection. The stronger argument doesn't need a GM audit:

  • darktable already relied on this. Worker threads set their own count with omp_set_num_threads(dt_get_num_threads()) (src/control/jobs.c:509, :554). The OpenMP thread count is per thread, so GM's init never touched those threads. In the GUI, every GM load has always run with darktable's count, not GM's. That is exactly Pascal's scenario (OMP_NUM_THREADS=2, -t 16). The PR only makes the main thread, which darktable-cli renders on, behave like the workers.
  • GM changes the count after init itself. SetMagickResourceLimit(ThreadsResource, n) calls omp_set_num_threads() (resource.c:1037), and gm benchmark uses it between iterations (command.c:2043). So GM can't assume the count it set at init stays.
  • A second init doesn't undo the reset. libgmic is loaded with the lut3d plugin after the reset. If it calls InitializeMagick() again, GM returns immediately because it is already initialized (magick.c:1202).

Separate problem found while checking (not reproduced)

G'MIC 3.x and 4.x call omp_set_num_threads(nb_cpus) every time a gmic instance is constructed, unless OMP_NUM_THREADS is set (gmic-v.4.0.5/src/gmic.cpp:3603, reached from :4183). darktable constructs one in lut3d for compressed LUTs (src/iop/lut3dgmic.cpp:48, :105). With -t below the CPU count, that would reset the count on the calling thread (a pipe worker, or darktable-cli's main thread) in the middle of a session. The reset in dt_init() can't cover that. The installed G'MIC is 2.9.4, which doesn't do this, so I couldn't reproduce it. This is the genuinely fragile part, and adding num_threads(dt_get_num_threads()) to DT_OMP_FOR would fix it. It belongs in its own issue, not this PR.

@kofa73

kofa73 commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Also: once #22456 and #22457 are fixed, a next fix can be added to allocate buffers based on the thread count, until that changing the allocation would change output. Over to Claude:

When OMP_THREAD_LIMIT is below -t, darktable still sizes its per-thread buffers for -t threads, although at most OMP_THREAD_LIMIT workers ever run. That wastes some memory but cannot overflow. Capping darktable.num_openmp_threads at omp_get_thread_limit() would avoid the waste, but the same count also decides how some modules split their work. X-Trans raw denoise builds its color planes in dt_get_num_threads() row chunks, and the chunk boundaries are where #22456 and #22457 change the output. With -t 4 and OMP_THREAD_LIMIT=1, the cap changed 35 pixels of the 0068 export (by up to 4/255). The integration runner (-t 4, OMP_THREAD_LIMIT=4) is not affected. So I'm keeping that change out of this PR. Once raw denoise no longer depends on the thread count, it can go in a separate PR, after checking other modules that split work by dt_get_num_threads().

@TurboGit

TurboGit commented Oct 1, 2026

Copy link
Copy Markdown
Member

@kofa73 : So if I understand correctly this works with current GM but will fail with GM > 3.x. That's not a nice future :)

What would be your proposal on this?

If we still want to merge this PR I would like to get more comments about the future GM 3.x and above issue that we are going to hit.

@kofa73

kofa73 commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

According to Claude Code's analysis:

This PR neither causes nor fixes it. The PR's reset runs once, in dt_init(). G'MIC changes the count later, whenever lut3d creates a gmic instance. The PR doesn't touch lut3d or G'MIC.

The risk applies to builds linked against a recent G'MIC:

  • cmake/modules/FindGMIC.cmake:10,35 accepts G'MIC versions from 2.7.0 up to (but not including) 10.0, so 3.x and 4.x builds are supported.
  • In 3.7.6 and 4.0.5, every gmic constructor reaches set_variable("_cpus", ...) (gmic.cpp:4183). That calls omp_set_num_threads(nb_cpus) when OMP_NUM_THREADS is unset (:3603). Claude Code hasn't checked when that behavior was introduced.

The path in darktable: commit_params() (src/iop/lut3d.c:1480) → _calculate_clut() → for a compressed .gmz LUT, lut3dgmic.cpp:48 or :105 → the gmic constructor. This runs on the pipe thread, so it affects the GUI too, not just darktable-cli. Each worker thread sets its count only once, when it starts (src/control/jobs.c:509, :554), so the reset count stays in place for the rest of that thread's life.

All three conditions are needed:

  • -t below the CPU count
  • OMP_NUM_THREADS not set
  • lut3d with a compressed LUT

The agent's proposal was above: adding num_threads(dt_get_num_threads()) to DT_OMP_FOR would fix it; a more detailed version:

If the aim is one count that libraries can't override, the fix belongs in darktable's own loops: give every parallel region that uses per-thread buffers an explicit num_threads(dt_get_num_threads()). Then OpenMP never starts more threads than the buffers hold, whatever a library sets, and both the GM reset and the G'MIC risk stop mattering.

  • Changing DT_OMP_FOR (darktable.h:131) covers most of these regions.
  • 18 bare parallel regions outside the macro would need the same clause. I counted them but didn't check which ones use per-thread buffers.

@kofa73 kofa73 mentioned this pull request Oct 2, 2026
6 tasks done
kofa73 added 3 commits October 3, 2026 14:13
GraphicsMagick resets the main thread's OpenMP limit to the CPU count.
With --threads set lower, darktable-cli can then run more threads than
its per-thread buffers hold. Reapply the requested limit after Magick
initialization.

Fixes darktable-org#22436
OpenMP never runs more threads than OMP_THREAD_LIMIT, but darktable
sized its per-thread buffers for the requested count (--threads, or
the CPU count), so the slots above the limit were allocated and never
used.

Lower darktable.num_openmp_threads to the limit before the first
omp_set_num_threads(). OMP_THREAD_LIMIT=N now acts like --threads N
wherever darktable uses its thread count: buffer sizes, the number of
work slices in algorithms such as the bilateral grid, the number of
background worker threads, and whether thumbnails are generated in
the background. Exports match those made with --threads N, which can
differ from the previous output by rounding.
@kofa73
kofa73 force-pushed the 22436-threads-option-undone-by-graphicsmagick-overflows-perthread-buffers branch from d060937 to 6778d31 Compare October 3, 2026 14:27
@kofa73

kofa73 commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

I rebased the branch onto current master. Claude Code added a new commit to cap the thread count and thus allocate memory according to the effective number of threads. We now have 3 commits: the original crash fix, the cap on the thread count, and the release note. Over to the bot:


The new commit, common: cap the OpenMP thread count at OMP_THREAD_LIMIT, limits darktable.num_openmp_threads to omp_get_thread_limit(). OpenMP never runs more threads than OMP_THREAD_LIMIT, so per-thread buffers sized for --threads held slots that no thread used.

With the cap, OMP_THREAD_LIMIT=N acts like --threads N everywhere darktable reads its thread count. The bilateral grid splits its work into N slices. When N is below 4, darktable also starts fewer background worker threads and does not generate thumbnails in the background. Exports match those made with --threads N, which can differ from the old output by rounding. In 0015-shadhi-bilateral with N = 1, 17 pixels differ by 1/255.

The cap needs #22475. Before that change, X-Trans raw denoise split its work by thread count, so the cap changed its output.

The cap has no separate release note (the issue is about the bug fix in the first commit).

Testing (Linux, 12 CPUs, Release build, CPU path):

  • 0068-rawdenoise-xtrans with -t 4, with and without OMP_THREAD_LIMIT=1, is identical to the output of the build without the cap.
  • 0015-shadhi-bilateral with -t 4 and OMP_THREAD_LIMIT=1 is identical to -t 1 without the cap. With OMP_THREAD_LIMIT=3, it is identical to -t 3 without the cap.

The cap change has been reviewed by Codex CLI and Antigravity.

@kofa73

kofa73 commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

@TurboGit : shall I create an issue for the GM >= 3 question? (#22438 (comment))

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

darktable-cli: -t N below the CPU count corrupts the heap

2 participants