Skip to content

feat(codecs): deprecate the host-byte-order default of BytesCodec on big-endian hosts - #4439

Draft
d-v-b wants to merge 9 commits into
zarr-developers:mainfrom
d-v-b:deprecate/bytes-codec-host-endian
Draft

d-v-b wants to merge 9 commits into
zarr-developers:mainfrom
d-v-b:deprecate/bytes-codec-host-endian

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

🤖 AI text below 🤖

This PR depends on #4435 and is based on its branch. Until #4435 merges, the diff here includes its commits too. Only the last two commits belong to this PR: the deprecation and the changelog rename.

A bare BytesCodec() defaults to endian=sys.byteorder. So the same code writes big-endian chunks on s390x and little-endian chunks everywhere else. That output is valid and readable on every host, but it is not identical across machines.

History: the bytes codec in the original zarrita import defaulted to "little". The host-order default came in with #1660 ("Remove attrs"), and neither that PR's description nor its comments discuss the switch. #3968 later pinned it with a test.

Changing the default back would change what existing code writes on big-endian hosts. So this PR only deprecates the current default; the switch itself is left for a later release.

Changes

  • The deprecation: when endian is omitted on a big-endian host, BytesCodec() raises a ZarrFutureWarning. The codec still uses "big". The warning says to pass endian="big" to keep the current behaviour, or endian="little" to adopt the new default now.
    • Nothing changes on little-endian hosts, where the current default is already "little".
    • BytesCodec.from_dict and zarr's default serializer always pass endian, so they never warn.
    • The warning uses ZarrFutureWarning rather than DeprecationWarning so that end users see it. It points at the caller's line.
  • Internal callers: callers that relied on the default now pass endian="little". These are the zarr.testing strategies and state machine, the create_array docstrings, and about 70 test call sites. That includes two tests that passed the class itself as a factory (default_factory=BytesCodec), which the s390x run caught.
  • Tests:
    • test_bytes_codec_endian checks the resolved endian for each combination of host byte order and argument.
    • test_bytes_codec_default_endian_warns_on_big_endian_host checks the warning.
    • Both monkeypatch sys.byteorder, so ordinary CI covers the big-endian case too.

Testing

  • Full suite on this Mac (little-endian): passes.
  • Full suite on emulated s390x, run by the agent: the only failures are the 3 test_examples tests, because the test image has no uv. The first run also exposed the two factory call sites mentioned above; the agent fixed them and re-ran those files on s390x, and they now pass.

Refs #3438.

🤖 Generated with Claude Code

… hosts

Running the suite on s390x (QEMU, Debian sid) gave 110 failures on main.
Most came from one asymmetry: V3 data types without an explicit byte
order defaulted to little-endian, while NumPy dtype names mean host order.
On a big-endian host `dtype="float64"` and `dtype=np.float64` produced
different data types, and a reopened array changed dtype.

- V3 data types default to the host byte order (`HasEndianness`), because
  V3 metadata carries none. Stored chunks stay little-endian by default.
- `ShardingCodec` defaults its inner and index `bytes` codecs to
  little-endian, so sharded output no longer depends on the host.
- `scale_offset` computes in native byte order and restores the input's
  byte order; `>f8` and `>u8` arrays failed on every host.
- Base64 fill values of structured data types in V3 metadata are
  little-endian bytes in both directions.
- Tests no longer assume a little-endian host: explicit `<` dtypes and
  `endian="little"` where the expectation is little, and a big-endian case
  for migration and pipeline parity.

Refs zarr-developers#3438

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-opus-5-5
A weekly and path-filtered job that runs the data type, codec and metadata
tests on an emulated big-endian host, in a Debian sid image that takes numpy
and numcodecs from Debian packages so nothing heavy compiles under emulation.

Assisted-by: ClaudeCode:claude-opus-5-5
`BytesCodec()` defaulted to `sys.byteorder`, so the same code wrote
big-endian chunks on s390x and little-endian chunks elsewhere. The bytes
codec originally defaulted to "little"; the host default came in as a side
effect of replacing attrs with dataclasses (zarr-developers#1660), and was later pinned
by a test that only restated it. `"little"` also matches
`default_serializer_v3`.

Chunks written on big-endian hosts with a bare `BytesCodec()` are now
little-endian. Chunks written before this change remain readable, because
the codec records its byte order in the metadata.

Refs zarr-developers#3438

Assisted-by: ClaudeCode:claude-opus-5-5
This reverts commit 8dd33e3. Changing the default of `BytesCodec()`
changes what existing code writes on big-endian hosts, so it needs a
deprecation cycle and will land in its own PR.

Assisted-by: ClaudeCode:claude-opus-5-5
…big-endian hosts

`BytesCodec()` without an `endian` argument means `sys.byteorder`, so the
same code writes big-endian chunks on s390x and little-endian chunks
elsewhere. The default will become "little" on every host, which is what
the codec originally defaulted to before zarr-developers#1660 and what the default
serializer uses. On big-endian hosts, omitting `endian` now raises a
ZarrFutureWarning; nothing changes on little-endian hosts.

Callers inside zarr (`zarr.testing` strategies, docstrings) and the tests
now pass `endian` explicitly, so they do not depend on the default.

Refs zarr-developers#3438

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-opus-5-5
@codecov

codecov Bot commented Sep 27, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.42%. Comparing base (0401a7f) to head (17d8190).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4439      +/-   ##
==========================================
+ Coverage   94.37%   94.42%   +0.04%     
==========================================
  Files          93       93              
  Lines       13174    13198      +24     
==========================================
+ Hits        12433    12462      +29     
+ Misses        741      736       -5     
Files with missing lines Coverage Δ
src/zarr/api/asynchronous.py 96.32% <ø> (ø)
src/zarr/api/synchronous.py 92.95% <ø> (ø)
src/zarr/codecs/bytes.py 98.91% <100.00%> (+0.13%) ⬆️
src/zarr/codecs/scale_offset.py 100.00% <100.00%> (ø)
src/zarr/codecs/sharding.py 95.53% <ø> (ø)
src/zarr/core/array.py 98.09% <ø> (ø)
src/zarr/core/dtype/common.py 92.22% <100.00%> (+0.08%) ⬆️
src/zarr/core/dtype/npy/structured.py 97.14% <100.00%> (+2.53%) ⬆️
src/zarr/testing/stateful.py 36.46% <ø> (ø)
src/zarr/testing/strategies.py 96.62% <ø> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant