Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
46 commits
Select commit Hold shift + click to select a range
cd7d264
fix(chunk-grids): enforce one zero-length-axis invariant across model…
d-v-b Sep 9, 2026
a171507
fix(metadata): read a legacy v2 zero chunk edge on an empty axis as 1
d-v-b Sep 9, 2026
921be0d
chore: rename changelog fragment to the PR number
d-v-b Sep 9, 2026
047a92e
docs: state what 2.x actually did with a zero chunk edge
d-v-b Sep 9, 2026
7a2a71a
Merge branch 'main' into fix/zero-length-single-invariant
d-v-b Sep 12, 2026
2297c62
docs: qualify zero-chunk compatibility history
d-v-b Sep 13, 2026
2873178
Merge branch 'main' into fix/zero-length-single-invariant
d-v-b Sep 16, 2026
d868213
Merge branch 'main' into fix/zero-length-single-invariant
d-v-b Sep 16, 2026
4456d80
Merge branch 'main' into fix/zero-length-single-invariant
d-v-b Sep 19, 2026
49564dc
fix(metadata): read a legacy zero chunk size on an empty axis in Zarr…
d-v-b Sep 19, 2026
5bb1f97
refactor(metadata): check stored chunk shapes against the array shape…
d-v-b Sep 19, 2026
f93193d
refactor(metadata): scope the stored chunk shape check to regular chu…
d-v-b Sep 19, 2026
64b3550
Merge branch 'main' into fix/zero-length-single-invariant
d-v-b Sep 25, 2026
8414ea3
fix(metadata): read a stored zero chunk size on a grown axis as one s…
d-v-b Sep 25, 2026
2bf32df
test: a stateful test of one array's create/append/resize/write life
d-v-b Sep 25, 2026
1723f0b
fix(metadata): stop the zero chunk size warning from naming writers
d-v-b Sep 25, 2026
3d223f1
chore(metadata): drop an orphaned comment in ArrayV2Metadata.__init__
d-v-b Sep 25, 2026
339c514
fix(chunk-grids): size a full-span shard as a multiple of the inner c…
d-v-b Sep 25, 2026
e49eb7e
refactor(metadata): read invalid stored chunk sizes in one module of …
d-v-b Sep 25, 2026
f6f5c7b
docs: rewrite the 4334 changelog fragment and trim restating docstrings
d-v-b Sep 25, 2026
f873358
test: model exactly what the store holds in the array lifecycle state…
d-v-b Sep 25, 2026
d9549b8
refactor(metadata): warn about an upgraded document only once it vali…
d-v-b Sep 25, 2026
5ef0f51
refactor(metadata): one per-axis rule for stored chunk sizes, one rul…
d-v-b Sep 26, 2026
192f86c
refactor(chunk-grids): start chunk guessing from the one full-span rule
d-v-b Sep 26, 2026
2e878cf
test: run the array lifecycle state machine in the slow Hypothesis job
d-v-b Sep 26, 2026
c0d4b3c
fix(array): store upgraded metadata before writing the first chunk
d-v-b Sep 26, 2026
121b6b2
refactor(metadata): name an array one way in upgrade warnings
d-v-b Sep 26, 2026
c1353bf
refactor(metadata): one rule for bare sizes and edge lists; one noun …
d-v-b Sep 26, 2026
f081f1b
fix(array): store the upgrade of the current stored document before w…
d-v-b Sep 26, 2026
9e29333
test: prefer zero-length axes for stored chunk size 0 in the lifecycl…
d-v-b Sep 26, 2026
632ac2b
fix(metadata): keep accepting the chunk sizes zarr 3.4.0 accepted in …
d-v-b Sep 26, 2026
245acae
fix(array): leave a valid stored document as written on a stale handl…
d-v-b Sep 26, 2026
2d0121d
docs: describe the 4334 changes to metadata constructors for a patch …
d-v-b Sep 26, 2026
4c4fe21
fix(metadata): read float edges only in stored rectilinear documents
d-v-b Sep 26, 2026
6fde402
test(codecs): pickle a ShardingCodec with an inner chunk size of 0
d-v-b Sep 26, 2026
b0ea125
fix(group): build every node before create_hierarchy deletes or store…
d-v-b Sep 26, 2026
a903d66
fix(metadata): warn only where the user must act; guard and refresh u…
d-v-b Sep 26, 2026
f78e115
test: expect the lifecycle machine's upgrade warning only on non-empt…
d-v-b Sep 26, 2026
bc306aa
fix(metadata): leave a sharded chunk size of 0 unread when the inner …
d-v-b Sep 26, 2026
3bc802f
perf(metadata): read each upgraded member once, concurrently, and ado…
d-v-b Sep 26, 2026
dc50ec6
fix(group): adopt refreshed consolidated members in place
d-v-b Sep 26, 2026
cddb877
fix(metadata): raise the rectilinear flag error when refreshing a con…
d-v-b Sep 26, 2026
4b9f82c
fix(group): store a group's documents only; keep upgraded members as …
d-v-b Sep 26, 2026
65f3f8e
docs(group): say what consolidated metadata stores for an upgraded array
d-v-b Sep 26, 2026
faa6efe
fix(array): clear the stored document whenever an array stores its me…
d-v-b Sep 26, 2026
30d7ca6
test(metadata): pin that a failed metadata save keeps the stored docu…
d-v-b Sep 26, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion Justfile
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ coverage-serve *args:

# Run slow Hypothesis tests and write coverage.xml
hypothesis *args:
hatch run {{ quote(hatch_env) }}:coverage run --source=src -m pytest -nauto --run-slow-hypothesis tests/test_properties.py tests/test_store/test_stateful* "$@"
hatch run {{ quote(hatch_env) }}:coverage run --source=src -m pytest -nauto --run-slow-hypothesis tests/test_properties.py tests/test_store/test_stateful* tests/test_array_stateful.py "$@"
hatch run {{ quote(hatch_env) }}:coverage xml

# Validate executable documentation code blocks
Expand Down
7 changes: 7 additions & 0 deletions changes/4334.bugfix.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
A chunk edge length is now always at least 1, while an array extent may be 0. Every spelling of one chunk spanning an axis (`chunks=-1`, `chunks=False`, `chunks="auto"`, `shards=-1`, `shards=False`) gives chunk size 1 on a zero-length axis, and a shard spanning an axis is a multiple of the inner chunk size, so `shards=-1` no longer fails when the axis length is not a multiple of it. Rectilinear chunk grids can be created on a zero-length dimension with a non-empty list of positive chunk sizes, which are kept for later growth.

`FixedDimension(size=0, ...)` now raises a `ValueError`. `ArrayV2Metadata(chunks=(0,))` is still accepted and written as given; an array built from such metadata (with `create_hierarchy`, for example) reads the chunk size as a stored chunk size of 0 is read (below). `create_hierarchy` now builds every array and group before it deletes or stores anything, so a node that cannot be built fails with the store untouched. The regular chunk grid metadata class reads a `bool` chunk edge length as the `int` it equals, and rejects a string or a mapping as a chunk shape as a whole. `RectilinearChunkGridMetadata`, which is experimental, now reads a `bool` edge length as the `int` it equals and rejects a NumPy integer or a float edge length such as `4.0` with a `TypeError`; it used to keep them as given, so that it could not store a NumPy integer, could not read back a `bool`, and stored a float as a JSON float, which the Zarr specification does not allow.

Stored metadata with a regular chunk size of 0 or JSON `false` (Zarr format 2 `chunks`, Zarr format 3 `regular` `chunk_shape`) is now read as one chunk spanning the axis (a multiple of the inner chunk size for sharded arrays, which previously failed to open). zarr-python wrote such sizes for arrays created with a zero-length axis until 3.4, and, from 2.18.7 to 3.2.1, for an explicit chunk size of 0 or `False` on an axis of any length, as in `chunks=(0,)` (those arrays could store no data). On a zero-length axis this opens silently, as before; appending to such an axis then stores chunks of size 1. On an axis of positive length it opens with a `ZarrUserWarning`, which says that the axis holds only the fill value and how to store valid metadata: `array.update_attributes({})`, then `zarr.consolidate_metadata` if the metadata is consolidated. JSON `true`, as zarr-python 3.0 and 3.2 wrote for a chunk size of `True`, is read as 1, silently, in the chunk grid (regular, and the explicit edges and run-length encoded sizes of a rectilinear one) and in the inner chunk shape of every sharding codec, nested or not. A stored rectilinear chunk grid whose edge lengths are integral JSON floats (`[[4.0, 2]]`), as zarr-python 3.2 wrote for float edges, is read with those edges as integers, silently; a float anywhere no release wrote one (a regular chunk shape, the inner chunk shape of a sharding codec, a run-length repeat count, an edge below 1) is rejected.

Writing data to an array read from such metadata (other than an empty selection) first stores the upgrade of the metadata the store then holds, unless another writer has stored valid metadata since, so that other readers find the chunks written; in a `ZipStore` this adds a second entry for the metadata document, as every metadata update does. If the metadata the store then holds lays out chunks differently (another writer resized the array keeping the chunk size of 0, say), the write raises a `ValueError` asking to reopen the array, and stores nothing. Apart from the array's own metadata writes (`update_attributes`, `resize`), which store the upgrade as before, this is the only operation that stores it: a group's consolidated metadata keeps such an array's metadata as it was stored (`zarr.consolidate_metadata` copies it as the array's own document holds it) until the array stores its upgrade through the group's handle, and changing a group reads and writes no metadata of its members.
6 changes: 6 additions & 0 deletions docs/user-guide/arrays.md
Original file line number Diff line number Diff line change
Expand Up @@ -707,6 +707,12 @@ z.append(np.arange(10, dtype='float64'))
print(f"After append: shape={z.shape}, chunk_sizes={z.write_chunk_sizes}")
```

A rectilinear array can also be created with a zero-length dimension: because no
non-empty list of positive chunk sizes can sum to 0, the chunk sizes given for such a
dimension are stored as-is and describe the chunks the dimension will grow into
on `append` or `resize` — the same state as resizing an existing rectilinear
dimension down to 0.

### Compressors and filters

Rectilinear arrays work with all codecs — compressors, filters, and checksums.
Expand Down
6 changes: 6 additions & 0 deletions src/zarr/core/_json.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,12 @@ def buffer_to_json(buffer: Buffer) -> JSON:
return cast("JSON", json.loads(buffer.to_bytes()))


def json_equal(a: JSON, b: JSON) -> bool:
"""Whether two JSON values have the same JSON encoding. Python compares `True` and
`1`, or `1.0` and `1`, as equal; JSON does not."""
return json.dumps(a) == json.dumps(b)


def buffer_to_json_object(buffer: Buffer) -> dict[str, JSON]:
"""Parse the contents of a `Buffer` as a JSON object (a `dict`).

Expand Down
169 changes: 86 additions & 83 deletions src/zarr/core/array.py
Original file line number Diff line number Diff line change
Expand Up @@ -118,7 +118,13 @@
ArrayV2MetadataDict,
ArrayV3Metadata,
)
from zarr.core.metadata.io import save_metadata
from zarr.core.metadata.io import (
ARRAY_DOCUMENTS,
parse_stored_array,
read_documents,
save_metadata,
upsert_metadata,
)
from zarr.core.metadata.v2 import (
CompressorLikev2,
get_object_codec_id,
Expand Down Expand Up @@ -201,13 +207,22 @@ def _chunk_sizes_from_shape(
return tuple(result)


def parse_array_metadata(data: Any) -> ArrayMetadata:
def parse_array_metadata(data: Any, path: str | None = None) -> ArrayMetadata:
"""Array metadata from a metadata object or a metadata document, naming the array at
`path` in warnings about how an invalid document was read.

`ArrayV2Metadata` accepts a chunk size of 0, as it always has, though only an
invalid document holds one: such metadata is read as the documents it would store
are (see `zarr.core.metadata.upgrades`), so an array can be built from it. No data
was read or written under that chunk size, so the reading is silent."""
if isinstance(data, ArrayV2Metadata) and 0 in data.chunks:
return parse_stored_array(data.to_buffer_dict(default_buffer_prototype()), 2)
if isinstance(data, ArrayMetadata):
return data
elif isinstance(data, dict):
if isinstance(data, dict):
zarr_format = data.get("zarr_format")
if zarr_format == 3:
meta_out = ArrayV3Metadata.from_dict(data)
meta_out = ArrayV3Metadata.from_dict(data, path=path)
if len(meta_out.storage_transformers) > 0:
msg = (
f"Array metadata contains storage transformers: {meta_out.storage_transformers}."
Expand All @@ -216,7 +231,7 @@ def parse_array_metadata(data: Any) -> ArrayMetadata:
raise ValueError(msg)
return meta_out
elif zarr_format == 2:
return ArrayV2Metadata.from_dict(data)
return ArrayV2Metadata.from_dict(data, path=path)
else:
raise ValueError(f"Invalid zarr_format: {zarr_format}. Expected 2 or 3")
raise TypeError # pragma: no cover
Expand Down Expand Up @@ -404,7 +419,7 @@ def __init__(
store_path: StorePath,
config: ArrayConfigLike | None = None,
) -> None:
metadata_parsed = parse_array_metadata(metadata)
metadata_parsed = parse_array_metadata(metadata, str(store_path))
config_parsed = parse_array_config(config)

object.__setattr__(self, "metadata", metadata_parsed)
Expand Down Expand Up @@ -765,7 +780,7 @@ def from_dict(
ValueError
If the dictionary data is invalid or incompatible with either Zarr format 2 or 3 array creation.
"""
metadata = parse_array_metadata(data)
metadata = parse_array_metadata(data, str(store_path))
return cls(metadata=metadata, store_path=store_path)

@classmethod
Expand Down Expand Up @@ -1610,10 +1625,50 @@ async def get_coordinate_selection(
return out_array

async def _save_metadata(self, metadata: ArrayMetadata, ensure_parents: bool = False) -> None:
"""
Asynchronously save the array metadata.
"""
"""Store `metadata` as this array's own documents, then clear the
`_stored_document` mark (see `_stored_document_replaced`)."""
await save_metadata(self.store_path, metadata, ensure_parents=ensure_parents)
self._stored_document_replaced()

def _stored_document_replaced(self) -> None:
"""Record that the store no longer holds a document of this array that needs an
upgrade: it holds the upgrade, a valid document, or none. The metadata this handle
holds, which a consolidated group handle may share, then stops standing for the
document it was read from (see `mark_upgraded`), so no later write through either
handle stores that document again."""
object.__setattr__(self.metadata, "_stored_document", None)

async def _store_upgraded_document(self) -> None:
"""Store the upgrade of this array's current stored document, if it needs one,
before chunks are written under this handle's metadata.

Only for metadata read from a document that had to be upgraded (see
`zarr.core.metadata.upgrades`). The document is read again, because the store may
hold a newer one than this handle's metadata. If that one lays out chunks
differently (the array was resized since by software that kept the invalid chunk
size), this handle would write chunks no reader finds, so it raises and stores
nothing. If it needs no upgrade (the array was re-saved since, possibly by another
implementation), it is left as written; if there is none, there is nothing to
upgrade. Storing the same upgrade twice is harmless, so concurrent callers need
no coordination.
"""
if self.metadata._stored_document is None:
return
zarr_format = self.metadata.zarr_format
documents = await read_documents(self.store_path, ARRAY_DOCUMENTS[zarr_format])
try:
current = parse_stored_array(documents, zarr_format)
except ArrayNotFoundError:
pass
else:
if _chunk_layout(current) != _chunk_layout(self.metadata):
raise ValueError(
f"The metadata stored for the array at {str(self.store_path)!r} has "
"changed since this array was opened: reopen the array to write to it."
)
if current._stored_document is not None:
await upsert_metadata(self.store_path, current, documents)
self._stored_document_replaced()

async def _set_selection(
self,
Expand All @@ -1623,6 +1678,10 @@ async def _set_selection(
prototype: BufferPrototype,
fields: Fields | None = None,
) -> None:
if product(indexer.shape) > 0:
# Chunks are about to be stored under the upgraded metadata, so store it
# first: every reader of the store then agrees with them.
await self._store_upgraded_document()
return await _set_selection(
self.store_path,
self.metadata,
Expand Down Expand Up @@ -1674,16 +1733,10 @@ async def setitem(
- This method is asynchronous and should be awaited.
- Supports basic indexing, where the selection is contiguous and does not involve advanced indexing.
"""
return await _setitem(
self.store_path,
self.metadata,
self.codec_pipeline,
self.config,
self._chunk_grid,
selection,
value,
prototype=prototype,
)
if prototype is None:
prototype = default_buffer_prototype()
indexer = BasicIndexer(selection, shape=self.metadata.shape, chunk_grid=self._chunk_grid)
return await self._set_selection(indexer, value, prototype=prototype)

@property
def oindex(self) -> AsyncOIndex[T_ArrayMetadata]:
Expand Down Expand Up @@ -4846,6 +4899,16 @@ async def create_array(
)


def _chunk_layout(
metadata: ArrayMetadata,
) -> tuple[tuple[int, ...] | ChunkGridMetadata, tuple[int, ...] | None]:
"""How an array's chunks are laid out: its chunk grid and, if it is sharded, the
inner chunk shape."""
grid = metadata.chunks if isinstance(metadata, ArrayV2Metadata) else metadata.chunk_grid
sharding = _sharding_codec(metadata)
return grid, None if sharding is None else sharding.chunk_shape


def _sharding_codec(metadata: ArrayMetadata) -> ShardingCodec | None:
"""The array's sharding codec, or None if the array is not sharded.

Expand Down Expand Up @@ -5827,58 +5890,6 @@ async def _set_selection(
)


async def _setitem(
store_path: StorePath,
metadata: ArrayMetadata,
codec_pipeline: CodecPipeline,
config: ArrayConfig,
chunk_grid: ChunkGrid,
selection: BasicSelection,
value: npt.ArrayLike,
prototype: BufferPrototype | None = None,
) -> None:
"""
Set values in the array using basic indexing.

Parameters
----------
store_path : StorePath
The store path of the array.
metadata : ArrayMetadata
The array metadata.
codec_pipeline : CodecPipeline
The codec pipeline for encoding/decoding.
config : ArrayConfig
The array configuration.
chunk_grid : ChunkGrid
The chunk grid.
selection : BasicSelection
The selection defining the region of the array to set.
value : npt.ArrayLike
The values to be written into the selected region of the array.
prototype : BufferPrototype or None, optional
A prototype buffer that defines the structure and properties of the array chunks being modified.
If None, the default buffer prototype is used.
"""
if prototype is None:
prototype = default_buffer_prototype()
indexer = BasicIndexer(
selection,
shape=metadata.shape,
chunk_grid=chunk_grid,
)
return await _set_selection(
store_path,
metadata,
codec_pipeline,
config,
chunk_grid,
indexer,
value,
prototype=prototype,
)


async def _resize(
array: AsyncArray[ArrayV2Metadata] | AsyncArray[ArrayV3Metadata],
new_shape: ShapeLike,
Expand Down Expand Up @@ -5928,7 +5939,7 @@ async def _delete_key(key: str) -> None:
)

# Write new metadata
await save_metadata(array.store_path, new_metadata)
await array._save_metadata(new_metadata)

# Update metadata and chunk_grid (in place)
object.__setattr__(array, "metadata", new_metadata)
Expand Down Expand Up @@ -5993,15 +6004,7 @@ async def _append(
slice(None) if i != axis else slice(old_shape[i], new_shape[i])
for i in range(len(array.shape))
)
await _setitem(
array.store_path,
array.metadata,
array.codec_pipeline,
array.config,
array._chunk_grid,
append_selection,
data,
)
await array.setitem(append_selection, data)

return new_shape

Expand All @@ -6028,7 +6031,7 @@ async def _update_attributes(
array.metadata.attributes.update(new_attributes)

# Write new metadata
await save_metadata(array.store_path, array.metadata)
await array._save_metadata(array.metadata)

return array

Expand Down
Loading
Loading