Repository navigation
(m)TLS to PG - #627
(m)TLS to PG#627shortishly wants to merge 564 commits into
Conversation
|
|
||
| #[test] | ||
| fn test_salted_password_256() -> Result<()> { | ||
| let password = "password"; |
Check failure
Code scanning / CodeQL
Hard-coded cryptographic value Critical
| #[test] | ||
| fn test_salted_password_256() -> Result<()> { | ||
| let password = "password"; | ||
| let salt = b"abcdef"; |
Check failure
Code scanning / CodeQL
Hard-coded cryptographic value Critical
|
|
||
| #[test] | ||
| fn test_salted_password_512() -> Result<()> { | ||
| let password = "password"; |
Check failure
Code scanning / CodeQL
Hard-coded cryptographic value Critical
| #[test] | ||
| fn test_salted_password_512() -> Result<()> { | ||
| let password = "password"; | ||
| let salt = b"abcdef"; |
Check failure
Code scanning / CodeQL
Hard-coded cryptographic value Critical
nisshi-io is on GitHub's Free plan (20 concurrent jobs). One ci.yml run is ~60 jobs, 31 of them with no `needs`, so every pull_request push starved every other PR of runners. This splits the workflow so pull_request pushes run a small Tier A and the full Tier B runs once per merge-queue entry. Tier A (every pull_request push): check, fmt, clippy, typos, third-party-license, test on postgres:17 only, the non-experimental (memory) leg of each compat suite, ci-gate. Tier B (merge_group and push): full build-storage/build-storage-lake matrix, test on pg16/pg18, experimental compat legs, cargo-publish-dry-run, src, package, smoke, release. Tier B jobs carry `if: github.event_name != 'pull_request'`; trimmed matrices choose the list by event so check-run names stay stable across tiers. `test` no longer `needs` build-storage/build-storage-lake, which would otherwise cascade a skip onto Tier A. ci-gate: `smoke` is dropped from `needs` (advisory, still runs). Its compose setup pulls a minio image with a recurring anonymous-pull auth failure that lives in nisshi-io/example-java; in a queue that flake would eject good entries. A second step fails the gate if any listed job is `skipped` on a non-pull_request run, so a misconfigured Tier B job can't pass as "didn't fail". Adds the concurrency group codeql.yml already uses (cancel superseded pull_request runs only; cancelling a merge_group run fails the entry). codeql.yml, workflow-lint.yml, dependencies.yml gain the same merge_group trigger: their jobs are required checks, and a required check that never reports on the queue branch hangs every entry until it's ejected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HjykxQXwhz6Bfkfte3bbcb
Review fold-in. The queue fast-forwards main to the exact commit it tested, so the merge_group run and the post-merge push run share a sha, and a `v*` tag on main's tip makes a third. Keyed on sha alone they'd share a concurrency group: the push run would pend behind the queue run, and GitHub cancels a pending run whenever a newer one arrives in the group regardless of `cancel-in-progress`. Off pull_request the group is now `ref@sha` in both ci.yml and codeql.yml. ci-gate gains a pull_request-side check that no Tier A job was `skipped`, which is what a stray `needs` onto a Tier B job would look like; the header comment warned about it, now the gate enforces it. The smoke comment stated a re-add condition that had already been met (example-java #4 builds minio from source as of today); it now states the actual condition. The header no longer claims the queue avoids a post-merge double run: the push run does repeat Tier B on the same tree, at the same cost as today's post-merge run. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HjykxQXwhz6Bfkfte3bbcb
Reverts the demotion of `smoke` to advisory from earlier in this branch. The minio anonymous-pull failure that motivated it was fixed at its source today (nisshi-io/example-java #4 builds minio from source) and main's next run was green on it, so there is nothing left to wait for. `smoke` is Tier B like the rest of the package chain: skipped on pull_request, required to run and succeed on merge_group and push. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HjykxQXwhz6Bfkfte3bbcb
A group request that finds nothing cached starts from an empty group with no version, relying on storage to reject the write as outdated. SlateDB compared versions only when one was given, so the empty group overwrote the stored one and dropped every member (UnknownMemberId). Overlapping transactions also surfaced as a raw "transaction conflict" error instead of Outdated. - Require the caller's version to match, with no version matching only a missing group, as the other backends already do. - Map a transaction-conflict commit to Outdated with the stored group, so the coordinator retries against it. - Document the version contract on Storage::update_group. - Add regression tests for overlapping heartbeats and for a versionless update of an existing group, on every backend. - Fix the misleading panic message in new_cg.rs. Fixes #757 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…java Review fold-in from #759. Tier B `if:`s and the two trimmed-matrix selectors now also key on `vars.MERGE_QUEUE == 'on'`. The split only makes sense while a merge queue runs Tier B before merge; with the variable unset, pull_request runs everything as before, so turning the queue off restores full PR coverage without a revert, and the window between this merge and the ruleset flip loses nothing. `release` gains `needs: build-storage, build-storage-lake` (both Tier B): a feature combination that doesn't compile now stops a queue entry before the release -> package -> smoke chain and never publishes on a `v*` tag. Those edges were cut from `test` because `test` is Tier A. ci-gate prints every job's result to the step summary and names the jobs behind each failure. The pull_request skip check is inverted to a Tier B exemption list, so a newly added job defaults to must-not-skip instead of depending on someone extending a Tier A list. `smoke`'s example-java checkout is pinned to 896d854 (minio built from source, example-java #4). It gates every merge, so changes there should arrive as a reviewed bump here rather than on push to its default branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HjykxQXwhz6Bfkfte3bbcb
With MERGE_QUEUE unset (the state from merge until the flip, and during any rollback) Tier B runs on pull_request again, including fork PRs. A fork PR's token is read-only, so `push: true` on package failed and took ci-gate with it. Restore main's guard: push only on repo-owned refs and same-repo PRs, and skip smoke on fork PRs so it doesn't pull an image that was never pushed. smoke is in the gate's Tier B exemption list, so that skip still passes ci-gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HjykxQXwhz6Bfkfte3bbcb
test: build each crate's integration tests as one binary
Review follow-ups for the SlateDB update_group fix: - Narrow the Storage::update_group doc to what every backend enforces; backends still disagree on a version for a group with no stored state. - Note in SlateDB that no version also matches the empty placeholder offset_commit stores, and log the group id on a commit conflict. - Rename concurrent_heartbeat to group_cache_miss and add heartbeats through a new coordinator over the same storage, as after a restart. - Add stale-version and same-version race cases to update_group_conditional, reading the stored group back through describe_groups. The race pins the transaction-conflict mapping and the group it re-reads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move group_cache_miss and update_group_conditional into tests/it/ and declare them in main.rs, following the single test binary layout from #756. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| "actual: {}, v.len: {}", | ||
| actual.maximum_allocation_size()?, | ||
| v.len() |
Check failure
Code scanning / CodeQL
Cleartext logging of sensitive information High test
| "actual: {}, v.len: {}", | ||
| actual.maximum_allocation_size()?, | ||
| v.len() |
Check failure
Code scanning / CodeQL
Cleartext logging of sensitive information High test
ci: two-tier split and merge_group trigger for a merge queue
rust-toolchain.toml and the Dockerfile's chef stage both pin 1.98; the line here still said 1.93. The CI Pipeline section described the pre-#759 single chain with the pg16/17/18 matrix on every run; it now describes the two tiers, ci-gate, the merge queue, and the MERGE_QUEUE switch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HjykxQXwhz6Bfkfte3bbcb
…e-only fix(slatedb): stop concurrent group requests from dropping members
docs: CLAUDE.md toolchain is 1.98, CI pipeline is two-tier
Bumps [taiki-e/install-action](https://github.com/taiki-e/install-action) from 2.87.13 to 2.87.18. - [Release notes](https://github.com/taiki-e/install-action/releases) - [Changelog](https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md) - [Commits](taiki-e/install-action@26e9283...dfae9bf) --- updated-dependencies: - dependency-name: taiki-e/install-action dependency-version: 2.87.18 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps [actions-rust-lang/setup-rust-toolchain](https://github.com/actions-rust-lang/setup-rust-toolchain) from 1.17.0 to 2.0.0. - [Release notes](https://github.com/actions-rust-lang/setup-rust-toolchain/releases) - [Changelog](https://github.com/actions-rust-lang/setup-rust-toolchain/blob/main/CHANGELOG.md) - [Commits](actions-rust-lang/setup-rust-toolchain@166cdcf...ecabd13) --- updated-dependencies: - dependency-name: actions-rust-lang/setup-rust-toolchain dependency-version: 2.0.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
…olchain v2.0.0 v2.0.0 stopped defaulting RUSTFLAGS to "-D warnings" and instead defaults the new build-warnings input to "deny" (CARGO_BUILD_WARNINGS), so the release job's existing rustflags: "" override no longer suppresses warning-denial on its own. Set build-warnings: "" alongside it to keep that job's original behavior. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FgYrQE9srCLhwPTsLn2Rjt
…/taiki-e/install-action-2.87.18 build(deps): bump taiki-e/install-action from 2.87.13 to 2.87.18
Bumps [digest](https://github.com/RustCrypto/traits) from 0.10.7 to 0.11.3. - [Commits](RustCrypto/traits@digest-v0.10.7...digest-v0.11.3) --- updated-dependencies: - dependency-name: digest dependency-version: 0.11.3 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com>
….11.3 build(deps): bump digest from 0.10.7 to 0.11.3
Bumps [console](https://github.com/console-rs/console) from 0.16.4 to 0.16.6. - [Release notes](https://github.com/console-rs/console/releases) - [Changelog](https://github.com/console-rs/console/blob/main/CHANGELOG.md) - [Commits](console-rs/console@0.16.4...0.16.6) --- updated-dependencies: - dependency-name: console dependency-version: 0.16.6 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps [futures](https://github.com/rust-lang/futures-rs) from 0.3.33 to 0.3.34. - [Release notes](https://github.com/rust-lang/futures-rs/releases) - [Changelog](https://github.com/rust-lang/futures-rs/blob/main/CHANGELOG.md) - [Commits](rust-lang/futures-rs@0.3.33...0.3.34) --- updated-dependencies: - dependency-name: futures dependency-version: 0.3.34 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps [http-body-util](https://github.com/hyperium/http-body) from 0.1.4 to 0.1.5. - [Release notes](https://github.com/hyperium/http-body/releases) - [Commits](hyperium/http-body@http-body-util-v0.1.4...http-body-util-v0.1.5) --- updated-dependencies: - dependency-name: http-body-util dependency-version: 0.1.5 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com>
…fetches A slatedb log start moved by compaction in v0.7.0-pre.1 or pre.2 stays where it is after the upgrade, and the new bounds check now resets a group behind it. The Changed entry names that cause and the check to run before upgrading; the Fixed entry describes what changes instead (ListOffsets earliest and log_start_offset in Fetch stay at the log start through compaction). A ReadCommitted fetch between the last stable offset and the high watermark returns no records, so the stage read before the fetch answers it again rather than being read once more on every round. The above_high_watermark help text says it shows only whether parked fetches exist, and that a short burst on multi-broker dynostore is expected. Tests: offset_stage_reads uses a stage that grows on each read and checks that the answer covers the last record, with a case ending one past the high watermark; compacted_header gains a slatedb leg, which fails against the engine before this branch's compaction change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011P97dHLhTRdJ35fMJFpYPg Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com>
Compaction by the old version keeps raising the log start between the pre-upgrade check and the upgrade, so the check alone can miss a group. A reset to earliest lands on the batch the old version would have served. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011P97dHLhTRdJ35fMJFpYPg Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com>
…-keys fix(dynostore): group ids containing / or empty collapsed onto the same key, letting an unfenced write overwrite committed offsets
…unds fix: answer OFFSET_OUT_OF_RANGE for a fetch offset outside the partition
Releases no longer build `x86_64-apple-darwin`. Apple silicon (`aarch64-apple-darwin`) stays as the only macOS target, and the Linux targets are unchanged. Removes the Intel leg from the `release` matrix in `ci.yml`, which saves one macOS job per release run, and fixes the rust-cache comment that mentioned two macOS legs. `ci-gate` lists `release` as a whole, so no required check changes. Nothing else in the repo referenced the removed target or its tarball. Heads up: the latest published release (v0.7.0-pre.2) ships `nisshi-x86_64-apple-darwin.tar.gz`, so a link pinned to that asset will 404 after the next release. The CHANGELOG notes it. Verified with `actionlint` on `ci.yml`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
A push to a fork PR runs all of Tier B today. Repository variables are empty on a fork's `pull_request` run, so `vars.MERGE_QUEUE != 'on'` is true and the Tier B guard lets everything through. **Evidence:** `MERGE_QUEUE` is `on`, yet run 37786893003 on fork PR #863 ran `build-storage`, `build-storage-lake`, `cargo-publish-dry-run`, `src` and the experimental compat legs. **Change** - The six Tier B `if:` conditions get the same same-repo check `smoke` already had. - The three matrices that pick a smaller list while the queue is on (`test`, both compat jobs) also pick it for fork PRs. - Comments and `CLAUDE.md` updated to match. **Behavior** | Run | Tier B | |---|---| | Same-repo PR, queue on | skipped | | Same-repo PR, queue off | runs | | Fork PR, either | skipped | | `merge_group`, push to `main`, tag | runs | A fork PR still gets full Tier B in the merge queue before it merges, and `ci-gate` already tolerates the skips on `pull_request`. Fork PRs now run less, so nothing is weakened. One trade-off: with the queue turned off as a rollback, a fork PR's Tier B runs only after it merges. **Verified:** `actionlint` passes, and I checked the expressions for the same-repo and fork cases, queue on and off, and for runs with no `pull_request` payload. A fork PR run on this branch hasn't happened yet; #863 will be the first. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
The `build-storage-lake` job is 3 builds instead of 12, and `turso` now
has its own storage build.
**Why**
- The lake job ran `cargo build --features <storage> --features <lake>`
from the workspace root. That leaves every member's default features on,
so all 12 legs were close to the same all-features build, and `test` and
`clippy` already compile that.
- The lake features (`nisshi-schema`) and the storage features
(`nisshi-broker`, `nisshi-storage`) don't depend on each other, so one
storage covers every lake format.
**Changes**
- `build-storage-lake`: three legs (delta, iceberg, parquet) through
`just build dev dynostore,<lake>`, the same recipe `build-storage` uses,
so only those features are on.
- `build-storage` and the `build-storage` just recipe: add `turso`. It
had no standalone build.
- Remove the unused `libkrb5-dev` install from those two jobs. Nothing
in `Cargo.lock` needs it.
- Refresh the stale tier comment in `ci.yml` (the org is on Enterprise
now) and the CI and feature-flag text in `CLAUDE.md`.
**Verified**
- `actionlint` passes on `ci.yml`.
- `just build dev turso` and `just build dev
dynostore,{delta,iceberg,parquet}` all build locally on macOS.
- Nothing keys off the removed legs: the `main` ruleset requires only
`ci-gate` and other jobs by name, and `ci-gate` lists jobs, not matrix
legs.
- Not yet verified: the new legs on ubuntu-latest. CI on this PR is the
first Linux run.
If the `libkrb5-dev` removal turns a leg red, I'll put that line back.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
## What does this change do, and why? DeleteGroups routed straight to storage and deleted a group's state and committed offsets regardless of whether it still had members, on every backend. Worse, deleting storage state out from under a live member did not even stay deleted: the next heartbeat's conditional group-detail write is an upsert keyed on a stale cached version, so it silently recreated the row — the delete wasn't even durable while a member was attached. ### Change - Check emptiness through the coordinator instead of storage alone, since member expiry here is lazy (a stale member is only evicted when some other request for the group happens to arrive — nothing proactively times it out): `Coordinator::delete_groups` checks a group's cached state out the same way join/sync/heartbeat/leave already do, applies `missed_heartbeat` to evict anyone past their session timeout, and only then decides. A still-non-empty group is put back and refused with `NON_EMPTY_GROUP`; an empty one is deleted from storage and its cached entry is dropped rather than reinserted, so nothing can resurrect it from a stale version. A group with no cached state (never joined on this broker) falls back to `Storage::describe_groups`. - That fallback exposed a real bug in `describe_groups` on Postgres and libSQL: a group that only ever committed offsets has a group row but no detail row, and the LEFT JOIN query returns NULL for it. Reading that NULL `detail` column errored, failing the whole DeleteGroups request rather than just that one group. Both backends (and limbo/turso, which shares the same query) now treat a NULL detail as an empty group — this also fixes the same latent gap in DescribeGroups. - `nisshi-storage`'s `DeleteGroupsService` is unchanged and still used directly by tests exercising the raw storage primitive; only the wire route for `DeleteGroupsRequest` moved, from the storage-only router to the coordinator-only one, since the check needs the coordinator's in-memory state and storage already lives behind it. ### Known limitation, filed as a follow-up rather than closed here The check-then-delete is not fully atomic. Checking the group out of the coordinator's cache narrows the window against a concurrent join/sync/heartbeat/leave for the same group on this broker, but doesn't close it: a request landing in that window self-heals via the existing `UpdateError::Outdated` retry path, reinserting a fresh, versioned cache entry — if that happens before the storage delete completes, the delete still goes ahead and that freshly-reinserted entry is left stale, reopening the same problem on the *same* broker (a second broker sharing the same storage without going through this coordinator is a separate, additional way to hit the same gap). Closing it fully needs a version-guarded delete added to the `Storage` trait across all 5 backends (pg, lite, limbo, dynostore, slatedb) — out of scope for this ticket's size. ## How was this tested? - Tests across all 4 broker-integration backends (in-memory, libSQL, SlateDB, Postgres): a joined consumer blocks the delete and its offsets survive; after it leaves, the delete succeeds; an expired session (no LeaveGroup) is still deletable, proving member expiry doesn't block a legitimately-abandoned group forever; an offsets-only group no longer trips the NULL-detail bug; and deleting an empty group evicts the coordinator's cached state, verified by asserting a brand-new member joining under the same, just-deleted group name gets a fresh generation rather than continuing a resurrected one. - Mutation-tested the cache-eviction fix specifically: reinserting the stale wrapper instead of dropping it on the empty path reproduces the original resurrection bug, caught by the generation-number assertion on every backend. - `just clippy`, `cargo fmt --all --check`, `just doc` clean. --------- Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
## What
Refreshes the checked-in Understand-Anything graph (`.ua/`) from
`b3d0ceaf` (Sep 10) to `da6d1160`, the current `main`.
| | Before | After |
|---|---|---|
| Nodes | 1,835 | 1,393 |
| Edges | 2,816 | 1,513 |
| Layers | 10 | 10 |
| Tour steps | 14 | 14 |
Only `.ua/knowledge-graph.json`, `.ua/fingerprints.json` and
`.ua/meta.json` change. No source files are touched.
## Why the node count dropped
The previous graph still held nodes for files that have since moved or
been deleted, mainly the old
`nisshi-storage/src/{dynostore,slate,pg,lite,limbo,null}.rs` modules and
the pre-split test files. These were removed (629 nodes, 1,051 edges).
Every file in the current scan has a node, and no node points at a
missing file.
The storage layer now reflects the crate split: `nisshi-storage`,
`nisshi-storage-sql` (DDL and query files), `nisshi-storage-slatedb` and
`nisshi-storage-null`.
## Confidence
- 386 changed files were re-analyzed; the rest of the graph carries over
unchanged.
- Inline validation found no issues: unique IDs, no dangling edges,
every file-level node in exactly one layer.
- Several batches wrote function and class summaries from extracted
structure and names without reading each source body. Treat those
summaries as lower confidence than the file-level ones.
- Eight `tested_by` edges were dropped by the merge step and not
restored.
- `.ua/` files are excluded from the graph so it does not describe its
own output.
## Review
Generated data, so a skim is enough. The diff in `fingerprints.json` is
large because the baseline is regenerated for 502 files.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com>
## What Adds `tier-c.yml`, an optional nightly workflow. It does not gate anything: it is not in ci-gate's `needs` and not a required check. - Triggers: a daily schedule and `workflow_dispatch`. Every job runs on `main` only. - Compat suites (librdkafka, franz-go) on the S3 engine, with the report jobs. A failure turns the run red. - `test` on Postgres 16, 17 and 19 beta (`postgres:19beta4`, the newest 19 tag on Docker Hub). The beta leg uses `continue-on-error`. - The Rust cache is restore-only, so a nightly run never overwrites the cache that `ci.yml`'s `test` job saves on main. - An `alert` job (the only one with `issues: write`) opens or updates one `nightly-failure` issue when any job does not succeed, and closes it on the next green run. No new secrets. ## Why Some checks are too slow or too unstable to gate a merge. Tier C runs them on a schedule and reports through one issue. This PR is additive and touches no existing file. Until a later PR moves the S3 compat legs and the extra Postgres versions out of Tier B, Tier C duplicates that coverage. ## Review notes An independent review found these, and the second commit fixes them: - A cancelled or timed-out job did not raise an alert. Now any non-success result does. - The concurrency group is now scoped by ref, so a manual run from another ref cannot replace a waiting scheduled run. - The cache comment now states that the job id and the `RUST*`/`CARGO*` env values must match `ci.yml`'s `test` job. Open follow-ups, not in this PR: the issue notifies only people who watch the repo or subscribe, and a failing beta leg leaves the run green with no signal besides the red leg. ## Testing `actionlint` (no output, exit 0) and `zizmor --offline` on the new file (`No findings to report`). `zizmor` at medium/medium over the whole workflows directory also reports no findings, and the `artipacked` check finds none. I ran `zizmor` offline, so the online audits did not run. The workflow itself has not run yet; after merge, dispatch it once from main and check the alert job, in particular that a failing beta leg does not count as a failure of `test`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
Adds `.claude/rules/naming.md`, rules for choosing names in the workspace, and links it from `CONTRIBUTING.md` next to the comment rules. The rules ask for names that a new contributor understands without reading the definition: name a value for what it means, a function for its result or effect, a test for the behaviour it requires, and a type for what it holds or does. They spell words out, avoid generic words such as `data` or `helper`, and use one name for one thing. Like `comments.md`, the file has `paths:` frontmatter, so Claude Code loads it only when it edits Rust, shell or justfile code. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: William Kourlas <156007774+solace-wkourlas@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Closes #886. Keeps secrets and record data out of logs, at every log level. ## Changes **Generated `Debug` hides secret fields.** `SENSITIVE_FIELDS` in `nisshi-sans-io/build.rs` lists fields by message, struct and field name. For a struct with a listed field, the generator leaves `Debug` out of the derive and writes a `Debug` impl that matches the derived output, except that the listed field reads `[hidden]`. This covers the public and the internal (mezzanine) structs, so logging a message, a `Body` or a `Frame` hides it. Field types, serde, `PartialEq` and `Hash` don't change. | Message | Struct | Field | |---|---|---| | `SaslAuthenticateRequest` | same | `AuthBytes` | | `SaslAuthenticateResponse` | same | `AuthBytes` | | `AlterUserScramCredentialsRequest` | `ScramCredentialUpsertion` | `SaltedPassword` | | `CreateDelegationTokenResponse` | same | `Hmac` | | `DescribeDelegationTokenResponse` | `DescribedDelegationToken` | `Hmac` | | `ExpireDelegationTokenRequest` | same | `Hmac` | | `RenewDelegationTokenRequest` | same | `Hmac` | | `AlterConfigsRequest` | `AlterableConfig` | `Value` | | `IncrementalAlterConfigsRequest` | `AlterableConfig` | `Value` | The build fails when an entry matches no field. It also fails on a `bytes` field that is in neither `SENSITIVE_FIELDS` nor `LOGGABLE_BYTES_FIELDS`, which lists the eight `bytes` fields that may be logged (the SCRAM salt, group protocol metadata and assignments, and client telemetry). A descriptor update that adds a `bytes` field therefore needs a decision before it builds. The build also fails if a listed field is tagged, because the internal struct keeps a tagged field's bytes in `tag_buffer`, which the generated `Debug` can't hide. **Raw bytes are logged by length.** Frame dumps in `nisshi-service`, `nisshi-client` and `Frame::request`/`Frame::response`, and the `serialize_bytes` spans (which recorded every bytes field at INFO), now log a length. The decoder no longer logs each `u8` it reads, which wrote a compressed record's key and value one byte per line, and the snappy inflator no longer logs its input. String fields are logged by length when they're encoded or decoded, because a string can hold a config value. When the broker decodes a request it logs the api key, version and correlation id read from the frame header, so a frame that fails to decode can still be identified. **Records are logged by length.** `deflated::Batch` writes `record_data_len` in place of `record_data`, so a logged produce or fetch message doesn't carry record contents. `Record` writes `key_len` and `value_len`, and a record `Header` writes its key and `value_len`. Record decoding, the dynostore batch decode and the SQL record reads log lengths. A failing test that compares records now shows lengths for keys and values; the other fields still show. **Storage.** `ScramCredential`'s `Debug`, and slatedb's stored form of it, show `salt` and `iterations` and hide `stored_key` and `server_key`. The libSQL and Turso query helpers log the SQL without its parameters, which hold record data and credentials. The SQL produce paths log key and value lengths when an insert fails, and the libSQL compaction spans skip the record key. The null backend's SCRAM upsert span skips the credential. **Schemas and lake tables.** Schema validation and Avro, JSON and protobuf conversion log the field, the schema and the kind of value, not the value. The error variants that held a value (`AvroToJson`, `InvalidValue`, `JsonToAvro`, `JsonToAvroFieldNotFound`, `UnsupportedSchemaRuntimeValue`) now hold its kind, and the `todo!`/`unimplemented!` messages for unsupported Avro types name the kind. An Avro record that fails to read or write against its schema becomes the new `Error::AvroRecord`, which holds no detail, because the `apache_avro` error for a record writes the record's values into its message. These are breaking changes to `nisshi_schema::Error`, listed in the CHANGELOG. ## Comparison with Kafka 3.9.1 Logs aren't visible to clients, so this changes nothing a client can observe. - **SASL auth bytes:** matches. Kafka handles SASL frames in the authenticator and never passes them to the request logger ([SaslServerAuthenticator.java#L425-L504](https://github.com/apache/kafka/blob/3.9.1/clients/src/main/java/org/apache/kafka/common/security/authenticator/SaslServerAuthenticator.java#L425-L504)). - **`SaltedPassword` and `Hmac`:** stricter than Kafka. Kafka's request logger redacts only config values ([RequestChannel.scala#L186-L212](https://github.com/apache/kafka/blob/3.9.1/core/src/main/scala/kafka/network/RequestChannel.scala#L186-L212)), so it logs these. We follow Kafka's own `DelegationToken.toString`, which writes `hmac=[*******]` ([DelegationToken.java#L74-L79](https://github.com/apache/kafka/blob/3.9.1/clients/src/main/java/org/apache/kafka/common/security/token/delegation/DelegationToken.java#L74-L79)). - **Config values:** stricter than Kafka. Kafka's request logger hides an `AlterConfigs` or `IncrementalAlterConfigs` value only when the config it names is sensitive ([RequestChannel.scala#L186-L212](https://github.com/apache/kafka/blob/3.9.1/core/src/main/scala/kafka/network/RequestChannel.scala#L186-L212)). We keep no list of sensitive config names, so we hide every value, topic configs included. - **Raw frames:** matches. Kafka logs request and response sizes, not their bytes ([RequestChannel.scala#L436-L450](https://github.com/apache/kafka/blob/3.9.1/core/src/main/scala/kafka/network/RequestChannel.scala#L436-L450)). - **Records:** matches. Kafka's generated JSON converters write a records field as its size ([JsonConverterGenerator.java#L404-L420](https://github.com/apache/kafka/blob/3.9.1/generator/src/main/java/org/apache/kafka/message/JsonConverterGenerator.java#L404-L420)). - **Placeholder:** `[hidden]`, as Kafka's `Password.HIDDEN` ([Password.java#L24](https://github.com/apache/kafka/blob/3.9.1/clients/src/main/java/org/apache/kafka/common/config/types/Password.java#L24)). ## Tests - `nisshi-sans-io/tests/it/redact.rs`: - For each hidden field, formats the message, its `Body` and a `Frame` with `{:?}`, and checks that a marker secret is absent as text, as a decimal byte list, as hex and as the decoder's one-line-per-byte form, that `[hidden]` is present, and that the other fields still show. - Round-trips `SaslAuthenticate` (request and response, every version), `AlterUserScramCredentials`, `AlterConfigs`, `IncrementalAlterConfigs` and a produce request through encode and decode under a TRACE subscriber with span events on, and checks that the captured log doesn't hold the marker. The SASL request test also checks that the `serialize_bytes` span ran, so it can't pass by capturing nothing; with the old span restored it fails. - Inflates a batch with each compression type under the same subscriber, and checks the `Debug` of a record and of an inflated and a deflated batch. - `nisshi-sans-io` unit test: the internal (mezzanine) `Debug` hides each listed field. - `nisshi-broker/tests/it/log_redaction.rs`: against the libSQL, SlateDB and PostgreSQL backends, creates a SCRAM user with `AlterUserScramCredentials`, logs in with SCRAM-SHA-256, creates a topic, produces an uncompressed and a gzip batch and fetches them back, all under a global TRACE subscriber. The in-memory backend has no leg, because it keeps no SCRAM credentials. It checks that the log holds none of the password, the salted password, the SASL messages, or the record key, value and header value. - `nisshi-storage`: `ScramCredential`'s `Debug` hides both keys. - Misspelling a `SENSITIVE_FIELDS` entry fails the build with: `SENSITIVE_FIELDS: no field SaltedPasswd in struct ScramCredentialUpsertion of message AlterUserScramCredentialsRequest; ...` ## Overlap #849 also touches `nisshi-service/src/frame.rs`. This PR changes the `debug!` lines around the request decode there and leaves the `debug!(?request)` line that #849 rewrites, since the generated `Debug` covers it. The branch is rebased on `main` after #850. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
### What is the purpose of this change? `build-storage` (5 legs) and `build-storage-lake` (3 legs) each run a cold debug build with the cache off. They exist to catch a feature combination that doesn't compile, which `cargo check` catches without codegen or linking. ### How was this change implemented? Both jobs run `cargo check --bin nisshi --no-default-features` with the same features `just build` enabled before, so the feature sets don't change. Calling cargo directly means neither job installs `just` any more. Job names are unchanged, so `ci-gate`, `release` and the Tier B list are untouched. On a private copy of the repo, before #895, a lake leg dropped from 10 to 18 minutes to 6 to 9, and a storage leg from 5 to 7 minutes to 4 to 5. A throwaway commit with a type error in a feature-gated module failed the lake legs as expected. ### Is there anything the reviewers should focus on/be aware of? `cargo check` doesn't link, so a failure that only shows up at link time would no longer fail these legs. `test`, `release` and `smoke` still link the all-features binary. The cache stays off: the repo already holds about 11.9 GB of Actions cache, and 8 new entries would evict the release, test and clippy caches. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…loses the connection (#845) ## Bug A Kafka producer configured with `acks=0` never reads a response. Nisshi's broker wrote one anyway for every Produce request. On success this served no purpose; on failure, the broker silently dropped the batch instead of signaling anything. Worse, the unsolicited success response sits on the connection and corrupts correlation-id matching on the client's next request, since a real acks=0 client never consumes it. ## Fix `ProduceService` (for `acks=0` requests only): - On success: marks a new `SuppressResponseExtension`, a self-resetting `AtomicBool`-backed marker, so `TcpBytesService::req` skips writing the already-assembled response. A plain bool can't be inserted once and reused because rama's `Extensions` is shared (not forked) across every request on a connection and has no way to remove an entry once inserted. - On any partition error anywhere in the assembled response, matching real Kafka's own `handleProduceRequest` (which closes the connection whenever `unauthorizedTopicResponses`/`nonExistingTopicResponses`/`invalidRequestResponses`/per-partition errors merge into a non-empty error set): returns a new `Error::AcksZeroProduceFailed`, reusing the existing error-closes-connection path instead of embedding the error in a response nobody reads. Both decisions (broad scope matching every error path, not just storage; a dedicated error variant) were checked against real Kafka source and independently mutation-tested: disabling the response suppression makes all 4 success tests fail with a stray Produce response's correlation id showing up where it shouldn't, and narrowing the error scope back to storage-only makes all 4 failure tests fail with a timeout since the connection incorrectly stays open. Both were restored to green afterward. `broker.rs` logs the new failure variant at `info`, matching real Kafka's own "Closing connection due to error during produce request..." level, instead of falling through to the generic `error!` path. A pre-existing deviation was also checked in this review: `TcpBytesService::process`'s `error!` on a connection failure was deliberately left alone rather than downgraded to `debug!`, because `TcpListenerService` (the proxy's own accept loop) only logs connection failures at `debug!` already -- downgrading the broker side too would remove the proxy operator's only error-level signal for this class of failure. ## Proxy batch-forwarding fix (included here) Audited every existing `acks=0` call site in the repo. `nisshi-generator` and `nisshi-perf` both built a default (`acks=0`) `ProduceRequest` and awaited a response that would now never arrive from a nisshi broker, so both now request `Ack::Leader` explicitly. One more call site needed the same fix: `nisshi-proxy/src/produce.rs`'s `produce_request()`, used by the batch-forwarding path (`tansu.batch=true`), also built a default `ProduceRequest` with no `acks(...)`. `send_pending_batch` awaits the origin's response to fan out base offsets to the waiting downstream clients, so forwarding with `acks=0` against an origin that correctly suppresses its response (this PR's fix) would hang that await forever. Fixed by requesting `Ack::Leader` explicitly, same pattern as generator/perf -- the batch path needs the origin's real response regardless of what kind of origin is on the other end. ## Deferred: proxy pass-through case The proxy's plain pass-through forwarding (not the batching path above) goes through `nisshi-client`'s `BytesConnectionService` (reached via `ConnectionManager`), which does `write_all` then an unconditional `read_exact` for every forwarded request. Forwarding a pass-through `acks=0` Produce to an origin that correctly never responds (a nisshi origin after this fix) will hang on that `read_exact`. This is a known, deliberately-deferred consequence of this fix for the pass-through path specifically, not something newly discovered and unrelated. Fixing it properly means adding acks-awareness to `BytesConnectionService` itself, a separate, larger piece of work, tracked as a separate follow-up (description corrected to name the right component and framing). ## Tests - `nisshi-broker/tests/it/produce_acks_zero.rs`: real-socket tests across all 4 storage engines (in-memory, libSQL, Postgres, SlateDB). - `success_is_not_written`: sends an `acks=0` Produce then a Metadata request, asserts the first (and only) frame read back carries the Metadata correlation id, not the Produce one -- positive proof nothing was interposed -- then confirms via `ListOffsets(Latest)` the record was still stored. - `failure_closes_connection`: a record-less batch, rejected by `ProduceService::partition`'s pre-storage `rejection()` validation, asserts the connection closes. - `storage_failure_closes_connection` (new in this round): an idempotent producer id never registered via `InitProducerId`, exercising a genuine `storage.produce()`-layer failure (`UnknownProducerId`) rather than pre-storage validation, across all 4 engines. - Bumped 5 `produce.rs` tests that relied on an `acks=0` response embedding an error code (their actual subject is idempotent-producer/batch-rejection error codes, not acks semantics) to a non-zero acks so they keep testing that. - `nisshi-proxy`'s batch tests updated to expect `acks=1` (`Ack::Leader`) on the forwarded request. Verified against a real client: librdkafka's compat suite's `0008_reqacks` (acks=-1/0/1) passes, along with the rest of the allowlist bar one pre-existing failure, `0060_op_prio` (reproduced identically against an unmodified `main` build) -- a legacy-consumer test that does call produce, but fails during its later consume phase, not in the produce path itself. Local run this round: ``` cargo nextest run -p nisshi-broker --all-features -E 'test(/^produce_acks_zero::/)' # 12 passed cargo nextest run -p nisshi-broker --all-features -E 'test(/^produce::/)' # 28 passed cargo nextest run -p nisshi-proxy --all-features # 10 passed just fmt / just clippy / just build-all # clean ``` ## CHANGELOG > The broker no longer writes a Produce response for an `acks=0` request. A storage or validation failure on an `acks=0` Produce now closes the connection (matching Apache Kafka) instead of silently dropping the batch. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01D2qdgPhMyLGZN5dLR8CVHs --------- Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Closes #909 crates.io requires a `description`; the dynostore, null, slatedb and sql storage crates had none, so `cargo publish` would fail. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…HANGELOG.md (#898) ## What changes, and why? Closes #897. Every PR edits the `[Unreleased]` section of `CHANGELOG.md`, so open PRs conflict with each other and rebase just for the changelog. With this change: - PRs don't edit `CHANGELOG.md`. The advisory `changelog` job and `.github/scripts/changelog-check.sh` are removed, and the file's header says it is generated from the commit history at release time. A follow-up adds that generation. - A new `pr-title` workflow requires a [Conventional Commits](https://www.conventionalcommits.org/en/v1.0.0/) PR title: `type(scope): summary`, with an optional scope and `!` for a breaking change. It is a workflow of its own, not a `ci.yml` job, because it must rerun on `edited`, and `edited` in `ci.yml` would rerun all of Tier A on every title or description edit. It skips its job on `merge_group`, so it can become a required check without blocking the queue. - Dependabot titles its PRs `build(deps): ...`. - The PR template has three short headings, because the description becomes the commit body. Its checklist moves to `CONTRIBUTING.md`, which also documents the title format, squash merging, and why each PR commit still needs `Signed-off-by`. ## Upgrade impact None for operators. For contributors: from the merge of this PR, don't add `CHANGELOG.md` entries, and use a Conventional Commits PR title. A maintainer switches the repository and merge queue to squash merging and makes `pr-title` a required check after this merges. ## How was this tested? - `actionlint`, `typos`, and `zizmor --min-confidence medium --min-severity medium` (online and offline) pass, and the offline `artipacked` run reports nothing. - On this PR: retitle to a non-conventional title and `pr-title` fails; restore the title and it passes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Samuel Gamelin <104787241+sgamelin@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
## What changes, and why? Bumps the workspace version from `0.7.0-pre.2` to `0.7.0-pre.3`, for the next pre-release tag. `0.7.0-pre.2` is already on crates.io, so the release workflow can't publish the workspace again until the version changes. `Cargo.lock` changes only the workspace crates' own versions. Merge after #910, so that the tagged commit has the storage crates' descriptions. Part of #911. ## Upgrade impact None. ## How was this tested? `cargo metadata --locked` passes, so `Cargo.lock` matches the new version. The merge queue's `cargo-publish-dry-run` job packages and builds every crate at the new version. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Samuel Gamelin <samuel.gamelin@solace.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…#887) ## What does this change do, and why? InitProducerId v0-2 and the v3+ KIP-360 epoch bump now work on every storage engine, following Kafka 3.9.1. Before, only Postgres handled both: libSQL and Turso panicked on a `todo!()`, the object store engines answered UNKNOWN_SERVER_ERROR, and SlateDB answered UNKNOWN_SERVER_ERROR without a transactional ID and didn't check the claim with one. Closes #878. The shared rules are in [`nisshi-storage/src/producer.rs`](https://github.com/nisshi-io/nisshi/blob/a50eedb955c82782a2386c1fdb05cae380f1f09d/nisshi-storage/src/producer.rs#L22-L64): | Transactional ID | Stored producer | Request | Answer | |---|---|---|---| | none | – | fresh or claim | new producer ID, epoch 0 | | set | none | fresh or claim | create, epoch 0 | | set | `(id, e)` | fresh | bump to `e+1` | | set | `(id, e)` | claim `(id, e)` | bump to `e+1` | | set | `(id, e)` | claim with another ID or epoch | PRODUCER_FENCED, id and epoch -1 | | any | – | only one of ID and epoch is -1 | INVALID_REQUEST, id and epoch -1 | v0-2 decode as `(None, None)`. Kafka gives those fields their default of -1, so they're a fresh request. `InitProducerIdService` applies the shape check once, and each engine repeats it so a direct `Storage` caller can't panic. A failed answer is logged at info with what the producer sent, because the response no longer echoes it. Kafka 3.9.1 references: the shape check in [KafkaApis.scala](https://github.com/apache/kafka/blob/3.9.1/core/src/main/scala/kafka/server/KafkaApis.scala#L2343-L2347), a null transactional ID ignoring the claim in [TransactionCoordinator.scala](https://github.com/apache/kafka/blob/3.9.1/core/src/main/scala/kafka/coordinator/transaction/TransactionCoordinator.scala#L114-L122), the producer ID check in [TransactionCoordinator.scala](https://github.com/apache/kafka/blob/3.9.1/core/src/main/scala/kafka/coordinator/transaction/TransactionCoordinator.scala#L212-L232), and the epoch check in [TransactionMetadata.scala](https://github.com/apache/kafka/blob/3.9.1/core/src/main/scala/kafka/coordinator/transaction/TransactionMetadata.scala#L275-L315). **Postgres changes beyond the shared rules** - The claim is checked inside the transaction that bumps the epoch, against the producer's newest epoch ([pg.rs](https://github.com/nisshi-io/nisshi/blob/a50eedb955c82782a2386c1fdb05cae380f1f09d/nisshi-storage-sql/src/pg.rs#L1520-L1548)). That includes the epoch the timeout sweep adds when it fences a producer, so a fenced producer can't claim its old epoch back. Kafka fences this case too: the fence sets the last epoch to -1. - The `txn` row is locked first ([`txn_select_name_for_update.sql`](https://github.com/nisshi-io/nisshi/blob/a50eedb955c82782a2386c1fdb05cae380f1f09d/nisshi-storage-sql/src/sql/txn_select_name_for_update.sql)). At read committed, two concurrent claims could both read the same epoch: one got the bump, and another either got the next epoch or failed on the `(producer, epoch)` unique key. - A claim for an unknown producer ID no longer answers UNKNOWN_PRODUCER_ID; it follows the table. **Deliberate differences from Kafka**, commented at [`check_claim`](https://github.com/nisshi-io/nisshi/blob/a50eedb955c82782a2386c1fdb05cae380f1f09d/nisshi-storage/src/producer.rs#L47-L64) and in the CHANGELOG: - Kafka accepts a claim of the previous epoch, from a producer retrying a bump whose response it lost. No engine stores the previous epoch, so that claim is fenced and the producer initialises again. Guessing `current - 1` would let a fenced producer take over after another producer's fresh init. - Kafka answers INVALID_PRODUCER_EPOCH instead of PRODUCER_FENCED below v4. The typed services don't see the API version yet, so that mapping is left for a separate change. **Not changed:** an init during an ongoing transaction still aborts it and bumps the epoch (Kafka aborts and answers CONCURRENT_TRANSACTIONS). Epoch exhaustion and the empty transactional ID are unchanged. ## How was this tested? - New broker tests in `nisshi-broker/tests/it/init_producer_id.rs`, on in_memory, lite, slatedb and pg (40 tests): - each version v0 to v5, encoded and decoded through `Frame`, with and without a transactional ID; - a claim without a transactional ID; a claim of the current epoch; a stale, newer and other-producer claim; a claim for an unknown transactional ID; mixed shapes through the service and through `Storage`; - a claim during an ongoing transaction aborts it (abort marker, last stable offset released) and bumps; - a fenced claim leaves the current producer's ongoing transaction open, and that producer still commits it. - New Postgres unit tests: a claim after the timeout sweep is fenced, and 8 concurrent claims over 10 rounds bump the epoch once per round. Each fails without its fix. - On `main`, 24 of the 32 shape tests fail: libSQL panics at the `todo!()`, the object store engines and SlateDB answer UNKNOWN_SERVER_ERROR, and Postgres applies the old rules. - The same tests are in a `turso` module, `#[ignore]`d because the Turso engine doesn't start in the broker tests yet (#866). Its change mirrors libSQL's. - `just fmt`, `just build-all`, `just clippy`, `just doc` pass. A full `just test` run passed except three `nisshi-schema::berg` tests that failed connecting to the local lakehouse catalog. After the last Postgres changes, the storage crates and the broker transaction, produce and InitProducerId tests pass (231 tests). ## Checklist - [x] Commits are signed off (`git commit -s`) — see [CONTRIBUTING.md](https://github.com/nisshi-io/nisshi/blob/main/CONTRIBUTING.md#sign-off-your-commits-dco) - [x] `just fmt`, `just clippy`, and `just test` pass locally - [x] Tests added or updated for the behavior changed - [x] Docs updated if user-facing behavior changed (CHANGELOG) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Samuel Gamelin <samuel.gamelin@solace.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…onses (#885) ## What does this change do, and why? `nisshi proxy` and the CLI tools send every request through the `nisshi-client` connection pool. When one broker stopped answering, that pool could hold up every request that went through it (#882): - **Dead connections:** a connection went back into use after the broker closed it, or after a request on it failed or was cancelled. - **Unbounded pool wait:** `pool.get()` waited forever for a free connection. - **No response deadline:** each response was read without a deadline, while the request held its connection. This PR: - **`recycle`** discards a connection whose last request did not complete, that the broker closed, that has bytes nobody read, or that was idle for longer than `max_idle` (default 9 minutes, the Java client's `connections.max.idle.ms`). - **Response deadline:** each request gets `request_timeout` (default 30s), or the wait it asks the broker for plus 5s when that is longer. That covers Fetch, Produce, JoinGroup, and the topic and partition admin requests. `max_request_wait` (default 5 minutes) caps the wait, because every client of a proxy shares its pool. This is the rule the Java client uses for JoinGroup ([`AbstractCoordinator`](https://github.com/apache/kafka/blob/3.9.1/clients/src/main/java/org/apache/kafka/clients/consumer/internals/AbstractCoordinator.java#L620-L628)). A request that misses its deadline fails with the new `Error::Timeout`. - **Pool timeouts:** a wait timeout (default 30s) and a create timeout (`connect_timeout`, default 30s). Each connect attempt stops after 10s, Kafka's `socket.connection.setup.timeout.ms`. - **Proxy:** `nisshi proxy` opens up to 256 connections to its origin, instead of twice the number of CPUs, because the origin holds a fetch or a JoinGroup on its connection until it answers. - **Settings:** all of these are on `nisshi_client::Builder`. No proxy command-line flags. - **Metrics:** new counters `request_timeouts`, `pool_get_errors` (by `error`) and `pool_connections_discarded` (by `reason`). Follow-ups, not in this PR: - The proxy does not cancel the origin request when its client disconnects. - Long-held requests (fetch, JoinGroup) share one pool with short ones. ## How was this tested? - New integration tests in `nisshi-client/tests/it/pool.rs`, against a fake broker: - a hung broker, through `Client` and through the frame layers that `nisshi proxy` uses - a connection the broker closed - a connection with unread bytes - a cancelled request whose late response arrives after the next request starts - a connection idle too long, and an idle connection that is reused - the pool wait timeout, and the configured wait and create timeouts - a fetch whose wait outlasts `request_timeout`, through both paths - Unit tests for the deadline rule. - With each fix removed in turn, its tests fail. - `cargo fmt --all --check` - `cargo clippy -p nisshi-client -p nisshi-proxy --all-features --all-targets -- -D warnings` - `cargo nextest run -p nisshi-client --all-features` - Before the review changes, `just build-all`, `just clippy` and `just doc` passed on the first version of this change. ## Checklist - [x] Commits are signed off (`git commit -s`) — see [CONTRIBUTING.md](https://github.com/nisshi-io/nisshi/blob/main/CONTRIBUTING.md#sign-off-your-commits-dco) - [x] `just fmt` and `just clippy` pass locally. `just test` ran for `nisshi-client` only. - [x] Tests added or updated for the behavior changed - [x] Docs updated if user-facing behavior changed (CHANGELOG) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Samuel Gamelin <samuel.gamelin@solace.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
## Summary - Follow-up to [#863](#863), split out per [Sam Gamelin's review comment](#863 (review)). - `is_address_family_unsupported` (used by the broker's IPv6→IPv4 listener fallback) only recognized the Unix errno `libc::EAFNOSUPPORT`. On Windows, socket errors come from `WSAGetLastError`, so `raw_os_error()` returns `WSAEAFNOSUPPORT` (10047) — a different number from the C runtime errno `libc::EAFNOSUPPORT` holds on Windows. A Windows host without IPv6 support would fail to bind `[::]` instead of falling back to `0.0.0.0`. - Adds a `#[cfg(unix)]`/`#[cfg(windows)]`-gated `address_family_unsupported_os_error()` helper, same style as #863, and reuses it from the existing `ipv4_fallback` unit tests so they exercise the right code on either platform. The Windows constant is now a named `WSAEAFNOSUPPORT`, not a bare `10047`. ## Test plan - [x] `cargo fmt --all --check`, `cargo test -p nisshi-broker --lib broker::` on Linux — all pass, no behavior change on Unix. - [x] Built and tested on **real Windows (MSVC)**, with #863's branch merged in locally (since `main` doesn't build on Windows without it yet): `cargo build -p nisshi-broker` and `cargo test -p nisshi-broker --lib broker::` both pass. Note on what this does and doesn't show, per @solace-aross's review: the unit tests construct a synthetic `io::Error` from `address_family_unsupported_os_error()`'s own value, so passing on Windows confirms the code builds and the `cfg` split/fallback branching is correct there — it does not independently confirm that a real Windows socket bind failure surfaces `raw_os_error() == 10047`. That part rests on `WSAEAFNOSUPPORT` (10047) being Microsoft's documented value for that error, not on a test exercising a real failing bind. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Bumps [nanoid](https://github.com/mrdimidium/nanoid) from 0.4.0 to 0.5.0. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/mrdimidium/nanoid/blob/main/CHANGELOG.md">nanoid's changelog</a>.</em></p> <blockquote> <h2>0.5.0</h2> <ul> <li>Bump <code>rand</code> to 0.9</li> <li>Add <code>rngs::thread_local</code> random source (<a href="https://redirect.github.com/mrdimidium/nanoid/issues/36">#36</a>)</li> <li><code>format</code> now accepts any <code>FnMut(usize) -> Vec<u8></code> random generator, enabling seeded and stateful RNGs (<a href="https://redirect.github.com/mrdimidium/nanoid/issues/32">#32</a>, <a href="https://redirect.github.com/mrdimidium/nanoid/issues/41">#41</a>). Non-capturing <code>fn(usize) -> Vec<u8></code> callers continue to work unchanged.</li> <li><code>nanoid!</code> macro size argument now accepts any expression, not only a single token (<a href="https://redirect.github.com/mrdimidium/nanoid/issues/28">#28</a>)</li> <li>Specialized fast path for alphabets whose size is a power of two (<a href="https://redirect.github.com/mrdimidium/nanoid/issues/35">#35</a>). Note: for seeded RNGs paired with a power-of-two alphabet (e.g. <code>SAFE</code>, the new <code>HEX_*</code> presets), the number of random bytes consumed per ID has changed — the output for a given seed will differ from 0.4.0.</li> <li>Add <code>alphabet::HEX_LOWERCASE</code> and <code>alphabet::HEX_UPPERCASE</code> presets (<a href="https://redirect.github.com/mrdimidium/nanoid/issues/39">#39</a>)</li> <li>Optional <code>smartstring</code> feature for small-string-optimized output (<a href="https://redirect.github.com/mrdimidium/nanoid/issues/29">#29</a>)</li> <li>Refreshed CI (GitHub Actions across OS matrix), drop Travis/AppVeyor</li> <li>Switched benchmarks to <code>criterion</code></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/mrdimidium/nanoid/commit/359c02d6f87260bd431e19374ccfca2890fdab1e"><code>359c02d</code></a> chore: 0.5.0 release</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/f0ad07fc16b96b4c00d76fa10853fa377ad8ee05"><code>f0ad07f</code></a> <a href="https://redirect.github.com/mrdimidium/nanoid/issues/39">#39</a>: Add hex alphabets</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/7f961f211fa6a1162bb0fd6bd5c6ff195005ec00"><code>7f961f2</code></a> Merge pull request <a href="https://redirect.github.com/mrdimidium/nanoid/issues/35">#35</a> from tmccombs/fast-impl</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/91a79fc9f11e463109dcba342d684ffdd5291862"><code>91a79fc</code></a> Update fast impl for actual format signature</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/ed800e971f01b589cad8d2c8b09973f8eacfd1a1"><code>ed800e9</code></a> feat: Use specialized implementation for alphabets with size 2^n</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/fef0b2eace7dbda294cca78fd5c8e96c188a00bc"><code>fef0b2e</code></a> Merge pull request <a href="https://redirect.github.com/mrdimidium/nanoid/issues/41">#41</a> from sidarth164/sid/fnmut</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/61e0606f9c558171238aa9caf29d02bf90fb7806"><code>61e0606</code></a> docs: update README and added an example</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/2004ff99bbe549e7323ac01bd43070aa1f11c33e"><code>2004ff9</code></a> feat: support passing mutable functions as random generators</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/3d405c51cd6932318cb5766d0a3adab0c0c704f5"><code>3d405c5</code></a> Fix ci for prs</li> <li><a href="https://github.com/mrdimidium/nanoid/commit/7011b102d07f7f2ba2f7c96d0b196a9a349dc9cd"><code>7011b10</code></a> Fixup readme, delete old example</li> <li>Additional commits viewable in <a href="https://github.com/mrdimidium/nanoid/compare/v0.4.0...v0.5.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
## What does this change do, and why? Adds `.github/dco.yml` with `allowOverrideAction: false`, which hides the dco2 app's "Set DCO to pass" button. Without it, anyone with write access can mark a PR as passing the DCO check from the check page, so a commit without a sign-off can still merge. dco2 reads this file from the default branch only, so it takes effect once merged. ## How was this tested? Config only. It takes effect after merge: the DCO check page on a PR with an unsigned commit should no longer offer the override button. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Reuben D'Souza <46090211+reubenjds@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…891) The published image is `FROM scratch` with a stripped Rust binary, so image scanners have no dependency list to read. This builds the release binaries with `cargo auditable build`, which embeds the crate list in the binary for Trivy, Grype and `cargo audit bin` to read. Trivy 0.69.2 on an x86_64 musl image built with this change reads 788 crates from the binary. On the current `ghcr.io/nisshi-io/nisshi:main` image it finds none. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…863) ## What changes, and why? The broker, `nisshi-generator` and `nisshi-perf` didn't compile on Windows. Each one imported `tokio::signal::unix` for SIGINT and SIGTERM handling, and that module exists only on Unix. The Unix code now builds only on Unix. On Windows, Ctrl+C (`tokio::signal::windows::ctrl_c()`) triggers the same graceful shutdown that SIGINT does, and so does closing the console window or ending the task in Task Manager (`ctrl_close()`), which stands in for SIGTERM. Windows ends the process 5 seconds after that event by default, whether or not shutdown has finished. Windows has no direct equivalent of SIGTERM for a console program, so a service manager's stop request doesn't trigger a graceful shutdown. Nothing changes on Unix. A new `build-windows` job in the nightly Tier C workflow runs `just build dev` on `windows-latest`, so a change that breaks the Windows build opens the `nightly-failure` issue. It runs nightly rather than in Tier B because a Windows build is slow and doesn't need to hold up the merge queue. ## Upgrade impact None. ## How was this tested? - `just build dev` (all features) for `--bin nisshi` on Windows with the MSVC toolchain. - On Windows, started the broker with `--storage-engine memory://nisshi/`, created a topic with `nisshi topic create`, produced a JSON message with `nisshi cat produce` and read it back with `nisshi cat consume`. - `cargo check` for `nisshi-broker`, `nisshi-generator` and `nisshi-perf`, and `cargo fmt --all --check`, on Linux after rebasing onto `main`. - The new Tier C job only runs on `main`, so it first runs in the nightly run after this merges. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: William Kourlas <156007774+solace-wkourlas@users.noreply.github.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
## What does this change do, and why? Adds `nisshi-smoke-test`, a suite that runs the real Kafka command-line tools against a broker on one storage engine, and makes the `smoke` CI job run it in place of the bats tests from a pinned `example-java` checkout. - `just smoke <postgres|sqlite|memory|s3>` starts the services the engine needs, a shared broker, and a container with the Kafka CLI tools, runs the tests, and removes everything it started. Each engine uses its own compose project, so runs on different engines don't interfere. - Tests use the shared broker or launch their own (`Broker::isolated()`), and use their own topic and group names, so they run in parallel. Every Kafka CLI call and blocking `docker` command has a time limit, so a hang fails the test that caused it and names the command. - A run fails if a test fails, or if the shared broker exited, restarted, panicked or didn't exit with 0 on SIGTERM. It writes a `name,PASS|FAIL|SKIP` row per test, the format the compat suites' report reads. - The first tests cover a topic's life: create, list, describe, produce with keys and headers, read offsets, consume through a group, describe the group's committed offsets, and delete. - `describe_topic_configs` is `#[ignore]`d: every storage backend labels a topic's own configs as defaults in its DescribeConfigs response, so `kafka-topics --describe` hides them. CI skips ignored tests; local runs still run them. - CI: Kafka 3.9.2 and 4.3.1 on postgres, sqlite, memory and s3, on x86 and arm. The sqlite legs build the broker from source until #796 is fixed. A new `smoke-report` job puts every leg's results in one grid in the step summary. - `just test` excludes the crate, since it needs Docker. ## How was this tested? - `just smoke memory` and `just smoke postgres` with Kafka 3.9.2, `just smoke memory` with Kafka 4.3.1, and `just smoke s3` against the `main` image, with `CI=true` and without (ignored test skipped in CI mode, run locally). - `cargo fmt --all --check`, and clippy (`-D warnings`) and rustdoc (`-D warnings`, private items) on `nisshi-smoke-test`. - `just smoke sqlite`, and the workspace-wide `just clippy` and `just test`. ## Checklist - [x] Commits are signed off (`git commit -s`) — see [CONTRIBUTING.md](https://github.com/nisshi-io/nisshi/blob/main/CONTRIBUTING.md#sign-off-your-commits-dco) - [x] `just fmt`, `just clippy`, and `just test` pass locally - [x] Tests added or updated for the behavior changed - [x] Docs updated if user-facing behavior changed 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: William Kourlas <156007774+solace-wkourlas@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…uce, consume, offsets and restarts (#902) Stacked on #862: review that one first. This PR adds the next set of smoke tests on top of its harness. ## What this adds Tests that run the Kafka CLI tools the way a user does, on every storage engine: - `broker`: ApiVersions names node 111, and the cluster id is the one the broker started with. - `topics`: describe, auto-create on produce, delete and re-create. - `configs`: describe, add and delete topic configs, and broker defaults. - `produce`: every `acks` setting, compression codecs, tombstones, `max.message.bytes`, idempotent and concurrent producers, `LogAppendTime`. - `consume`: reading from an offset, a group resuming where it stopped, a pattern subscription that picks up a topic created later. - `offsets`: lookups by time. - `restart`: topics, records, committed offsets, topic configs and SCRAM users survive a SIGTERM restart on PostgreSQL and SQLite, and an in-memory broker starts again empty. - `storage_url`: an unparsable or invalid storage URL stops the broker with an error that names it. The harness now has one module per Kafka tool. Each command whose output a test reads has its own `Output<Marker>` type, so a reader can only be called on its own command's output. Brokers can restart on the same storage, and tests can start isolated brokers. The CI change gives a smoke leg that fails before `just smoke` runs (setup, toolchain) a FAIL row, so it still shows in the report. `run.sh` also fails the leg if no nextest status line parses, instead of showing a green leg with no tests. `.claude/rules/smoke-tests.md` sets the rules for writing a smoke test. ## Ignored tests A test that fails because of a broker bug is ignored with the bug in its reason, so CI stays green and the fix enables it. Every reason names the issue or PR that fixes it, e.g. #798 for the five produce tests that read the latest offset on in-memory storage, and #849 for the `max-timestamp` lookup. ## Testing - `just clippy`-equivalent for `nisshi-smoke-test` (all features, and each engine feature alone), `cargo fmt`, the crate's unit tests, rustdoc with warnings denied, ShellCheck on `run.sh` and actionlint on `ci.yml` all pass. - `just smoke postgres`, `sqlite` and `memory` on this tree, rebased on main: every test that isn't ignored passes, and the shared broker passes its checks. Local runs also run the ignored tests: postgres 55 of 64 pass, sqlite 58 of 67, memory 45 of 58, and every failure is an ignored test. - The `maintenance_interval`, timestamp-lookup and `acks=0` tests were ignored for bugs that #850, #838 and #845 fixed, and are now enabled. `acks=0` stays ignored on in-memory storage for #798, like the other produce tests. - `s3` is unchanged: it still runs only `topic_lifecycle`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: William Kourlas <156007774+solace-wkourlas@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
## What changes, and why? Bumps `opentelemetry`, `opentelemetry_sdk`, `opentelemetry-otlp` and `opentelemetry-stdout` to 0.33.0 in the workspace `Cargo.toml`. The four crates are co-versioned, so the separate Dependabot bumps (#914, #915) cannot compile alone. This PR replaces them. The lock no longer carries `opentelemetry` 0.32, because `tracing-opentelemetry` and `opentelemetry-semantic-conventions` were already on 0.33. The lock also re-points ten crates to the already-locked windows-sys 0.61.2 and windows-link 0.2.1. No crate versions were added or removed; this affects Windows targets only. No source changes were needed; the workspace builds against 0.33 as is. ## Upgrade impact opentelemetry-otlp 0.33 turns OTLP retries on by default: an export that gets 429, 502, 503 or 504 is retried up to 3 times with exponential backoff (100 ms initial, 1.6 s max, Retry-After honoured on 503). Transport errors are not retried, so a down collector behaves as before; an overloaded collector can hold an export cycle and shutdown flush a few seconds longer. OTLP/HTTP bodies are now capped at 64 MiB; our metric payloads are far below that. No configuration or env var change for operators. ## How was this tested? - `cargo check --workspace --all-features --all-targets`: clean - `just clippy`: clean - `just doc`: clean - `cargo nextest run -p nisshi-otel -p nisshi-perf -p nisshi-generator -p nisshi-proxy --all-features`: 27 passed - Broker integration tests that need postgres and minio run in CI, not locally. - CI does not run the broker against an OTLP collector; the exporter change is compile-verified only. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Andrea Ross <168456375+solace-aross@users.noreply.github.com> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
## What changes, and why? A PR that only touches docs or CI config that `ci.yml`, `codeql.yml` and `dependencies.yml` don't read still runs the full Rust suite. A new `changes` job in each of those workflows lists the PR's files and classifies them with `.github/scripts/ci-changes.sh`. When every path is on its skip list, the Rust jobs skip and the required checks still report: - `ci.yml`: every job except `typos` and `ci-gate` gets `needs.changes.outputs.rust == 'true'`. `ci-gate` needs `changes`, so a failed classification fails the gate, and its "Tier A ran on pull_request" step only applies when the Rust jobs ran. - `codeql.yml`: the language matrix becomes `analyze-actions` and `analyze-rust`, keeping the `Analyze (actions)` / `Analyze (rust)` check names. Only the rust job waits on `changes`. - `codeql.yml`'s `Analyze (rust)` and `dependencies.yml`'s `cargo-deny` run unless `changes` said `false`, so a failed `changes` job can't turn a required check into a passing skip. The skip list is `.github/*` (except these three workflows and the scripts they run), `docs/*.md`, top-level `*.md`, `LICENSE` and `NOTICE`. Any other path, or an empty or truncated file list, runs everything. The `changes` jobs check out the base branch's copy of the script, so a PR can't widen its own skip list. `workflow-lint.yml` runs the PR's copy of `ci-changes.test.sh`. Push, `merge_group` and scheduled runs always run everything. ## Upgrade impact None for operators. Contributors see the Rust jobs as skipped on docs-only and CI-config-only PRs. ## How was this tested? - `.github/scripts/ci-changes.test.sh` passes (14 cases). - `actionlint` and `zizmor --min-confidence medium --min-severity medium` (offline) pass. - On this PR the `changes` jobs find no `ci-changes.sh` on the base, so they run everything. The skip path first takes effect on the PR after this one. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Reuben D'Souza <46090211+reubenjds@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Bumps [slatedb](https://github.com/slatedb/slatedb) from 0.14.1 to 0.17.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/slatedb/slatedb/releases">slatedb's releases</a>.</em></p> <blockquote> <h2>v0.17.0</h2> <h2>What's Changed</h2> <ul> <li>Scale SsTableView::estimate_size by the visible_range fraction by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2043">slatedb/slatedb#2043</a></li> <li>RFC-0033: Rewrite CachedObjectStore by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2031">slatedb/slatedb#2031</a></li> <li>docs: renumber local object mirroring RFC to 0034 by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2062">slatedb/slatedb#2062</a></li> <li>refactor(sst): remove WAL-specific table handling by <a href="https://github.com/rodesai"><code>@rodesai</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2058">slatedb/slatedb#2058</a></li> <li>feat(compactor): configure checkpoint lifetime by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2063">slatedb/slatedb#2063</a></li> <li>chore(foyer): bump to foyer v0.22.4 by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2066">slatedb/slatedb#2066</a></li> <li>Add Triplox to the list of adopters by <a href="https://github.com/FiV0"><code>@FiV0</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2075">slatedb/slatedb#2075</a></li> <li>feat(uniffi): expose with_db_cache and with_db_cache_disabled on DbReaderBuilder by <a href="https://github.com/flexorRegev"><code>@flexorRegev</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2070">slatedb/slatedb#2070</a></li> <li>add tracing spans slatedb.read and slatedb.read.memtable by <a href="https://github.com/cadonna"><code>@cadonna</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2048">slatedb/slatedb#2048</a></li> <li>Add per-Db block-cache evacuation to disk for shared Foyer hybrid caches by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2034">slatedb/slatedb#2034</a></li> <li>Add repository-wide Simple English guidance by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2077">slatedb/slatedb#2077</a></li> <li>Add ObjectStoreBuilder in Uniffi Binding by <a href="https://github.com/LGouellec"><code>@LGouellec</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2068">slatedb/slatedb#2068</a></li> <li>Add AI disclosure and enum guidance by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2078">slatedb/slatedb#2078</a></li> <li>Probe point-lookup sources concurrently instead of one at a time by <a href="https://github.com/hawkaa"><code>@hawkaa</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2039">slatedb/slatedb#2039</a></li> <li>feat: tag object store calls with segments by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2074">slatedb/slatedb#2074</a></li> <li>feat(manifest): enforce the L0 ULID cutoff invariant by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2079">slatedb/slatedb#2079</a></li> <li>[2021] Add DbBuilder::with_write_runtime to place the batch-writer on the caller's runtime by <a href="https://github.com/1996fanrui"><code>@1996fanrui</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2082">slatedb/slatedb#2082</a></li> <li>Enable filters in range scans by <a href="https://github.com/yiming-fang"><code>@yiming-fang</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2076">slatedb/slatedb#2076</a></li> <li>Return cache hit or miss results from DbCache fetch methods by <a href="https://github.com/cadonna"><code>@cadonna</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2083">slatedb/slatedb#2083</a></li> <li>Stream L0 flush blocks instead of building the whole SST in memory by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2064">slatedb/slatedb#2064</a></li> <li>chore(foyer): bump to foyer v0.22.6 by <a href="https://github.com/yiming-fang"><code>@yiming-fang</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2087">slatedb/slatedb#2087</a></li> <li>Fix pause ordering in writer fencing test by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2088">slatedb/slatedb#2088</a></li> <li>fix clippy issues when no cache feature enabled by <a href="https://github.com/cadonna"><code>@cadonna</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2090">slatedb/slatedb#2090</a></li> <li>add tracing spans slatedb.read.read_filters and slatedb.read.evaluate_filters by <a href="https://github.com/cadonna"><code>@cadonna</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2054">slatedb/slatedb#2054</a></li> <li>perf(wal): pipeline replay reads with shared limits by <a href="https://github.com/rodesai"><code>@rodesai</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2059">slatedb/slatedb#2059</a></li> <li>Fix descending scan behavior in the sorted run iterator by <a href="https://github.com/nomiero"><code>@nomiero</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2067">slatedb/slatedb#2067</a></li> <li>feat(compactions): enforce RFC 0029 ULID invariants by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2081">slatedb/slatedb#2081</a></li> <li>Retry multipart parts with object_store 0.14.2 by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2094">slatedb/slatedb#2094</a></li> <li>Snapshot DST filesystem listings before yielding by <a href="https://github.com/criccomini"><code>@criccomini</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2095">slatedb/slatedb#2095</a></li> <li>Let applications evict the cache entries of retired SSTs by <a href="https://github.com/yiming-fang"><code>@yiming-fang</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2092">slatedb/slatedb#2092</a></li> <li>add tracing spans slatedb.read.read_index by <a href="https://github.com/cadonna"><code>@cadonna</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2093">slatedb/slatedb#2093</a></li> <li>Raise the key size limit to u32::MAX by <a href="https://github.com/agavra"><code>@agavra</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2099">slatedb/slatedb#2099</a></li> <li>fix: refresh GC version gauges before boundary updates by <a href="https://github.com/RanaPriyansh"><code>@RanaPriyansh</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2100">slatedb/slatedb#2100</a></li> <li>Fix clock inheritance in union clones by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2107">slatedb/slatedb#2107</a></li> <li>Add Tasklet to adopters by <a href="https://github.com/rockwotj"><code>@rockwotj</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2114">slatedb/slatedb#2114</a></li> <li>fix(clone): don't carry inherited external dbs that hold no SSTs by <a href="https://github.com/sinbad-io"><code>@sinbad-io</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2113">slatedb/slatedb#2113</a></li> <li>feat(uniffi): expose Admin::delete_db by <a href="https://github.com/sinbad-io"><code>@sinbad-io</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2116">slatedb/slatedb#2116</a></li> <li>Add an optional sorted-run count trigger for compaction by <a href="https://github.com/geeknarrator"><code>@geeknarrator</code></a> in <a href="https://redirect.github.com/slatedb/slatedb/pull/2108">slatedb/slatedb#2108</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/LGouellec"><code>@LGouellec</code></a> made their first contribution in <a href="https://redirect.github.com/slatedb/slatedb/pull/2068">slatedb/slatedb#2068</a></li> <li><a href="https://github.com/yiming-fang"><code>@yiming-fang</code></a> made their first contribution in <a href="https://redirect.github.com/slatedb/slatedb/pull/2076">slatedb/slatedb#2076</a></li> <li><a href="https://github.com/RanaPriyansh"><code>@RanaPriyansh</code></a> made their first contribution in <a href="https://redirect.github.com/slatedb/slatedb/pull/2100">slatedb/slatedb#2100</a></li> <li><a href="https://github.com/sinbad-io"><code>@sinbad-io</code></a> made their first contribution in <a href="https://redirect.github.com/slatedb/slatedb/pull/2113">slatedb/slatedb#2113</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/slatedb/slatedb/compare/v0.16.0...v0.17.0">https://github.com/slatedb/slatedb/compare/v0.16.0...v0.17.0</a></p> <h2>v0.16.0</h2> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/slatedb/slatedb/commit/c1e36fc07fb9b759d31484dda18c1c1aad8a2853"><code>c1e36fc</code></a> Bump version to 0.17.0</li> <li><a href="https://github.com/slatedb/slatedb/commit/052b1af577785c0805e1ee40324260db02b23483"><code>052b1af</code></a> Add an optional sorted-run count trigger for compaction (<a href="https://redirect.github.com/slatedb/slatedb/issues/2108">#2108</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/c4868b79b9201c1f360b3e8587491207392e9e76"><code>c4868b7</code></a> feat(uniffi): expose Admin::delete_db (<a href="https://redirect.github.com/slatedb/slatedb/issues/2116">#2116</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/a644322b01885b207bd7135c2b6d9861500b183b"><code>a644322</code></a> fix(clone): don't carry inherited external dbs that hold no SSTs (<a href="https://redirect.github.com/slatedb/slatedb/issues/2113">#2113</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/b24c488a956a36d783cc62580d5bd1e51a19f2a2"><code>b24c488</code></a> Add Tasklet to adopters (<a href="https://redirect.github.com/slatedb/slatedb/issues/2114">#2114</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/2866a351e39d89f91fda1a6a3ec2b10ecf9bfb71"><code>2866a35</code></a> Fix clock inheritance in union clones (<a href="https://redirect.github.com/slatedb/slatedb/issues/2107">#2107</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/0fb36809427d047f8899427cd1c4d15804509b14"><code>0fb3680</code></a> fix: refresh GC version gauges before boundary updates (<a href="https://redirect.github.com/slatedb/slatedb/issues/2100">#2100</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/85199863ffc140a69ab5616c01e8e2d557b2ca89"><code>8519986</code></a> Raise the key size limit to u32::MAX (<a href="https://redirect.github.com/slatedb/slatedb/issues/2099">#2099</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/3c67092fab7c9c35d366a6a7c12acb00a46afff5"><code>3c67092</code></a> add tracing spans slatedb.read.read_index (<a href="https://redirect.github.com/slatedb/slatedb/issues/2093">#2093</a>)</li> <li><a href="https://github.com/slatedb/slatedb/commit/a04b941dd68bd154f1a40b45a50104a77862eaf3"><code>a04b941</code></a> Let applications evict the cache entries of retired SSTs (<a href="https://redirect.github.com/slatedb/slatedb/issues/2092">#2092</a>)</li> <li>Additional commits viewable in <a href="https://github.com/slatedb/slatedb/compare/v0.14.1...v0.17.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…te-records and SCRAM (#923) Stacked on #902: review that one first. ## What changes, and why? It adds smoke tests that run the Kafka CLI tools against Nisshi for the broker's maintenance and authentication: - **Retention** (PostgreSQL and SQLite): expired records deleted and newer ones kept, the next offset after retention, a topic without `cleanup.policy`, `retention.ms=-1`, a 30-day `retention.ms`, `kafka-configs --add-config` and `--delete-config retention.ms`, `retention.bytes`, offsets after retention empties a partition. - **Compaction** (PostgreSQL and SQLite): the latest value per key, a tombstone as the key's only record, and `compact,delete`. - **`kafka-delete-records`** (every engine): the new low watermark, the earliest offset, the latest offset. - **SCRAM:** users created with `kafka-configs` and described with it; login with SCRAM-SHA-256 and SCRAM-SHA-512, a wrong password, and a client without credentials, all on PostgreSQL and SQLite, because a broker with `--authentication` can't create its first user, so each test restarts the broker. That restart duplicated three SCRAM tests in `restart.rs`, which this PR moves into `scram.rs`. - **SQLite `vacuum_into`:** a broker started on the snapshot, after its own database has been deleted, has the topics and records. Harness changes: `kafka-delete-records`, `kafka-configs --describe --entity-type users`, a producer that sets old timestamps, a broker that runs maintenance every 2 seconds, and a restart onto other storage that can delete files first. `Broker::file_exists` now copies the file out with `docker cp`. It used `docker exec … test`, and the broker image has no `test` command, so it returned `false` for every file in a broker container. ### Ignored tests Each is waiting for: - #918: a topic without `cleanup.policy` is never cleaned - #919: `retention.ms=-1` deletes every record - #920: a `retention.ms` above 2147483647 stops retention on PostgreSQL (2 tests) - #921: `retention.bytes` is ignored - #922: `kafka-configs` can't describe users without `DescribeClientQuotas` - #864: offsets go back to 0 when retention empties a partition - #904: `kafka-configs --delete-config` refuses a topic's own config - PR #835, which fixes DeleteRecords: the three `delete_records` tests ## Upgrade impact None. This changes only the smoke tests. ## How was this tested? Local runs of `just smoke <engine>` with every test, the ignored ones included, and the Kafka 3.9 tools: | Engine | Tests run | Passed | Failed, all ignored | |---|---|---|---| | SQLite | 86 | 69 | 17 | | Memory | 62 | 46 | 16 | | PostgreSQL | 82 | 63 | 19 | On each engine, the failed tests are exactly the tests ignored on that engine, so CI, which skips them, passes. The failures include the ignored tests from #902. - [x] `cargo clippy -p nisshi-smoke-test --all-features --all-targets -- -D warnings`, and with `--features memory` - [x] `cargo fmt --check`, and rustdoc with warnings denied - [x] each commit builds and passes the crate's unit tests on its own - [ ] S3 and the Kafka 3.7 and 4.3 tools: CI's smoke jobs 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: William Kourlas <156007774+solace-wkourlas@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
work in progress