Skip to content

Commit e2fcc87

Browse files
committed
docs(rfc): propose join-free artifact retrieval
1 parent ba7dc8d commit e2fcc87

2 files changed

Lines changed: 406 additions & 0 deletions

File tree

Lines changed: 221 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,221 @@
1+
---
2+
title: Artifact Search Projections and Join-Free Retrieval
3+
---
4+
5+
- Proposal Name: `artifact_search_projections`
6+
- Start Date: 2026-09-30
7+
- RFC PR: [oceanbase/powercontext#1803](https://github.com/oceanbase/powercontext/pull/1803)
8+
- Migration dependency: [Unified Versioned Database Migrations, #1771](https://github.com/oceanbase/powercontext/pull/1771)
9+
- Amends RFCs: [0014](0014_memory_layer_design.md), [0051](0051_experience_skill_artifact_families.md),
10+
[0080](0080_memory_search_reranking.md), [1417](1417_topic_memory.md)
11+
- Related RFCs: [1396](1396_handoff_access_control.md), [1467](1467_artifact_tags.md),
12+
[1549](1549_artifact_family_unification.md), [1652](1652_memory_quality_and_lifecycle.md)
13+
14+
# Summary
15+
16+
This RFC addresses index problems in OceanBase full-text and vector search by proposing a common query boundary for
17+
searchable Artifact Families: **full-text and vector retrieval statements must not use business-table joins**. Matching,
18+
eligibility filtering, candidate ordering, and truncation cannot depend on joins to authoritative history, heads, tags,
19+
or other business tables.
20+
21+
OceanBase must follow this query boundary. Whether it is also mandatory for other backends remains an open question.
22+
23+
Each Family should use a search wide table or separate current search projections. Tables may be separated by full-text
24+
and vector capabilities or by search granularity, such as topics and chunks. All data need not occupy one physical table.
25+
After candidates are selected, content and display fields may be fetched in batches using exact references.
26+
27+
Authoritative revisions and heads retain their identity, history, and current-version responsibilities. Search projections
28+
are rebuildable derived data maintained synchronously with authoritative state. The standard covers Memory, Topic Memory,
29+
Experience, and Skill; each Family's adaptation proceeds independently.
30+
31+
# Motivation
32+
33+
## Joins affect search indexes and retrieval results
34+
35+
Families already maintain partial search projections, but search still involves cross-table associations:
36+
37+
| Family | Current search structure |
38+
| --- | --- |
39+
| Memory | Separate full-text and vector projections; retrieval joins entry versions for content and uses correlated tag filters |
40+
| Topic Memory | Separate Topic/chunk full-text and vector projections; vector candidates join display content after truncation |
41+
| Experience | Full-text matching on common heads, followed by a join to authoritative revision content |
42+
| Skill | The same full-text retrieval path as Experience |
43+
44+
In OceanBase, a distance-ordered vector scan participating in a merge join causes nearest-neighbor loss. Joining full-text
45+
search with other business tables also degrades the full-text index path. Topic Memory already limits ANN candidates
46+
before joining content to work around the vector merge-join issue, showing that candidate selection and content reads
47+
can be organized separately.
48+
49+
Families need a clear retrieval boundary that keeps business associations for content, current state, and tags out of
50+
full-text and vector retrieval plans.
51+
52+
## Search data has different maintenance costs
53+
54+
Full text, vectors, and multivalued tags use different indexes. Topics and chunks also have different search granularities.
55+
Requiring one physical table couples content duplication, vector configuration, tag updates, and index lifecycles.
56+
Colocating fields alone does not ensure that multivalued tags and vector filtering combine efficiently.
57+
58+
Wide tables suit direct return of complete results. Separate projections suit independently maintained capabilities and
59+
granularities. The common contract should allow Families to choose an appropriate layout.
60+
61+
## Eligibility and content completion need different boundaries
62+
63+
Content can be fetched in batches after candidates are selected. Scope, tags, lifecycle, and access eligibility affect
64+
which objects qualify as candidates. Applying them only after taking a global top-k lets ineligible objects consume
65+
candidate slots and can exclude relevant eligible objects.
66+
67+
Allowing separate tables therefore requires guarantees for eligibility, exact references, and read consistency.
68+
69+
# Guide-level explanation
70+
71+
## Use a wide table or separate projections
72+
73+
A Family may store search fields and response content in its own wide table and return complete results directly.
74+
Alternatively, it may maintain separate current full-text and vector projections, retrieve exact references, match
75+
positions, and scores, then fetch content in batches.
76+
77+
Topic Memory can retain separate topic and chunk search granularities. Chunk hits still reference the exact Topic
78+
revision, preserving candidate budgets, result folding, and scoring semantics. Full-text-only Experience and Skill do
79+
not need vector projections.
80+
81+
The corresponding Family owns these layouts. Families do not have to share one search table.
82+
83+
## Return complete, consistent results
84+
85+
Search interfaces continue returning their promised content, metadata, and exact version references. With separate
86+
projections, the service performs batch content reads; callers need not request details per hit. Content reads use the
87+
version selected during retrieval and cannot replace it by resolving latest again.
88+
89+
After publication, deactivation, or tag changes, search selects candidates from consistent current state. Historical
90+
references remain readable under access control. Search projection layout does not affect exact historical reads.
91+
92+
## Preserve tag filtering semantics
93+
94+
Tag management may retain independent authoritative assignments. Search uses synchronously maintained tag projections
95+
or complete, bounded filter inputs established before retrieval to apply existing all/any matching semantics during
96+
candidate selection.
97+
98+
It cannot arbitrarily truncate the tag-matching object set before ranking or join tags after retrieval to filter results.
99+
Deployments without embeddings retain existing full-text search. Backend capability boundaries explicitly define the
100+
supported combinations of filtering and retrieval modes.
101+
102+
# Reference-level explanation
103+
104+
## Core contract
105+
106+
The following query organization and projection maintenance requirements apply to backends covered by this standard.
107+
Whether non-OceanBase backends are included remains an unresolved question. Every backend must preserve existing
108+
authorization, filtering, current-version, and returned-content semantics.
109+
110+
1. **No business joins in retrieval statements.** Full-text and vector candidate selection must not join authoritative
111+
history, heads, tags, or other business tables. Correlated subqueries cannot bring those associations back into retrieval.
112+
2. **Multiple current projections are allowed.** Family-owned wide tables or separate projections are recommended.
113+
The Family and backend determine physical table counts, whether full text and vectors share storage, and search granularity.
114+
3. **Eligibility participates in candidate selection.** Scope, tags, lifecycle, and access control retain their semantics.
115+
Filtering after truncation cannot replace retrieval of eligible candidates.
116+
4. **Content completion runs separately.** Once candidates exist, separate batch reads may complete their content.
117+
Business tables must not be joined in the same statement that performs full-text or vector retrieval. Content reads
118+
do not redetermine eligibility, the current version, or relevance ordering.
119+
5. **Current state is consistent.** Authoritative changes and affected projections update atomically. All retrieval
120+
channels and content reads in one search share a consistent committed state.
121+
6. **History and budgets are preserved.** Rebuilding changes no authoritative identities, historical content, exact
122+
references, or Source processing positions. Candidates, filter-input transfer, content reads, and fusion respect
123+
budgets, without per-hit queries or unbounded object lists connecting stages.
124+
125+
## Responsibilities and scope
126+
127+
`pc_artifacts` and `pc_artifact_heads` retain their existing roles. Projections maintained by Family writers establish
128+
currentness; retrieval does not join heads to identify the latest revision. Candidates carry exact identities sufficient
129+
to identify their original content and match positions.
130+
131+
Content completion may read the corresponding current content projection or authoritative exact-version records.
132+
It must not silently substitute another revision, omit missing content, or interpret inconsistencies as empty results.
133+
Returned content must correspond to the candidates used for scoring.
134+
135+
Independent changes to tags and other eligibility information also update every affected projection synchronously.
136+
Metadata changes need not all create content revisions. Existing Scope-level authorization can run before search;
137+
changing search storage does not require changing the authorization model.
138+
139+
Backends covered by this standard may implement search with native or auxiliary indexes. Internal database index access
140+
is not a business-table join prohibited by this RFC. Application queries still respect the boundary between retrieval
141+
and content reads. A shared standard allows backend-specific DDL and index implementations.
142+
143+
This RFC covers existing and future searchable Families and defines only the common search standard. Each Family's
144+
adaptation is designed and delivered independently; simultaneous completion is not required. Domain contracts own content
145+
generation, evidence semantics, and semantic merging. Public return content, filters, and scoring contracts remain unchanged.
146+
147+
## Upgrade and compatibility
148+
149+
Schema changes and search projection rebuilding follow
150+
[RFC #1771: Unified Versioned Database Migrations](https://github.com/oceanbase/powercontext/pull/1771),
151+
using its unified maintenance entry point for the offline upgrade.
152+
153+
Search projections are rebuilt from authoritative data without creating content revisions or re-extracting processed
154+
Sources. Existing exact references and access control are preserved. Historical content reads do not depend on historical
155+
vectors.
156+
157+
# Drawbacks
158+
159+
- **Projection maintenance costs.** Wide tables may duplicate content. Separate projections still duplicate some
160+
eligibility attributes and require synchronous maintenance of multiple current representations.
161+
- **Read coordination.** Batch content completion adds a read and must use the same versions and consistent state as retrieval.
162+
- **Index combination limits.** Backend capabilities determine how multivalued tags combine with vector retrieval.
163+
Removing joins does not eliminate filtering or scanning costs.
164+
- **Upgrade downtime.** Projection migration and index rebuilding require extra space and a maintenance window.
165+
Authoritative history continues to grow independently.
166+
167+
# Rationale and alternatives
168+
169+
## Constrain retrieval while allowing projection choices
170+
171+
Wide tables reduce content lookups and coordination between copies, suiting Families whose response fields align with
172+
search units. Separate projections allow independent maintenance of full text, vectors, topics, and chunks, reducing
173+
large-field duplication and separating index management. This proposal recommends both layouts under the same query
174+
boundary and consistency responsibilities.
175+
176+
## Require exactly one wide table per Family
177+
178+
This simplifies direct result return but restricts independent evolution of search granularities, vector configurations,
179+
and backend indexes. It also cannot replace multivalued-tag index design, so it is not a mandatory physical requirement.
180+
181+
## Keep business associations in OceanBase retrieval
182+
183+
This reduces some duplication but retains OceanBase nearest-neighbor loss and full-text index degradation. Expressing
184+
associations as correlated subqueries or outer joins in the same retrieval statement still combines business tables
185+
with the retrieval plan. Separate batch content reads establish an explicit boundary between the two stages.
186+
187+
# Prior art
188+
189+
[RFC 1417](1417_topic_memory.md) provides current Topic/chunk projections, separate retrieval channels, and atomic
190+
publication as a foundation for independent projections. [RFC 0051](0051_experience_skill_artifact_families.md) defines
191+
Experience/Skill content and admission. [RFC 1467](1467_artifact_tags.md) and
192+
[RFC 1396](1396_handoff_access_control.md) define tag and authorization semantics that must be preserved.
193+
194+
# Unresolved questions
195+
196+
1. Determine the minimum supported OceanBase version.
197+
2. Decide whether this standard is also mandatory for non-OceanBase backends.
198+
199+
The second question has two options; the backend scope remains undecided:
200+
201+
| Option | Benefits | Costs |
202+
| --- | --- | --- |
203+
| Require every backend to follow the standard | Each Family changes its shared data model once and reuses common projection maintenance and retrieval logic, reducing long-term maintenance branches. | A common retrieval design may sacrifice some retrieval performance on backends such as SQLite and limit backend-specific optimization. |
204+
| Require OceanBase to follow the standard; let other backends choose | Each backend can choose queries and storage layouts suited to its index capabilities, retaining room for independent optimization. | Divergent retrieval paths add maintenance branches. If storage models also diverge, they require mappings to a common business model and long-term maintenance of multiple read, write, migration, and verification paths. |
205+
206+
SQLite and seekDB currently primarily serve embedded, local use, with relatively small expected datasets. Given this
207+
usage, a common retrieval design with fewer maintenance branches is preferred. Whether compliance is mandatory for
208+
every backend remains undecided.
209+
210+
A shared standard does not require identical DDL or indexes across backends. Backends that are not required to follow it
211+
may still reuse the same models and retrieval implementation voluntarily.
212+
213+
If backends adopt different storage models, how to isolate those differences through a common business model and storage
214+
adapters remains a deferred question for separate discussion. This RFC does not design that architecture or make its
215+
construction a prerequisite for the index changes.
216+
217+
# Future possibilities
218+
219+
A unified search could combine each Family's results, for example returning Memory, Experience, and Skill in one request
220+
and ranking them together. Scoring and ranking across Families would need separate decisions. This capability is not a
221+
delivery requirement of this proposal.

0 commit comments

Comments
 (0)