|
| 1 | +--- |
| 2 | +title: Artifact Search Projections and Join-Free Retrieval |
| 3 | +--- |
| 4 | + |
| 5 | +- Proposal Name: `artifact_search_projections` |
| 6 | +- Start Date: 2026-09-30 |
| 7 | +- RFC PR: [oceanbase/powercontext#1803](https://github.com/oceanbase/powercontext/pull/1803) |
| 8 | +- Migration dependency: [Unified Versioned Database Migrations, #1771](https://github.com/oceanbase/powercontext/pull/1771) |
| 9 | +- Amends RFCs: [0014](0014_memory_layer_design.md), [0051](0051_experience_skill_artifact_families.md), |
| 10 | + [0080](0080_memory_search_reranking.md), [1417](1417_topic_memory.md) |
| 11 | +- Related RFCs: [1396](1396_handoff_access_control.md), [1467](1467_artifact_tags.md), |
| 12 | + [1549](1549_artifact_family_unification.md), [1652](1652_memory_quality_and_lifecycle.md) |
| 13 | + |
| 14 | +# Summary |
| 15 | + |
| 16 | +This RFC addresses index problems in OceanBase full-text and vector search by proposing a common query boundary for |
| 17 | +searchable Artifact Families: **full-text and vector retrieval statements must not use business-table joins**. Matching, |
| 18 | +eligibility filtering, candidate ordering, and truncation cannot depend on joins to authoritative history, heads, tags, |
| 19 | +or other business tables. |
| 20 | + |
| 21 | +OceanBase must follow this query boundary. Whether it is also mandatory for other backends remains an open question. |
| 22 | + |
| 23 | +Each Family should use a search wide table or separate current search projections. Tables may be separated by full-text |
| 24 | +and vector capabilities or by search granularity, such as topics and chunks. All data need not occupy one physical table. |
| 25 | +After candidates are selected, content and display fields may be fetched in batches using exact references. |
| 26 | + |
| 27 | +Authoritative revisions and heads retain their identity, history, and current-version responsibilities. Search projections |
| 28 | +are rebuildable derived data maintained synchronously with authoritative state. The standard covers Memory, Topic Memory, |
| 29 | +Experience, and Skill; each Family's adaptation proceeds independently. |
| 30 | + |
| 31 | +# Motivation |
| 32 | + |
| 33 | +## Joins affect search indexes and retrieval results |
| 34 | + |
| 35 | +Families already maintain partial search projections, but search still involves cross-table associations: |
| 36 | + |
| 37 | +| Family | Current search structure | |
| 38 | +| --- | --- | |
| 39 | +| Memory | Separate full-text and vector projections; retrieval joins entry versions for content and uses correlated tag filters | |
| 40 | +| Topic Memory | Separate Topic/chunk full-text and vector projections; vector candidates join display content after truncation | |
| 41 | +| Experience | Full-text matching on common heads, followed by a join to authoritative revision content | |
| 42 | +| Skill | The same full-text retrieval path as Experience | |
| 43 | + |
| 44 | +In OceanBase, a distance-ordered vector scan participating in a merge join causes nearest-neighbor loss. Joining full-text |
| 45 | +search with other business tables also degrades the full-text index path. Topic Memory already limits ANN candidates |
| 46 | +before joining content to work around the vector merge-join issue, showing that candidate selection and content reads |
| 47 | +can be organized separately. |
| 48 | + |
| 49 | +Families need a clear retrieval boundary that keeps business associations for content, current state, and tags out of |
| 50 | +full-text and vector retrieval plans. |
| 51 | + |
| 52 | +## Search data has different maintenance costs |
| 53 | + |
| 54 | +Full text, vectors, and multivalued tags use different indexes. Topics and chunks also have different search granularities. |
| 55 | +Requiring one physical table couples content duplication, vector configuration, tag updates, and index lifecycles. |
| 56 | +Colocating fields alone does not ensure that multivalued tags and vector filtering combine efficiently. |
| 57 | + |
| 58 | +Wide tables suit direct return of complete results. Separate projections suit independently maintained capabilities and |
| 59 | +granularities. The common contract should allow Families to choose an appropriate layout. |
| 60 | + |
| 61 | +## Eligibility and content completion need different boundaries |
| 62 | + |
| 63 | +Content can be fetched in batches after candidates are selected. Scope, tags, lifecycle, and access eligibility affect |
| 64 | +which objects qualify as candidates. Applying them only after taking a global top-k lets ineligible objects consume |
| 65 | +candidate slots and can exclude relevant eligible objects. |
| 66 | + |
| 67 | +Allowing separate tables therefore requires guarantees for eligibility, exact references, and read consistency. |
| 68 | + |
| 69 | +# Guide-level explanation |
| 70 | + |
| 71 | +## Use a wide table or separate projections |
| 72 | + |
| 73 | +A Family may store search fields and response content in its own wide table and return complete results directly. |
| 74 | +Alternatively, it may maintain separate current full-text and vector projections, retrieve exact references, match |
| 75 | +positions, and scores, then fetch content in batches. |
| 76 | + |
| 77 | +Topic Memory can retain separate topic and chunk search granularities. Chunk hits still reference the exact Topic |
| 78 | +revision, preserving candidate budgets, result folding, and scoring semantics. Full-text-only Experience and Skill do |
| 79 | +not need vector projections. |
| 80 | + |
| 81 | +The corresponding Family owns these layouts. Families do not have to share one search table. |
| 82 | + |
| 83 | +## Return complete, consistent results |
| 84 | + |
| 85 | +Search interfaces continue returning their promised content, metadata, and exact version references. With separate |
| 86 | +projections, the service performs batch content reads; callers need not request details per hit. Content reads use the |
| 87 | +version selected during retrieval and cannot replace it by resolving latest again. |
| 88 | + |
| 89 | +After publication, deactivation, or tag changes, search selects candidates from consistent current state. Historical |
| 90 | +references remain readable under access control. Search projection layout does not affect exact historical reads. |
| 91 | + |
| 92 | +## Preserve tag filtering semantics |
| 93 | + |
| 94 | +Tag management may retain independent authoritative assignments. Search uses synchronously maintained tag projections |
| 95 | +or complete, bounded filter inputs established before retrieval to apply existing all/any matching semantics during |
| 96 | +candidate selection. |
| 97 | + |
| 98 | +It cannot arbitrarily truncate the tag-matching object set before ranking or join tags after retrieval to filter results. |
| 99 | +Deployments without embeddings retain existing full-text search. Backend capability boundaries explicitly define the |
| 100 | +supported combinations of filtering and retrieval modes. |
| 101 | + |
| 102 | +# Reference-level explanation |
| 103 | + |
| 104 | +## Core contract |
| 105 | + |
| 106 | +The following query organization and projection maintenance requirements apply to backends covered by this standard. |
| 107 | +Whether non-OceanBase backends are included remains an unresolved question. Every backend must preserve existing |
| 108 | +authorization, filtering, current-version, and returned-content semantics. |
| 109 | + |
| 110 | +1. **No business joins in retrieval statements.** Full-text and vector candidate selection must not join authoritative |
| 111 | + history, heads, tags, or other business tables. Correlated subqueries cannot bring those associations back into retrieval. |
| 112 | +2. **Multiple current projections are allowed.** Family-owned wide tables or separate projections are recommended. |
| 113 | + The Family and backend determine physical table counts, whether full text and vectors share storage, and search granularity. |
| 114 | +3. **Eligibility participates in candidate selection.** Scope, tags, lifecycle, and access control retain their semantics. |
| 115 | + Filtering after truncation cannot replace retrieval of eligible candidates. |
| 116 | +4. **Content completion runs separately.** Once candidates exist, separate batch reads may complete their content. |
| 117 | + Business tables must not be joined in the same statement that performs full-text or vector retrieval. Content reads |
| 118 | + do not redetermine eligibility, the current version, or relevance ordering. |
| 119 | +5. **Current state is consistent.** Authoritative changes and affected projections update atomically. All retrieval |
| 120 | + channels and content reads in one search share a consistent committed state. |
| 121 | +6. **History and budgets are preserved.** Rebuilding changes no authoritative identities, historical content, exact |
| 122 | + references, or Source processing positions. Candidates, filter-input transfer, content reads, and fusion respect |
| 123 | + budgets, without per-hit queries or unbounded object lists connecting stages. |
| 124 | + |
| 125 | +## Responsibilities and scope |
| 126 | + |
| 127 | +`pc_artifacts` and `pc_artifact_heads` retain their existing roles. Projections maintained by Family writers establish |
| 128 | +currentness; retrieval does not join heads to identify the latest revision. Candidates carry exact identities sufficient |
| 129 | +to identify their original content and match positions. |
| 130 | + |
| 131 | +Content completion may read the corresponding current content projection or authoritative exact-version records. |
| 132 | +It must not silently substitute another revision, omit missing content, or interpret inconsistencies as empty results. |
| 133 | +Returned content must correspond to the candidates used for scoring. |
| 134 | + |
| 135 | +Independent changes to tags and other eligibility information also update every affected projection synchronously. |
| 136 | +Metadata changes need not all create content revisions. Existing Scope-level authorization can run before search; |
| 137 | +changing search storage does not require changing the authorization model. |
| 138 | + |
| 139 | +Backends covered by this standard may implement search with native or auxiliary indexes. Internal database index access |
| 140 | +is not a business-table join prohibited by this RFC. Application queries still respect the boundary between retrieval |
| 141 | +and content reads. A shared standard allows backend-specific DDL and index implementations. |
| 142 | + |
| 143 | +This RFC covers existing and future searchable Families and defines only the common search standard. Each Family's |
| 144 | +adaptation is designed and delivered independently; simultaneous completion is not required. Domain contracts own content |
| 145 | +generation, evidence semantics, and semantic merging. Public return content, filters, and scoring contracts remain unchanged. |
| 146 | + |
| 147 | +## Upgrade and compatibility |
| 148 | + |
| 149 | +Schema changes and search projection rebuilding follow |
| 150 | +[RFC #1771: Unified Versioned Database Migrations](https://github.com/oceanbase/powercontext/pull/1771), |
| 151 | +using its unified maintenance entry point for the offline upgrade. |
| 152 | + |
| 153 | +Search projections are rebuilt from authoritative data without creating content revisions or re-extracting processed |
| 154 | +Sources. Existing exact references and access control are preserved. Historical content reads do not depend on historical |
| 155 | +vectors. |
| 156 | + |
| 157 | +# Drawbacks |
| 158 | + |
| 159 | +- **Projection maintenance costs.** Wide tables may duplicate content. Separate projections still duplicate some |
| 160 | + eligibility attributes and require synchronous maintenance of multiple current representations. |
| 161 | +- **Read coordination.** Batch content completion adds a read and must use the same versions and consistent state as retrieval. |
| 162 | +- **Index combination limits.** Backend capabilities determine how multivalued tags combine with vector retrieval. |
| 163 | + Removing joins does not eliminate filtering or scanning costs. |
| 164 | +- **Upgrade downtime.** Projection migration and index rebuilding require extra space and a maintenance window. |
| 165 | + Authoritative history continues to grow independently. |
| 166 | + |
| 167 | +# Rationale and alternatives |
| 168 | + |
| 169 | +## Constrain retrieval while allowing projection choices |
| 170 | + |
| 171 | +Wide tables reduce content lookups and coordination between copies, suiting Families whose response fields align with |
| 172 | +search units. Separate projections allow independent maintenance of full text, vectors, topics, and chunks, reducing |
| 173 | +large-field duplication and separating index management. This proposal recommends both layouts under the same query |
| 174 | +boundary and consistency responsibilities. |
| 175 | + |
| 176 | +## Require exactly one wide table per Family |
| 177 | + |
| 178 | +This simplifies direct result return but restricts independent evolution of search granularities, vector configurations, |
| 179 | +and backend indexes. It also cannot replace multivalued-tag index design, so it is not a mandatory physical requirement. |
| 180 | + |
| 181 | +## Keep business associations in OceanBase retrieval |
| 182 | + |
| 183 | +This reduces some duplication but retains OceanBase nearest-neighbor loss and full-text index degradation. Expressing |
| 184 | +associations as correlated subqueries or outer joins in the same retrieval statement still combines business tables |
| 185 | +with the retrieval plan. Separate batch content reads establish an explicit boundary between the two stages. |
| 186 | + |
| 187 | +# Prior art |
| 188 | + |
| 189 | +[RFC 1417](1417_topic_memory.md) provides current Topic/chunk projections, separate retrieval channels, and atomic |
| 190 | +publication as a foundation for independent projections. [RFC 0051](0051_experience_skill_artifact_families.md) defines |
| 191 | +Experience/Skill content and admission. [RFC 1467](1467_artifact_tags.md) and |
| 192 | +[RFC 1396](1396_handoff_access_control.md) define tag and authorization semantics that must be preserved. |
| 193 | + |
| 194 | +# Unresolved questions |
| 195 | + |
| 196 | +1. Determine the minimum supported OceanBase version. |
| 197 | +2. Decide whether this standard is also mandatory for non-OceanBase backends. |
| 198 | + |
| 199 | +The second question has two options; the backend scope remains undecided: |
| 200 | + |
| 201 | +| Option | Benefits | Costs | |
| 202 | +| --- | --- | --- | |
| 203 | +| Require every backend to follow the standard | Each Family changes its shared data model once and reuses common projection maintenance and retrieval logic, reducing long-term maintenance branches. | A common retrieval design may sacrifice some retrieval performance on backends such as SQLite and limit backend-specific optimization. | |
| 204 | +| Require OceanBase to follow the standard; let other backends choose | Each backend can choose queries and storage layouts suited to its index capabilities, retaining room for independent optimization. | Divergent retrieval paths add maintenance branches. If storage models also diverge, they require mappings to a common business model and long-term maintenance of multiple read, write, migration, and verification paths. | |
| 205 | + |
| 206 | +SQLite and seekDB currently primarily serve embedded, local use, with relatively small expected datasets. Given this |
| 207 | +usage, a common retrieval design with fewer maintenance branches is preferred. Whether compliance is mandatory for |
| 208 | +every backend remains undecided. |
| 209 | + |
| 210 | +A shared standard does not require identical DDL or indexes across backends. Backends that are not required to follow it |
| 211 | +may still reuse the same models and retrieval implementation voluntarily. |
| 212 | + |
| 213 | +If backends adopt different storage models, how to isolate those differences through a common business model and storage |
| 214 | +adapters remains a deferred question for separate discussion. This RFC does not design that architecture or make its |
| 215 | +construction a prerequisite for the index changes. |
| 216 | + |
| 217 | +# Future possibilities |
| 218 | + |
| 219 | +A unified search could combine each Family's results, for example returning Memory, Experience, and Skill in one request |
| 220 | +and ranking them together. Scoring and ranking across Families would need separate decisions. This capability is not a |
| 221 | +delivery requirement of this proposal. |
0 commit comments