Repository navigation
[python] Execute local vector search through Rust core - #10476
JingsongLi wants to merge 2 commits into
Conversation
leaves12138
left a comment
There was a problem hiding this comment.
Reviewed at c846e956d4751d58b08673ff4b9a690ec53787fc using a freshly built apache/paimon-rust#1087 extension at cd2e2d1f14e83893513c4b487e432b1da04ad0f1. Requesting changes: the new Native delegation does not yet preserve the Java scalar-filter policies that this PR correctly implements on its classic path. The semantic fixes principally belong in the companion core; the dependent Python path should not be enabled until they are available.
Independent validation on the fixed heads:
- Rebuilt and installed the companion Native extension from the reviewed Rust head.
- Full Rust core unit suite: 3877 passed, 6 ignored.
- Python binding vector/read/table/resolved-table suites: 197 passed.
- New real-REST Native vector suite: 30 passed.
- Fourteen related PyPaimon suites with all five Native switches enabled: 997 passed and 96 subtests passed; every update-family counter was nonzero.
- Formatting, changed-file Flake8, and Python 3.6 syntax checks passed.
- Additional reviewer Java/API parity cases: 13 failed, 6 passed. These expose the two DE scalar-filter semantic differences plus the single-vector keyword mismatch described inline.
The passing baseline suites do not cover these parity cases. Java references are the companion head's AbstractDataEvolutionVectorRead.prepareSplits (lines 159-181) and scalarMatchedRows (lines 246-266). The reproduced reference outputs use the classic Python path, checked against those Java implementations; I did not run a separate Java process.
| results = native.execute_batch_local() | ||
| return [DictBasedScoredIndexResult(result.row_ids()) for result in results] | ||
| native.with_query_vector(queries) | ||
| result = native.execute_local() |
There was a problem hiding this comment.
[P1] Do not enable Native DE filtering before matching Java's raw fallback
For a scalar predicate that cannot be evaluated by a scalar index, the classic reader added in this PR correctly follows Java AbstractDataEvolutionVectorRead.prepareSplits: it reroutes vector-index-covered ranges to raw scoring. This new delegation bypasses that policy and accepts an empty Native ANN result as a successful result, so the exception fallback does not help.
Using a real REST IVF-flat index with four clusters, ivf.nprobe=1, query [0, 0], limit 3, and id >= 100 without an id scalar index, the classic path returns 3 hits while the new public Native path returns 0. This reproduces for single/batch and FAST/FULL/DETAIL; probing all four clusters recovers the three Native hits. The core issue is detailed on apache/paimon-rust#1087. Please land the Java-equivalent core routing before enabling these filtered searches, or keep unsupported cases on the classic lane until then.
| native.with_vector_column(builder._vector_column.name) | ||
| native.with_limit(builder._limit).with_options(builder._options) | ||
| if builder._filter is not None: | ||
| native.with_filter(_predicate_to_native(builder._filter)) |
There was a problem hiding this comment.
[P2] Preserve global-index.filter.refine-from-data semantics during delegation
With an inexact scalar-index answer, Java and this PR's classic reader exclude candidates when global-index.filter.refine-from-data=false. The delegated core path instead reads the data to build an exact allow-list, ignoring that switch and changing the existing public API's result and I/O behavior just by enabling Native reading.
On a real REST table with a BTREE index on name, an IVF-flat vector index, name LIKE '%zeta%', and limit 1, classic returns 0 hits with refinement disabled whereas the new Native path returns 1 (batch: [0, 0] versus [1, 1]). All FAST/FULL/DETAIL cases reproduce it; enabled-refinement controls agree. Please require the companion core fix, or decline Native delegation for these unsupported refinement semantics.
leaves12138
left a comment
There was a problem hiding this comment.
Re-reviewed at 7e6453d1a8ba8a80da647bc5e520260061248006, with a freshly rebuilt apache/paimon-rust#1087 extension at 2a2692f5b85ab9317f4257982b37a7af3945ab6a. Both previous DE-filter findings and the companion vector= keyword mismatch are fixed; the independent 19-case regression now passes.
Keeping Request Changes for one remaining companion-core Java-parity blocker: missing raw scoring of scalar-unindexed ranges when scalar coverage is older than vector coverage and scalar-index.search-mode=full/detail. The public Native path silently accepts the incorrect empty result, so exception fallback does not protect callers. The root fix belongs in apache/paimon-rust#1087; this PR should depend on that corrected core and add the coverage regression.
Reproduction: use the real REST vector_table fixture, build BTREE on id, append {id: 7, embedding: [3, 3], pt: 1}, then build IVF-flat on embedding with dimension 2 and nlist 1. With query [0, 0], filter id = 7, limit 10 and scalar mode FULL or DETAIL, classic returns 1 hit while Native returns 0. Batch returns [1, 1] versus [0, 0]. The direct Rust binding and the public Native builder both reproduce it across all three vector modes, including FAST. Scalar FAST controls agree. Python single/batch search methods were patched to raise during Native calls.
This follows Java DataEvolutionVectorScan lines 174-191: scalar-unindexed ranges are added to raw splits according to scalar search mode, and intersected with vector coverage only when vector mode is FAST. The executable reference is classic PyPaimon on the same REST tables; Java source was checked at this fixed head, but no separate Java process was run.
Validation: Rust core 3882 passed / 6 ignored; binding suites 198 passed; this PR's real-REST vector suite 54 passed; fourteen related Native suites 997 passed / 96 subtests passed with all Native counters exercised; previous parity cases 19 passed. Additional scalar/vector coverage matrix: 12 failed / 6 scalar-FAST controls passed. Rust fmt and changed Python files' Python 3.6 syntax passed.
Small CI cleanup: configured Flake8 (--config=./dev/cfg.ini) reports E501 on the newly added native_vector_search_test.py lines 484 (122 characters) and 500 (124 characters), exceeding the repository's 120-character limit.
Purpose
Use Rust core for the local vector-search capabilities PyPaimon already supports on REST tables. Single-vector searches delegate for DE and configured PK vector indexes; batch delegation uses PyPaimon's existing DE result API.
Depends on apache/paimon-rust#1087. This PR is Draft until that binding API is available on Rust main. Native CI continues to install Apache paimon-rust main.
Changes
AbstractDataEvolutionVectorRead.prepareSplits: when a scalar filter cannot be evaluated by an index, search its covered ranges from the data rather than silently excluding those rows. Validate persisted metrics for every rerouted shard.typestring and omitting the extranullablemember that blocked Native REST loading.Tests
Real REST regression cases compare Python and Native single/batch results for all three metrics, scalar/partition predicates, indexed and uncovered ranges, FAST/FULL/DETAIL, empty results, Java-produced PK indexes, refinement, result materialization, snapshot-pinned scan/read reuse, mixed metric failures and fallback behavior.
Initial local verification:
Native execution counters verified plan/read/write/commit use and every update family.
Scope
This enables existing local search APIs. Distributed Ray scan/read and PyPaimon's PK batch API retain their current behavior. No new dependencies or CI ref pins are added.
Rust core API and semantics are supplied by apache/paimon-rust#1087.
Review fixes
The companion Rust commit
2a2692ffixes both Native delegation findings in core: unevaluable scalar filters route indexed ranges to raw scoring in every search mode, and inexact scalar answers respectglobal-index.filter.refine-from-data. This PR continues to depend on #1087 being merged and available from Apache paimon-rust main.Added real REST single/batch regressions for the four-cluster
nprobe=1reproduction, BTREE LIKE with refinement enabled/disabled, matching unindexed tails, exact indexes combined with partition filters, and unsupported predicates sharing an FM-indexed field. FM test indexes are built through Rust's existing SQL procedure. Direct binding checks and classic-reader failure guards ensure the tests exercise Native execution.Verification for commit
7e6453d1a8, using the rebuilt companion Rust extension:The prior broader verification listed above predates these review repairs; this round reran the affected vector and index paths. The PR remains Draft until the companion core API and fixes are available on Rust main.