Skip to content

fix: fall back to brute force in HNSW AnnIterator at high filter ratios - #1863

Open
rere950303 wants to merge 2 commits into
zilliztech:mainfrom
rere950303:fix/hnsw-iterator-brute-force-high-filter-ratio
Open

rere950303 wants to merge 2 commits into
zilliztech:mainfrom
rere950303:fix/hnsw-iterator-brute-force-high-filter-ratio

Conversation

@rere950303

@rere950303 rere950303 commented Sep 30, 2026 •

Copy link
Copy Markdown

issue: #1862

Summary

  • The HNSW AnnIterator had no brute-force fallback at high filter ratios. At high filter ratios FaissHnswIterator kept traversing the graph and visited almost every node to find the few unfiltered ones.
  • A new kHnswSearchIteratorBFFilterThreshold (97%, same value as the range-search threshold) decides the switch. At or above it, the iterator scans the unfiltered points once and hands them to IndexIterator in a single batch. IndexIterator keeps them in its heap and returns them in order.
    • The existing accumulated_alpha behavior at the kNN threshold (93%) is unchanged.
    • Per review, the threshold is a static constant for now and can become a search parameter in a follow-up PR.
    • Distances come from the same storage distance computer, or from the refine index when present, as the brute-force kNN Search does. Label mapping and result id mapping are unchanged.
  • This is the iterator counterpart of fix: check brute-force threshold before iterator path in HNSW RangeSearch #1535, which did the same for RangeSearch.

Behavior change

At high filter ratios the iterator now returns every unfiltered point, including points that are unreachable from the entry point in the filtered graph. RaBitQ filtered results and exhausted iterators match full-code references pinned the reachable set, so it now expects all unfiltered points once the threshold is reached.

Tests

  • New section Test HNSW iterator falls back to brute force at high filter ratio (HNSW, HNSW_SQ, HNSW_PQ, with and without refine; 98% / 99%; first-N and random bitsets; L2 / IP / COSINE). It checks that:
    • every unfiltered point is returned and no filtered one is;
    • the top-k matches brute force (exact for HNSW and refined indexes, the index's own Search for the quantized ones).
  • knowhere_tests (Release, aarch64): 286 test cases; the only failure before the test update was the RaBitQ reachability assertion above.

Benchmark

Synthetic, 200k × 768 fp32, COSINE, HNSW M=16 / efConstruction=200, ef=750, random bitset, 20 queries, time to get the first 16 results from AnnIterator (best of 3), aarch64 10 cores.

filter ratio before (ms/q) after (ms/q) speedup recall@16 before recall@16 after
90.0% 21.48 22.66 — (unchanged path) 0.847 0.847
95.0% 28.05 unchanged — (below the 97% threshold) 0.872 unchanged
99.0% 56.01 1.33 42.2x 0.931 1.000
99.9% 97.75 1.08 90.9x 0.969 1.000

Milvus search iterator V2 calls AnnIterator on every page, so each page of a filtered iterator search pays the traversal. In our deployment, the p99 latency of filtered HNSW search_iterator calls rose by roughly an order of magnitude after moving from the pymilvus V1 iterator to V2. Note that this also came with a larger effective ef: pymilvus 2.4's V1 iterator clamps ef to the batch size before issuing RangeSearch, while V2 passes the configured ef through.

When the bitset filters out at least kHnswSearchKnnBFFilterThreshold (93%) of
the points, FaissHnswIterator kept traversing the graph and visited almost every
node to find the few unfiltered ones. kNN Search already switches to brute force
at this threshold, and RangeSearch does since zilliztech#1535.

At or above the threshold the iterator now scans the unfiltered points once and
hands them to IndexIterator in a single batch. Distances come from the same
storage distance computer (or the refine index when present, as the brute-force
kNN Search does); label mapping and result id mapping are unchanged.

A side effect is that points unreachable from the entry point are now returned
too; the RaBitQ acceptance test is updated accordingly.

Signed-off-by: Hyungwook Yang <yhwjjang1995@naver.com>
@sre-ci-robot

Copy link
Copy Markdown
Collaborator

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: rere950303
To complete the pull request process, please assign foxspy after the PR has been reviewed.
You can assign the PR to them by writing /assign @foxspy in a comment when ready.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@rere950303

Copy link
Copy Markdown
Author

Follow-up (not in this PR): DiskANN has the same gap. PQFlashIndex::getIteratorNextBatch has no brute-force switch (// todo: switch to quant-bf, thirdparty/DiskANN/src/pq_flash_index.cpp). DiskANN does not override RangeSearch, so both RangeSearch (IndexNode::RangeSearch → AnnIterator) and Milvus search iterator V2 go through that iterator. kNN Search does switch at kFilterThreshold (0.93, brute_force_beam_search). Fixing the iterator would cover both callers.

This one is larger than the HNSW change, because the brute-force scan has to read the vectors from disk (or scan PQ codes and then refine from disk). I'd like to open a separate issue to agree on the approach first.

@sre-ci-robot

Copy link
Copy Markdown
Collaborator

Welcome @rere950303! It looks like this is your first PR to zilliztech/knowhere 🎉

@alexanderguzhva

alexanderguzhva commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

@rere950303 I'm taking a look

@alexanderguzhva

alexanderguzhva commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

@rere950303 I've did a benchmark for our typical datasets that we use for testing. Sure, I do confirm performance gains for the provided operating point (ef=750), which is a bit unusual operating point for the search. However, for typical search scenario, 93% and even 95% is way too "early" threshold, because a switch to the brute-force overall slows down iterators significantly (but may improve the recall rate). For 99% - sure, the brute force clearly wins in all of the cases that have been tried, although I believe that it also depends on the dimensionality of the dataset. I'd rather bet of 97% as a default threshold that better suits a wider set of use cases.

I think that the provided PR makes sense if the ANN iterator threshold would go up from 93% to 97%. The reasonable way to do so might be to introduce a new constant here

struct HnswSearchThresholds {
.

Given your particular search scenario, where ef is quite high and that benefits from a lower level of such a threshold, you could also consider making such a threshold parameter configurable. So, it needs to be transformed from a static field into a regular passable search parameter. I think that it is acceptable to introduce a new static float threshold first in this PR, and transition the threshold into a regular config parameter in the following PR.

Please let me know what you think.

…witch

Add kHnswSearchIteratorBFFilterThreshold (0.97) and use it to decide when
FaissHnswIterator scans the unfiltered points. At typical ef values the graph
traversal of an iterator stays cheaper than the brute force below ~97%. The
accumulated_alpha behavior at the kNN threshold (93%) is unchanged.

Signed-off-by: Hyungwook Yang <yhwjjang1995@naver.com>
@rere950303

Copy link
Copy Markdown
Author

@alexanderguzhva Thanks for running it on your datasets. Agreed. Pushed 29bfcc8:

  • Adds HnswSearchThresholds::kHnswSearchIteratorBFFilterThreshold = 0.97f and uses it only for the iterator's brute-force switch.
  • The existing accumulated_alpha behavior at kHnswSearchKnnBFFilterThreshold (93%) is unchanged.
  • The new test now uses 98% / 99% filter ratios, and the RaBitQ acceptance test checks the reachable set below the new threshold.

I'm happy to follow up with a separate PR that turns it into a search parameter.

You're right that ef=750 is unusual, and I need to correct the issue description on that point. On our previous version (knowhere 2.3.14), pymilvus 2.4's V1 iterator clamped ef to the batch size (16) before issuing RangeSearch. That made it a short, ef-bounded search, and brute force only applied at 97%. So our regression came from the move to AnnIterator together with the effective ef jumping from 16 to our configured 750, not from losing a 93% brute-force switch. We'll lower ef for those iterator calls on our side. This PR still helps at ≥97% regardless of ef: in the same synthetic setup at 99% filtered, the graph iterator takes 13.6 ms with recall@16 0.66 at ef=16, versus 1.6 ms with recall 1.0 for the brute-force scan.

@alexanderguzhva

Copy link
Copy Markdown
Collaborator

/lgtm

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants