Documentation
README
Scaling for Query Volume
Problem: When a query has a large limit (e.g. 1000) and there are multiple shards (e.g. 10), naively each shard must return the full 1000 results β totaling 10,000 scored points transferred and merged. This is wasteful since data is randomly distributed across auto-shards.
Core idea
Instead of asking every shard for the full limit, ask each shard for a smaller limit computed via Poisson distribution statistics, then merge. This is safe because auto-sharding guarantees random, independent data distribution.
This is the opening of the README. Read the full README on GitHub.