Elasticsearch kNN alternatives: when to move vector search out

Jordan Cole
Published
AI DEVELOPER TOOLSElasticsearch kNNalternatives: when to movevector search out

Already on Elasticsearch? Compare quantization, filters, hybrid search tiers, AGPL licensing, OpenSearch and dedicated vector databases before moving vectors.

Find an AI market worth building in before anyone big claims it.

Every Monday we run every tracked search through four checks: buyers are looking for a tool, demand is rising, advertisers pay real money for every click, and a focused new site can still reach the first page. The few that pass are that week's openings.

Two searches and two growing AI companies each week, free. No card needed.

Plans from $49 a month

If you already run Elasticsearch, test a better-configured Elasticsearch before any alternative. Several old reasons to leave, such as a supposed limit of three similarity metrics or quantization that always hurts results, don't describe current releases. What remains is a narrower list: vector memory competing with the rest of your cluster, hybrid ranking features your subscription doesn't include, a licensing requirement, or a product where vector search is the whole job.

Leaving also has a cost that is easy to miss. In Elasticsearch, BM25 keyword relevance, filters, aggregations and vector similarity can run in one request against one index. Move the vectors elsewhere and your application has to join two systems, keep them in sync and secure both.

This guide is based on Elastic's and OpenSearch's documentation, checked on September 24, 2026. It is not a benchmark we ran.

Match the move to the problem you actually have

What is prompting the changeTry first inside ElasticsearchWhen leaving makes sense
Vector memory is too large or searches slow down as data growsCheck the field's quantization and size the page cacheQuantized vectors still cost more than a separate service, or you can't isolate the workload
Filtered searches return fewer results than requestedMove the filter inside the kNN queryFiltered queries still miss your latency target at your real filter selectivity
You need rank fusion for hybrid searchWeighted knn plus query scoringYou need RRF or linear retrievers and won't buy the Enterprise tier
You need an OSI-approved license for the whole distributionRead what the 2024 AGPL option coversOpenSearch's Apache 2.0 license is a hard requirement
You want managed vector search but prefer Elastic's query languageElastic's Vector Database project typeIt fits, but remains a separate project to keep in sync
Vector retrieval is the product and keyword search isn't neededNothing, if the current setup already worksA dedicated vector database suits your team's operations better

Current Elasticsearch already quantizes most vectors

On Elasticsearch 9.1 and later, a float dense_vector field with 384 or more dimensions defaults to bbq_hnsw, a binary-quantized HNSW index. Smaller float vectors default to int8_hnsw, and 9.0 used int8_hnsw for all float vectors. From 9.4, and on Serverless, the default becomes bbq_disk (DiskBBQ) when your license includes it, which on a self-managed cluster means Enterprise. Elastic describes the memory reductions as roughly 4x for int8, 8x for int4 and 32x for BBQ, each at some cost to accuracy. Rescoring with oversampling is available to recover some of that accuracy. See the dense_vector reference.

Memory is a common reason to ask whether vectors should move. Elastic's kNN tuning guide says HNSW works efficiently only when most vector data is in memory, and gives a formula for estimating it. Applied to a hypothetical 10 million vectors of 1,024 dimensions with the default graph setting, that formula gives about 42 GB unquantized, about 11 GB with int8 and about 2 GB with BBQ, graph included (decimal gigabytes). That is arithmetic, not a measurement, but it shows why an unquantized index can look far more expensive than it needs to be.

Quantized fields keep the raw float vectors on disk, adding roughly 25% (int8), 12.5% (int4) or 3.1% (BBQ) to disk use. And you can change a field's index type in place along documented upgrade paths, but vectors already indexed keep their original type until they are reindexed. Check what your existing mappings actually use.

Elasticsearch also lists a 4,096-dimension cap for dense_vector. That matters only if your embedding model produces more; OpenSearch's documented limit is 16,000.

Filters and aggregations are what you would give up

A common filtered-search complaint has a documented cause. A filter placed inside the kNN query is applied during the approximate search, so Elasticsearch looks for enough matching candidates. Filters elsewhere in the request run afterwards and can leave fewer than k results even when enough matching documents exist. See the kNN query reference.

Restrictive filters can also slow approximate search, because the engine explores more of the graph to find eligible candidates. Elastic's kNN search guide explains that Lucene switches to a brute-force search over the filtered documents when the filtered set is small enough. Test your real filter selectivity before concluding the engine is at fault.

The same guide shows what a separate vector store would cost you: one request can combine knn with a standard query, blending lexical relevance, filters and aggregations, with boosts weighting each score.

Suppose a marketplace search for "waterproof hiking boots" must match product text, respect size and stock filters, rank similar items and show brand counts. In Elasticsearch that can be one request. With vectors in a separate database, you either copy the filter fields into both systems and keep them current, or fetch candidate IDs from one system and filter or aggregate them in the other. Both approaches add code, latency and ways for results to disagree.

Two searches and two growing AI companies each week, free. No card needed.

Plans from $49 a month

Check what your subscription includes before paying for a migration

Elastic's self-managed subscription comparison, checked September 24, 2026, lists vector search at every tier, including the free Basic tier. It lists Reciprocal Rank Fusion, the linear and RRF retrievers, DiskBBQ and the Inference API only under Enterprise. Gold is marked discontinued and Platinum is for existing customers only, so a new self-managed buyer is choosing between Basic and Enterprise. Elastic Cloud Hosted's pricing page likewise places hybrid search with retrievers, and DiskBBQ, in its Enterprise tier.

If you want rank fusion and Enterprise is out of budget, weighted score combination remains, but it is a different method. Elastic's retriever reference also marks the RRF and linear retrievers as generally available on Serverless, which is priced separately from these tiers, so a Serverless project is another route to test. OpenSearch's hybrid query offers both a score-normalization processor and an RRF-based processor under Apache 2.0.

The 2024 licensing change is narrower than it sounds. In August 2024 Elastic announced AGPLv3 as an option alongside SSPL and the Elastic License 2.0. Its licensing FAQ says the option covers the free portions of the source code, and that the default distribution continues under the Elastic License 2.0. The AGPL option doesn't unlock paid features. If your legal requirement is an Apache 2.0 license for everything you run, OpenSearch is the path to test.

OpenSearch k-NN is the closest move, but not a drop-in

OpenSearch is an Apache 2.0 fork of Elasticsearch 7.10.2. Its FAQ says upgrading from Elasticsearch 7.11 or later isn't supported, so moving from an 8.x or 9.x cluster is a data migration. The vector field type is different too: knn_vector rather than dense_vector, with Faiss as the default engine, Lucene as an option and NMSLIB deprecated. See methods and engines. OpenSearch's Migration Assistant documents a step for converting dense_vector fields. Its compatibility table lists Elasticsearch sources up to 8.x, not 9.x, so check support before planning a 9.x migration around it.

Filtering follows the same pattern as Elasticsearch. Efficient filtering inside the k-NN query is supported for the Lucene and Faiss engines; post-filtering can return far fewer than k results. See k-NN filtering. For memory, disk-based vector search defaults to 32x compression and rescores with full-precision vectors from disk.

Choose OpenSearch for its license or engine choices, not as an unmeasured performance fix.

Elastic now sells a separate vector project

Elastic announced Elasticsearch Vector Database on September 11, 2026, a Serverless project type for vector workloads. Its documentation says each project stores up to 1 TB and is billed on storage, search, ingest and infrastructure rather than compute units. It supports semantic and hybrid search but not time series or LogsDB index modes.

It suits teams that want managed vector search without a new query language. Next to a self-managed or Hosted cluster, it is still a second system to synchronize and secure.

When a dedicated vector database earns its place

A separate vector database is easier to justify when vector retrieval is the product, keyword relevance and aggregations aren't needed in the same request, and your team would rather operate or pay for a service built around that job. Common candidates include Qdrant, Milvus and Pinecone, or pgvector if Postgres already holds the records; their filtering, hosting and migration details deserve separate evaluation.

Before moving, run one comparison on your own data:

  • The same embeddings and documents in both systems, with no model change.
  • Recall against an exact baseline, such as Elasticsearch's script_score exact kNN on a sample.
  • Your real filters, including the restrictive ones, and your normal concurrency.
  • Updates and deletes arriving while searches run.
  • The full bill, including the synchronization job and the engineering time to maintain two systems.

If a properly quantized Elasticsearch index with filters in the right place meets that test, keep it. Move when a specific requirement still fails and the new system solves it with a cost you can name.

Find an AI market worth building in before anyone big claims it.

Every Monday we run every tracked search through four checks: buyers are looking for a tool, demand is rising, advertisers pay real money for every click, and a focused new site can still reach the first page. The few that pass are that week's openings.

Two searches and two growing AI companies each week, free. No card needed.

Plans from $49 a month

Jordan Cole

Creator of NightWatcher AI. Specializes in data-driven insights for AI product development, market validation, and competitive analysis.

More from Vector Databases