ML Research

17% faster search, zero config: auto-calibrating vector quantization in Elasticsearch

Automatic calibration at merge time picks vector quantization parameters for each segment by predicting recall from a small sample. Here's how we built it into Elasticsearch's merge path.

17% faster search, zero config: auto-calibrating vector quantization in Elasticsearch
56% faster, up to 50% better retrieval performance: What's inside Jina's new 600 million parameter listwise reranker

July 27, 2026

56% faster, up to 50% better retrieval performance: What's inside Jina's new 600 million parameter listwise reranker

Jina Reranker 3.5 beats v3 by 50%+ on case law, closes the gap with models 7x its size on legal, medical, and financial benchmarks, and beats them outright on structured data. It's a drop-in replacement for v3, with no API changes.

How Elasticsearch detects multiple change points in time series with 0.99 recall

July 24, 2026

How Elasticsearch detects multiple change points in time series with 0.99 recall

ES|QL's CHANGE_POINT command finds structural shifts, variance changes and spikes in any metric in ~1ms, without tuning anything per series.

How Elasticsearch auto-tunes vector quantization to hit your recall target

How Elasticsearch auto-tunes vector quantization to hit your recall target

Learn the geometric model that lets Elasticsearch predict recall with R² > 0.98 accuracy and auto-select vector quantization parameters from a small data sample.

How BBQ shrinks Jina v5 embeddings by 29x without losing recall in Elasticsearch

July 10, 2026

How BBQ shrinks Jina v5 embeddings by 29x without losing recall in Elasticsearch

A hands-on test comparing BBQ and float32 vector indices in Elasticsearch, measuring memory, disk and recall@10 across five languages.

Elasticsearch DiskBBQ delivers 7x faster vector search than Qdrant on network-attached storage

June 24, 2026

Elasticsearch DiskBBQ delivers 7x faster vector search than Qdrant on network-attached storage

Elasticsearch DiskBBQ achieves up to 7x higher vector search throughput than Qdrant at comparable recall on network-attached storage. Explore the benchmark methodology and full results.

Is your ML job's datafeed losing a race it cannot win?

April 15, 2026

Is your ML job's datafeed losing a race it cannot win?

Learn how switching from scroll-based to aggregation-based datafeeds optimizes machine learning jobs for large-scale deployments.

Unsupervised document clustering with Elasticsearch + Jina embeddings

Unsupervised document clustering with Elasticsearch + Jina embeddings

A practical, reproducible approach to unsupervised document clustering with Elasticsearch and Jina embeddings.

Automating log parsing in Streams with ML

January 2, 2026

Automating log parsing in Streams with ML

Learn how a hybrid ML approach achieved 94% log parsing and 91% log partitioning accuracy through automation experiments with log format fingerprinting in Streams.

Ready to build state of the art search experiences?

Sufficiently advanced search isn’t achieved with the efforts of one. Elasticsearch is powered by data scientists, ML ops, engineers, and many more who are just as passionate about search as you are. Let’s connect and work together to build the magical search experience that will get you the results you want.

Try it yourself