Research

Databricks Launches NEAREST BY SQL Join for Vector Search

Databricks has launched a native SQL join called NEAREST BY to accelerate batch vector search directly within its runtime, eliminating the need for external vector databases.

Databricks AI1 day agoResearch
Image: Databricks AI

Databricks has integrated batch vector search directly into its native C++ Photon engine using a new SQL join syntax called NEAREST BY. Instead of treating batch searches as millions of individual queries dispatched to an external database, the system processes them as a single top-k ranking join. This architecture allows practitioners to run massive vector workloads directly on their existing Delta tables without setting up or syncing a separate vector database.

To maximize hardware efficiency, the update introduces a fused Photon operator that utilizes a custom blocked GEMM kernel. This design increases arithmetic intensity, allowing the system to transition from being memory-bound to compute-bound once a batch size of 64 queries is reached. The engine leverages SIMD-accelerated distance functions—such as L2 distance, cosine similarity, and inner product—compiled into instruction-set-specific clones for different CPU architectures. For massive datasets, an optional inverted file index, stored as a liquid-clustered Delta table, prunes unnecessary data partitions before they are read.

In benchmark tests targeting at least 96% recall@K, the native implementation demonstrated significant speedups across common workloads. A classification task matching 10 million queries against 100,000 reference vectors finished in under a minute. For larger operations, a semantic deduplication self-join of 10 million records, a recommendation batch of 1 million queries against 10 million vectors, and an entity resolution task matching 1 million queries against 1 billion reference vectors all completed within minutes.

For data engineers and AI practitioners, this release simplifies the data pipeline by keeping embeddings inside the Lakehouse. It eliminates the operational overhead, synchronization lag, and licensing costs of maintaining a secondary vector store. Because the search runs within the standard Databricks Runtime, it automatically inherits enterprise-grade features like automatic task retries, disk-spilling under memory pressure, and elastic scaling across hundreds of CPU cores.

This is our own summary of reporting by Databricks AI

More in Research