We demonstrate 210x average speedup over linear semantic memory search by projecting high-dimensional embeddings into 3D geometric space and using BVH traversal — the algorithm executed by idle RT Cores in consumer GPUs. Tested on real sentence embeddings with 48% Recall@10, improving to near-100% with a re-ranking step. Results obtained on an RTX 3070 using CPU KDTree as RT Core proxy. Conservative hardware estimates suggest 1,000x-3,000x speedup with true RT Core traversal.
Mohamed Mahmoud Ahmed Mohamed (2026) studied this question.