RAG and retrieval
Chunking, embeddings, rerankers, evaluation and what holds up in production.
66 links, newest first.
- RAG and retrievalPaper
CAMI Frames Multi-Index Construction as Budgeted Selection
The CAMI paper presents a framework for treating multi-index construction as a budgeted, multi-objective portfolio selection problem. It addresses the cost of adding semantic enrichment indices to RAG ingestion pipelines.
It offers a way to reason about retrieval quality and the computational cost of enrichment indices.
- RAG and retrievalPaper
RAG Retrieval Enhancements with a Strong Reranker
The paper evaluates query expansion, summarization, graph-based expansion, routing, rank fusion, and corrective re-retrieval alongside a strong cross-encoder reranker. It examines whether reported benefits hold for mixed-format collections.
It tests whether common retrieval additions provide reliable gains in collections that better reflect production data.
- RAG and retrievalPaper
As We May Search: A Local-First Design for Information Retrieval
The paper proposes local-first information retrieval, with indexes, models, and inference on user devices and remote services treated as optional. It presents a framework for organizing retrieval architectures around privacy and other dimensions.
It offers a design framework for engineers considering privacy-sensitive retrieval systems.
- RAG and retrievalPaper
SHIFT steers activations to mitigate knowledge conflicts in RAG
SHIFT adds learnable gates to FFN activations to help LLMs balance retrieved context against parametric knowledge. The paper addresses conflicts between retrieved evidence and a model’s internal knowledge.
It explores an internal model intervention for handling conflicting evidence in RAG.
- RAG and retrievalPaper
A Framework for Comparing RAG Approaches on Semi-Structured Data
The paper introduces a framework to evaluate regular, Graph, Modular, and Agentic RAG on semi-structured knowledge bases. It provides implementations for nine standardized scenarios and reports a comparison.
It offers a structured way to compare RAG variants across scenarios with data and domain restrictions.
- RAG and retrievalPost on X
FlashMemory-DeepSeek-V4 Uses Sparse Retrieval for KV Cache
The post describes a paper that predicts which past KV chunks a model will need, keeping selected chunks on GPU and offloading the rest. It reports a 13.5% average KV cache footprint and up to 90% memory reduction at 500K context.
The approach may help engineers evaluating memory trade-offs for long-context inference.
- RAG and retrievalPaper
Paper studies geometric limits on RAG compression
“The Geometry of Consolidation” is a NeurIPS 2026 paper about RAG compression and embedding geometry. Its authors claim a spectral quantity sets a hard floor on compression.
Engineers evaluating embedding dimensions and RAG compression can inspect the paper’s analysis and claims.
- RAG and retrievalPaper
Skill-RAG routes retrieval based on detected knowledge gaps
Skill-RAG uses hidden-state probing to detect potential knowledge failures and route queries to specialized retrieval strategies. The paper evaluates it on HotpotQA, Natural Questions, and TriviaQA.
It explores how adaptive retrieval routing can address query-evidence misalignment rather than simply retrying retrieval.
- RAG and retrievalPost on X
Binary quantization reduces embedding storage
The post describes binary quantization, which maps each embedding dimension from a 32-bit float to one bit based on its sign. It says this reduces a 4,096-byte vector to 128 bytes.
Lower embedding storage can matter when building and operating retrieval systems.
- RAG and retrievalPost on X
Legal RAG Bench evaluates retrieval on criminal law questions
The paper introduces a legal RAG benchmark with 4,876 criminal law passages and 100 complex questions drafted by legal experts. It reports that retrieval problems can contribute to fabricated answers.
It gives engineers a domain-specific benchmark for evaluating retrieval in legal RAG systems.
- RAG and retrievalPaper
LightRetriever speeds up LLM-based text retrieval
LightRetriever keeps LLM-based document processing offline and uses a simpler query encoder at inference time. The post reports 1,000× faster query encoding and 10× higher end-to-end throughput while retaining 95% of benchmark performance.
The paper explores a way to reduce online inference costs in LLM-based retrieval.
- RAG and retrievalArticle
Querying 3 Billion Vectors
An article about querying 3 billion vectors; its preview notes that requirements are hard.
Useful as a starting point for thinking about the requirements behind large-scale vector retrieval.
- RAG and retrievalPost on X
Post Claims ColBERTv2 Outperforms Qwen3-Embed-8B
The post claims that 100M-parameter ColBERTv2 outperforms Qwen3-Embed-8B. A quoted post says a study found BM25 with passage-level retrieval and reranking remains competitive in deep research.
It raises a retrieval-model comparison and points to BM25 plus reranking as a production-relevant baseline.
- RAG and retrievalPost on X
How BM25 Scores Documents Without Embeddings
The post explains BM25’s use of term rarity, diminishing returns for repeated terms, and document-length normalization to score search results. It says BM25 requires no training, embeddings, or fine-tuning.
Understanding BM25 helps engineers assess a traditional lexical retrieval method alongside vector search.
- RAG and retrievalArticle
UltraRAG 3.0 builds RAG workflows with MCP servers and YAML
UltraRAG is a lightweight framework for building RAG systems, designed for research. It packages core components as MCP servers and uses YAML to define workflows with sequential, loop, and conditional control structures.
Engineers can use its standardized components and configurable workflows to build and debug complex RAG systems.
- RAG and retrievalPaper
Less LLM, More Documents: A RAG Scaling Approach
The CMU researchers’ paper, “Less LLM, More Documents: Searching for Improved RAG,” studies scaling the document database instead of the language model. The post says a mid-sized model with a large corpus can match a larger model with a smaller corpus.
It examines whether expanding a retrieval corpus can be an alternative to scaling up the LLM.
- RAG and retrievalPaper
UniversalRAG Routes Retrieval Across Modalities and Granularities
UniversalRAG is a framework for retrieving knowledge from different modalities and at different granularities. It routes queries to a modality-specific corpus, then performs targeted retrieval within it.
Engineers building multimodal RAG can assess an approach that avoids cross-modal comparisons through modality-aware routing.
- RAG and retrievalPaper
MegaRAG builds multimodal knowledge graphs for RAG
MegaRAG presents a graph-based RAG method that automatically constructs knowledge graphs from visual documents to support cross-modal reasoning. The linked abstract describes knowledge graphs as a way to address limits in holistic understanding of long-form content.
Engineers working with visual documents can assess a graph-based approach to retrieval and cross-modal reasoning.
- RAG and retrievalPost on X
Iceberg benchmarks vector search on downstream tasks
Researchers from Alibaba and partners introduce Iceberg, a suite of task-centric benchmarks across eight datasets. It examines vector similarity search beyond recall and latency, including its impact on RAG and image classification.
It highlights how standard recall-latency benchmarks may miss effects that matter in downstream applications.
- RAG and retrievalPaper
RouteRAG uses reinforcement learning for adaptive text and graph retrieval
RouteRAG is an RL-based framework for RAG over text and graphs. It trains a model to choose retrieval actions during reasoning and adds an efficiency reward to discourage unnecessary retrieval.
Adaptive retrieval could help balance answer correctness against the cost of retrieving graph data.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor
