Skip to content
EN

RAG and retrieval

Chunking, embeddings, rerankers, evaluation and what holds up in production.

66 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: RAG and retrieval

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. CAMI Frames Multi-Index Construction as Budgeted Selection

    The CAMI paper presents a framework for treating multi-index construction as a budgeted, multi-objective portfolio selection problem. It addresses the cost of adding semantic enrichment indices to RAG ingestion pipelines.

    It offers a way to reason about retrieval quality and the computational cost of enrichment indices.

  2. RAG Retrieval Enhancements with a Strong Reranker

    The paper evaluates query expansion, summarization, graph-based expansion, routing, rank fusion, and corrective re-retrieval alongside a strong cross-encoder reranker. It examines whether reported benefits hold for mixed-format collections.

    It tests whether common retrieval additions provide reliable gains in collections that better reflect production data.

  3. As We May Search: A Local-First Design for Information Retrieval

    The paper proposes local-first information retrieval, with indexes, models, and inference on user devices and remote services treated as optional. It presents a framework for organizing retrieval architectures around privacy and other dimensions.

    It offers a design framework for engineers considering privacy-sensitive retrieval systems.

  4. SHIFT steers activations to mitigate knowledge conflicts in RAG

    SHIFT adds learnable gates to FFN activations to help LLMs balance retrieved context against parametric knowledge. The paper addresses conflicts between retrieved evidence and a model’s internal knowledge.

    It explores an internal model intervention for handling conflicting evidence in RAG.

  5. A Framework for Comparing RAG Approaches on Semi-Structured Data

    The paper introduces a framework to evaluate regular, Graph, Modular, and Agentic RAG on semi-structured knowledge bases. It provides implementations for nine standardized scenarios and reports a comparison.

    It offers a structured way to compare RAG variants across scenarios with data and domain restrictions.

  6. FlashMemory-DeepSeek-V4 Uses Sparse Retrieval for KV Cache

    The post describes a paper that predicts which past KV chunks a model will need, keeping selected chunks on GPU and offloading the rest. It reports a 13.5% average KV cache footprint and up to 90% memory reduction at 500K context.

    The approach may help engineers evaluating memory trade-offs for long-context inference.

  7. Paper studies geometric limits on RAG compression

    “The Geometry of Consolidation” is a NeurIPS 2026 paper about RAG compression and embedding geometry. Its authors claim a spectral quantity sets a hard floor on compression.

    Engineers evaluating embedding dimensions and RAG compression can inspect the paper’s analysis and claims.

  8. Skill-RAG routes retrieval based on detected knowledge gaps

    Skill-RAG uses hidden-state probing to detect potential knowledge failures and route queries to specialized retrieval strategies. The paper evaluates it on HotpotQA, Natural Questions, and TriviaQA.

    It explores how adaptive retrieval routing can address query-evidence misalignment rather than simply retrying retrieval.

  9. Binary quantization reduces embedding storage

    The post describes binary quantization, which maps each embedding dimension from a 32-bit float to one bit based on its sign. It says this reduces a 4,096-byte vector to 128 bytes.

    Lower embedding storage can matter when building and operating retrieval systems.

  10. Legal RAG Bench evaluates retrieval on criminal law questions

    The paper introduces a legal RAG benchmark with 4,876 criminal law passages and 100 complex questions drafted by legal experts. It reports that retrieval problems can contribute to fabricated answers.

    It gives engineers a domain-specific benchmark for evaluating retrieval in legal RAG systems.

  11. LightRetriever speeds up LLM-based text retrieval

    LightRetriever keeps LLM-based document processing offline and uses a simpler query encoder at inference time. The post reports 1,000× faster query encoding and 10× higher end-to-end throughput while retaining 95% of benchmark performance.

    The paper explores a way to reduce online inference costs in LLM-based retrieval.

  12. Querying 3 Billion Vectors

    An article about querying 3 billion vectors; its preview notes that requirements are hard.

    Useful as a starting point for thinking about the requirements behind large-scale vector retrieval.

  13. Post Claims ColBERTv2 Outperforms Qwen3-Embed-8B

    The post claims that 100M-parameter ColBERTv2 outperforms Qwen3-Embed-8B. A quoted post says a study found BM25 with passage-level retrieval and reranking remains competitive in deep research.

    It raises a retrieval-model comparison and points to BM25 plus reranking as a production-relevant baseline.

  14. How BM25 Scores Documents Without Embeddings

    The post explains BM25’s use of term rarity, diminishing returns for repeated terms, and document-length normalization to score search results. It says BM25 requires no training, embeddings, or fine-tuning.

    Understanding BM25 helps engineers assess a traditional lexical retrieval method alongside vector search.

  15. UltraRAG 3.0 builds RAG workflows with MCP servers and YAML

    UltraRAG is a lightweight framework for building RAG systems, designed for research. It packages core components as MCP servers and uses YAML to define workflows with sequential, loop, and conditional control structures.

    Engineers can use its standardized components and configurable workflows to build and debug complex RAG systems.

  16. Less LLM, More Documents: A RAG Scaling Approach

    The CMU researchers’ paper, “Less LLM, More Documents: Searching for Improved RAG,” studies scaling the document database instead of the language model. The post says a mid-sized model with a large corpus can match a larger model with a smaller corpus.

    It examines whether expanding a retrieval corpus can be an alternative to scaling up the LLM.

  17. UniversalRAG Routes Retrieval Across Modalities and Granularities

    UniversalRAG is a framework for retrieving knowledge from different modalities and at different granularities. It routes queries to a modality-specific corpus, then performs targeted retrieval within it.

    Engineers building multimodal RAG can assess an approach that avoids cross-modal comparisons through modality-aware routing.

  18. MegaRAG builds multimodal knowledge graphs for RAG

    MegaRAG presents a graph-based RAG method that automatically constructs knowledge graphs from visual documents to support cross-modal reasoning. The linked abstract describes knowledge graphs as a way to address limits in holistic understanding of long-form content.

    Engineers working with visual documents can assess a graph-based approach to retrieval and cross-modal reasoning.

  19. Iceberg benchmarks vector search on downstream tasks

    Researchers from Alibaba and partners introduce Iceberg, a suite of task-centric benchmarks across eight datasets. It examines vector similarity search beyond recall and latency, including its impact on RAG and image classification.

    It highlights how standard recall-latency benchmarks may miss effects that matter in downstream applications.

  20. RouteRAG uses reinforcement learning for adaptive text and graph retrieval

    RouteRAG is an RL-based framework for RAG over text and graphs. It trains a model to choose retrieval actions during reasoning and adds an efficiency reward to discourage unnecessary retrieval.

    Adaptive retrieval could help balance answer correctness against the cost of retrieving graph data.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor