Skip to content
EN

RAG and retrieval

Chunking, embeddings, rerankers, evaluation and what holds up in production.

66 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: RAG and retrieval

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. MTEB v2 updates the embedding evaluation suite

    The MTEB team released version 2 of its evaluation suite for embedding models. The accompanying blog post covers easier evaluation, multimodal support, rerankers, new interfaces, documentation, dataset statistics, and a migration guide.

    Engineers evaluating embedding models can review the suite’s updated capabilities and migration guidance.

  2. Open-source RLM implementation for long-context processing

    A Python implementation of Recursive Language Models stores long input in a Python REPL for analysis. The project says it supports 100+ LLM providers through LiteLLM.

    Engineers evaluating alternatives to RAG for processing very long documents can inspect the implementation.

  3. Information Retrieval Papers: Generative Embeddings and Document Retrieval

    Vol. 124 of a weekly information retrieval paper roundup, covering generative embeddings with test-time scaling and efficient visual document retrievers.

    A paper roundup can help engineers track research relevant to retrieval systems.

  4. REFRAG Replaces Retrieved Tokens with Reusable Chunk Embeddings

    A paper introduces REFRAG, which replaces most retrieved tokens with precomputed, reusable chunk embeddings. The post says it achieves 30× speed and supports 16× longer contexts without accuracy loss.

    It describes a retrieval change aimed at improving RAG speed and context capacity.

  5. Curated resources on GraphRAG

    Awesome-GraphRAG is a curated list of surveys, papers, benchmarks, and open-source projects on graph-based retrieval-augmented generation.

    It offers engineers a starting point for exploring research, benchmarks, and projects focused on GraphRAG.

  6. LogicRAG Builds a Reasoning Graph at Query Time

    LogicRAG plans a per-query DAG of dependent subproblems, retrieves evidence in order, and carries forward a rolling summary. The post says it outperforms GraphRAG systems on three multi-hop QA sets without offline graph building.

    The approach offers an alternative to prebuilt knowledge graphs for multi-hop retrieval.

  7. Survey Reviews Memory Mechanisms in LLM-Based Agents

    A post describes a survey by Renmin University of China and Huawei covering memory design and evaluation, applications, limitations, and future directions for LLM-based agents.

    Its overview of memory design and evaluation may inform engineers building agents for long-term tasks.

  8. VideoRAG retrieves videos for multimodal RAG

    VideoRAG is a framework proposed by research teams from KAIST for retrieving query-relevant videos and incorporating visual and textual information into generated responses. It uses automatic speech recognition to create auxiliary text when videos lack subtitles.

    It offers an approach to bringing video content into RAG retrieval and response generation.

  9. Graphiti builds temporal knowledge graphs for evolving data

    Graphiti builds and queries dynamic knowledge graphs that represent relationships between entities over time. It handles temporal data and supports multiple search methods.

    Engineers exploring retrieval beyond conventional document search may find its temporal knowledge-graph approach relevant.

  10. Auto-RAG Uses Iterative Retrieval for RAG

    The post describes Auto-RAG, a fine-tuned LLM that plans retrievals and refines queries through multiple turns until it has enough external information. It says the method adjusts the number of iterations based on question difficulty.

    It highlights an approach to automating retrieval planning and query refinement in RAG systems.

  11. RARe Uses Retrieved Examples to Adapt Retrieval Models

    The paper proposes RARe, which augments queries with task instructions and semantically similar query-document examples retrieved with BM25. The post reports that similar examples outperform random ones, with gains as the number of examples increases up to 10 tested.

    It explores a way to apply in-context examples to retrieval models and reports retrieval performance changes.

  12. HijackRAG Uses Injected Text to Target RAG Systems

    The post describes HijackRAG, an attack that combines retrieval, attention-hijacking, and instruction text to manipulate RAG outputs. It outlines black-box and gradient-based white-box modes.

    The attack highlights risks from malicious content in a RAG knowledge base and the limits of existing defenses.

  13. AssistRAG separates memory and retrieval from answer generation

    AssistRAG uses a trainable assistant LLM for memory and external-knowledge management while keeping the main answer-generating LLM frozen. Its training combines Curriculum Assistant Learning and Reinforced Preference Optimization.

    The design offers an approach to adding memory and knowledge management without retraining the main LLM.

  14. LazyGraphRAG combines vector and graph retrieval

    The post describes LazyGraphRAG as combining vector RAG with iterative breadth-first and best-first search, while avoiding index summarization. It says the approach targets one-off queries, exploratory analysis, and streaming data.

    Its retrieval and indexing trade-offs may help engineers assess alternatives to full GraphRAG.

  15. Notebook: RAG over PDFs with Docling and Weaviate

    A recipe notebook demonstrates RAG over PDF files using Docling and Weaviate. Docling is described as an open-source Python package that converts documents including PDFs to Markdown or JSON.

    Useful as a practical starting point for building a PDF-based RAG pipeline with Docling and Weaviate.

  16. Autoflow is a Graph RAG knowledge base built with TiDB

    The post describes pingcap/autoflow as a conversational knowledge base based on Graph RAG, using TiDB Serverless for vector storage.

    Engineers evaluating Graph RAG tools can note its use of TiDB Serverless as vector storage.

  17. OpenScholar Releases Code, Models, and a 45M-Paper Datastore

    OpenScholar releases code and model checkpoints, an OpenScholar Datastore with 45M+ papers through October 2024, and ScholarQABench. Its project page describes a research assistant for finding, summarizing, and analyzing scientific evidence.

    Engineers can inspect an open research-assistant stack and its paper datastore, models, and evaluation benchmark.

  18. KAR Uses Knowledge Graphs for Query Expansion

    KAR parses query entities with an LLM, retrieves relevant documents, propagates through knowledge-graph relations, scores and filters those relations, and generates knowledge-aware query expansions.

    The workflow offers a concrete approach to expanding retrieval queries with document context and knowledge-graph relations.

  19. SimpleQA: A Benchmark for Factuality and Hallucinations

    SimpleQA contains 4,000 human-written fact-seeking questions with single, indisputable answers. An autograder labels responses correct, incorrect, or not attempted; reference answers were verified by two annotators.

    Engineers can use it to evaluate model factuality with a human-written question set and verified reference answers.

  20. FACT uses iterative context rewriting for multi-fact retrieval

    The paper introduces Find All Crucial Texts (FACT), an iterative approach to context rewriting for multi-fact retrieval. It examines how models can lose track of critical information while generating from extended contexts.

    It explores a retrieval failure mode that can lead to incomplete answers when a task requires combining multiple facts.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor