Skip to content
EN

RAG and retrieval

Chunking, embeddings, rerankers, evaluation and what holds up in production.

66 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: RAG and retrieval

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. GraphRAG for query-focused summarization

    The paper presents GraphRAG, an approach for answering questions over document collections. It addresses global questions about a corpus, which standard RAG retrieval does not handle well.

    It describes an approach for corpus-level questions that may not be answered by retrieving individual passages.

  2. Tencent Open-Sources WeKnora, an LLM Knowledge Platform

    WeKnora is an open-source framework for document understanding, semantic retrieval and autonomous reasoning. The post describes source-linked answers, file-generating skills, opt-in long-term memory and self-hosting.

    Engineers can inspect a RAG framework that combines document retrieval with agent skills and memory.

  3. ripwire builds structural repository context for coding agents

    ripwire is a zero-dependency C++23 CLI and MCP server that parses 21 languages with Tree-sitter and returns task-relevant symbols, relationships, and tests in a token-budgeted response without embeddings or a vector database.

    It offers an alternative to repeated repository searches for agents that need relevant code, potential change impact, and tests.

  4. Weixin Open-Sources WeMM-Embedding Multimodal Model

    Weixin says its Vision team developed WeMM-Embedding to support search and recommendations across text, images, and video. The 9B model is deployed across several Weixin services and ranked first on MMEB-v2 and MMEB-v3, according to the announcement.

    Engineers evaluating multimodal retrieval can inspect an embedding model used in production search and recommendations.

  5. Tencent Releases WeMM-Embedding Multimodal Models

    Tencent’s WeChat Vision team released WeMM-Embedding, a family of multimodal embedding models for understanding and retrieval. The post says the models support text, image, and video search and recommendations.

    Engineers can evaluate the open-source models for multimodal retrieval workloads.

  6. Tencent’s WeMM-Embedding-9B maps multiple modalities to a shared space

    Tencent’s WeMM-Embedding-9B maps text, images, videos, and visual documents into a shared space. The post says it achieves SOTA on MMEB-v2 and MMEB-v3.

    A shared embedding space may be relevant when building retrieval systems across text and visual content.

  7. Capacity Allocation in Hierarchical Search Agents

    The paper factorizes hierarchical search agents into delegation and execution roles and examines how model capacity should be distributed. The post says decomposition capacity is the key performance bottleneck.

    Engineers building search agents can consider whether delegation and execution need models of different capacities.

  8. NapMem treats long-term memory as an agent action space

    NapMem organizes user history into a linked pyramid of raw conversations, typed records, topic tracks, and user profiles. The agent learns to choose which memory granularity to inspect using memory-tool reinforcement learning.

    The approach offers an alternative to supplying agents with only evidence preselected by a retriever.

  9. SearchEyes trains multimodal search agents in simulated worlds

    SearchEyes uses a typed knowledge graph as the backbone of a simulated search world for training multimodal agents. The paper addresses disconnected training data, search environments, and reward signals in multi-hop reasoning.

    The approach connects search-world structure and training signals for multimodal search agents.

  10. NapMem treats long-term user memory as an action space

    The paper introduces NapMem, a framework that organizes user history into a linked, multi-granularity memory pyramid. It frames memory use as structured navigation rather than passive retrieval.

    Engineers evaluating conversational memory systems can compare active navigation with pre-selected retrieval.

  11. Theoretical Capacity of MaxSim Retrieval Models

    The paper studies the representation power of MaxSim and shows it can exactly replicate inner products between non-negative k-sparse vectors. The preview notes strong empirical performance for late-interaction models but limited prior theoretical understanding.

    It gives engineers a theoretical result for assessing MaxSim in late-interaction retrieval.

  12. CMDR benchmarks cross-page multimodal document retrieval

    The paper introduces CMDR and CMDR-Bench for retrieving relevant pages from multimodal documents, addressing queries that require context across multiple pages. The post also describes an embedding model for this task.

    It highlights how page-level retrieval can miss context needed to answer queries spanning a document.

  13. Relevance-Based Embeddings for Candidate Retrieval

    The paper describes representing queries and items with embeddings to retrieve candidates efficiently when the relevance function is expensive. The post says Yandex bases these representations on relevance to selected support items or queries.

    It may be useful for engineers exploring efficient candidate retrieval with expensive relevance models.

  14. A Study of In-Context Retrieval at Million-Token Scale

    The paper studies language models as in-context retrievers on million-token corpora and examines length generalization. It introduces BlockSearch, a 0.6B-parameter retriever that the post says generalizes up to 10 times beyond its training length.

    Engineers can compare in-context retrieval with vector-based retrieval at corpus scales practical systems face.

  15. Studying In-Context Retrieval at Million-Token Scale

    The paper studies whether language models can retrieve answers directly from in-context corpora, examining million-token corpora and length generalization. The post says it finds attention dilution and proposes length-aware fixes.

    It examines the limits of using long in-context corpora as an alternative to vector-based retrieval.

  16. LLM-Based Hard Negative Sampling for Two-Tower Retrieval

    The paper proposes a self-supervised hard negative sampling technique for two-tower recommendation models. It uses LLM-derived item clusters to generate harder negatives in real time.

    Harder training negatives may help engineers address a limitation of standard negative sampling in large-scale retrieval.

  17. Trie-based execution plans for IR pipeline experiments

    The paper presents a radix-trie execution plan for PyTerrier that reuses shared pipeline prefixes. The post reports experiment-time reductions of up to 26%.

    Reusing shared pipeline work may make retrieval experiments faster to run.

  18. Diffusion-GR2 speeds up generative reasoning reranking

    Diffusion-GR2 converts an autoregressive reasoning reranker into a block-diffusion model. The post reports near-AR ranking accuracy and 2.4–3.5× faster decoding.

    It explores parallel decoding as a way to reduce inference cost in reasoning-based rerankers.

  19. STEB benchmarks style text embeddings across 96 datasets

    STEB is an open-source benchmark for evaluating style embeddings across 96 datasets and 7 languages. The post says semantic embeddings underperform on stylistic tasks and no single model dominates.

    It offers a standardized way to evaluate embeddings for style-focused retrieval and related tasks.

  20. KbSD uses self-distillation to calibrate agentic search

    KbSD proposes a self-distillation framework for agentic search, where a hint-augmented teacher provides dense, token-level supervision. It targets decisions about using model memory, retrieved evidence, or abstaining.

    Engineers building retrieval agents can examine an approach to calibrating when a model relies on search or its own knowledge.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor