Table of Contents
- The Problem with Small Chunks
- The Opportunity with Long-Context Models
- LongRAG Architecture
- Redesigning the Chunking Strategy
- Retrieval Efficiency Considerations
- Comparison with Traditional RAG
- Implementation Guide
- Use Cases
- Limitations and Challenges
- What Later Research Says
- Summary
- Further Reading
- Changelog
- References
๐ ไธญๆ็
The core assumption of traditional RAG is that LLM context windows are limited, so we must split documents into small pieces and only feed the most relevant fragments into the model.
This assumption was reasonable before 2024. But mainstream commercial models moved past the 100K-token mark long ago (the exact numbers shift every quarter โ check each vendor's official model page). When usable context grows by two orders of magnitude, RAG's design math deserves to be redone.
LongRAG is the embodiment of this thinking: Stop splitting documents into fragments โ retrieve larger units and let long-context models do the understanding.
Up front: this article describes a tradeoff, not a settled winner. In the year-plus since the LongRAG paper, the community has argued hard over whether long context makes RAG obsolete, and the evidence points both ways โ the "What Later Research Says" section near the end lays out the counter-evidence.
The Problem with Small Chunks
Traditional RAG typically splits documents into small chunks of 256โ512 tokens. This approach seems reasonable, but in practice it creates a series of problems.
Information Fragmentation
A complete argument gets split across multiple chunks, with each chunk containing only partial information.
Consider this example:
Original paragraph:
"Article 12 of the contract stipulates that Party A must complete payment
within 30 business days of receiving the acceptance report. If Party A
delays payment, a penalty of 0.05% per day shall be applied. However,
if the delay is caused by Party B's failure to provide complete acceptance
documents, Party A bears no liability for the delay."
Split with a 128-token chunk size:
Chunk 1: "Article 12 of the contract stipulates that Party A must complete payment within 30 business days of receiving the acceptance report."
Chunk 2: "If Party A delays payment, a penalty of 0.05% per day shall be applied."
Chunk 3: "However, if the delay is caused by Party B's failure to provide complete acceptance documents, Party A bears no liability for the delay."
The user asks: "How much penalty does Party A pay for late payment?"
The retrieval system might only find Chunk 2. But the correct answer requires Chunk 2 + Chunk 3 โ because there's an exception clause. If only Chunk 2 is retrieved, the LLM will produce an answer that appears correct but is incomplete.
This is information fragmentation: each chunk's semantics are incomplete and must be combined with other chunks to reconstruct the full meaning.
Boundary Context Loss
Chunk boundaries are often arbitrary. A chain of reasoning that crosses a boundary gets broken:
End of Chunk A: "...therefore, we adopted the Transformer architecture. Specifically,"
Start of Chunk B: "we used a 6-layer encoder with Rotary Position Encoding (RoPE),"
Chunk B lacks the context of "why we adopted Transformer." If the user asks "Why was this architecture chosen?", Chunk B alone cannot answer the question.
Over-reliance on Retrieval Precision
Small chunks mean a large number of candidate fragments. A 100,000-character Chinese document is roughly 150,000 tokens, which at 512 tokens per chunk produces about 300 chunks. The retrieval system must precisely find the 3โ5 most relevant ones out of several hundred candidates.
This places extremely high demands on retrieval:
- Embeddings must accurately capture the semantics of each small chunk
- Ranking must be precise, because the difference between top-3 and top-10 could mean the difference between "has the answer" and "doesn't have the answer"
- Multi-hop reasoning (answers scattered across multiple chunks) is nearly impossible to handle well
Small chunks transfer all the "understanding" pressure onto "retrieval." And retrieval is never perfect.
Uneven Semantic Density
Different paragraphs have vastly different semantic densities. A 512-token legal clause might contain 5 important points, while a 512-token background introduction might contain only 1. Fixed-size chunks cannot reflect this variation.
The Opportunity with Long-Context Models
Between 2023 and 2025, mainstream LLM context windows went from a few thousand tokens to hundreds of thousands, even millions. No model comparison table here โ that kind of table expires in three months, and a vendor's stated maximum and the length at which a model still performs well are two different numbers. For current figures, check the vendor model docs; for whether you can actually use that much, see "What Later Research Says" below.
This changes the fundamental tradeoff in RAG:
Before: Context windows are scarce resources โ must retrieve precisely โ small chunks โ high retrieval pressure.
Now: Context windows are abundant โ can include more content โ large chunks โ shift pressure from retrieval to comprehension.
Key Insight
Long-context models excel at finding relevant information within large amounts of text. Needle-in-a-haystack (NIAH) tests show that even within a 100K-token context, good models can accurately locate specific facts embedded within it.
This means: We don't need perfect retrieval โ just good-enough retrieval. Throw roughly relevant content at the LLM and let it find the answer itself.
That insight has since been heavily qualified. Part of why NIAH looks so good is that the needle and the question share literal wording, letting models shortcut via string matching. The NoLiMa benchmark (2025) removed that lexical overlap and required models to locate the needle by latent association instead: of 13 models claiming 128K+ support, 11 dropped below half their short-context baseline at just 32K, and even GPT-4o โ one of the best performers โ fell from an almost-perfect 99.3% to 69.7%. So the range in which "good-enough retrieval is enough" holds is narrower than NIAH scores suggest.
From Precise Retrieval to Coarse-Grained Retrieval
Traditional RAG mindset: "Only give the LLM the most relevant fragments; don't waste tokens."
LongRAG mindset: "Give the LLM enough context and let it decide what's relevant."
This isn't a regression โ it's leveraging advances in model capability. Rather than investing heavily in perfecting retrieval (better embeddings, more precise reranking, smarter query expansion), leverage the model's own comprehension ability.
LongRAG Architecture
LongRAG's core design is simple: increase retrieval granularity.
Traditional RAG vs LongRAG
Traditional RAG:
โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โ Document A โ
โ Query โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโ... โ
โ โ โ โ c1 โโ c2 โโ c3 โโ c4 โ โ
โ โ โ โ512t โโ512t โโ512t โโ512t โ โ
โ โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โโโโโโฌโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โ Vector search From hundreds of chunks
โ (top-k=5) precisely find 5
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ c2 โโ c17 โโ c43 โโ c8 โโ c91 โ โ May miss key fragments
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ Feed to LLM (~2,500 tokens)
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ LLM answers from fragmented โ
โ context โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
LongRAG:
โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โ Document A โ
โ Query โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ โ โ Section 1 โโ Section 2 โ โ
โ โ โ โ (~6,000t) โโ (~8,000t) โ โ
โ โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โโโโโโฌโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โ Vector search From a small number of
โ (top-k=3) segments, find roughly
โผ relevant ones
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Section 2 โโ Section 5 โโ Section 11 โ โ Complete semantic units
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ Feed to LLM (~20,000 tokens)
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ LLM answers from complete โ
โ context โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The fundamental differences:
- Traditional RAG: Hundreds of small chunks โ precise retrieval โ fragmented context โ LLM pieces together an answer
- LongRAG: Dozens of large segments โ coarse-grained retrieval โ complete context โ LLM directly comprehends
Choosing Retrieval Units
LongRAG's "large chunks" aren't just small chunks stitched together arbitrarily. It uses the document's own structure as splitting boundaries:
| Retrieval Unit | Size Range | Use Case |
|---|---|---|
| Paragraph groups | 1,000โ3,000 tokens | Documents with unclear structure |
| Sections | 3,000โ10,000 tokens | Documents with heading structure |
| Sub-documents | 10,000โ30,000 tokens | Independent chapters in long documents |
| Entire documents | 30,000โ100,000 tokens | Short to medium-length standalone documents |
Key principle: Splitting boundaries should align with semantic boundaries, not fixed token counts.
Redesigning the Chunking Strategy
LongRAG isn't about "not chunking" โ it's about "chunking smarter." Here are three main strategies.
Strategy 1: Document-Level Retrieval
The most extreme approach: each document is a single retrieval unit.
Index structure:
Document A โ one embedding (representing the entire document's semantics)
Document B โ one embedding
Document C โ one embedding
...
Retrieval: find the 1โ3 most relevant documents, feed full text to LLM.
Pros:
- No chunking needed at all
- Zero information fragmentation
- Simplest implementation
Cons:
- A single document may exceed the context window
- One embedding struggles to represent all topics in a long document
- Lowest retrieval precision (document-level is too coarse)
Best for:
- Document collections where each document is relatively short (< 30K tokens)
- Documents with a single topic that don't cover multiple unrelated subjects
- Small document collections (< 1,000 documents)
Strategy 2: Section-Level Retrieval
Split by the document's section structure, with each section as a retrieval unit.
Index structure:
Document A / Section 1 โ one embedding
Document A / Section 2 โ one embedding
Document A / Section 3 โ one embedding
Document B / Section 1 โ one embedding
...
Retrieval: find the 2โ5 most relevant sections, combine and feed to LLM.
Pros:
- Preserves semantic completeness (sections are typically complete units of discourse)
- Moderate retrieval granularity, more precise than document-level
- Leverages the document's own structure โ no manual judgment needed for split points
Cons:
- Requires documents with clear section structure (headings, table of contents, etc.)
- Section sizes can vary dramatically (some 500 tokens, others 20,000 tokens)
- Cross-section references are still lost
Best for:
- Technical documentation, academic papers, legal documents
- Documents with Markdown/HTML heading structure
- Documents with multiple independently understandable topics
Strategy 3: Hybrid Strategy (Coarse Retrieval + Fine Reading)
This is LongRAG's recommended approach: a two-stage architecture.
Stage 1 - Coarse-grained retrieval:
Use section/document-level embeddings to find roughly relevant segments
Stage 2 - Fine-grained reading:
LLM finds precise answers within the large segments
Equivalent to:
Traditional RAG's "retriever + reader"
But the retriever is coarser, and the reader is stronger
Process:
- During indexing, split at the section level (3,000โ10,000 tokens)
- During retrieval, take top-k=3โ5 sections (roughly 15,000โ50,000 tokens total)
- Combine these sections into a single large context and feed to a long-context LLM
- The LLM finds precise answers within the large context
Pros:
- Combines high recall from coarse retrieval with precise comprehension from the LLM
- Doesn't require perfect retrieval (the LLM filters out irrelevant content for you)
- Highly adaptable โ different documents can use different granularities
Strategy Comparison
| Dimension | Document-Level | Section-Level | Hybrid Strategy |
|---|---|---|---|
| Splitting granularity | Entire document | 3Kโ10K tokens | 3Kโ10K tokens |
| Index size | Smallest | Medium | Medium |
| Retrieval precision | Low | Medium | Medium + LLM compensation |
| Context completeness | Highest | High | High |
| Token consumption | Highest | Medium-high | Medium-high |
| Implementation complexity | Lowest | Medium | Medium-high |
| Required model | 1M context | 100K+ context | 100K+ context |
Retrieval Efficiency Considerations
Large chunks change the performance characteristics of the retrieval system.
Dramatically Fewer Candidates
For the same corpus:
Traditional RAG (512 tokens/chunk):
100K-character document ร 100 documents = ~30,000 chunks
Vector search space: 30,000 vectors
LongRAG (section-level, ~5,000 tokens/section):
100K-character document ร 100 documents = ~8,000 sections
Vector search space: 8,000 vectors
A 10x reduction in candidates directly provides:
- Faster vector search: ANN algorithms like HNSW are faster on smaller indexes
- Lower storage costs: Fewer vectors = less memory and disk space
- Simpler index maintenance: Fewer vectors to update when adding/removing documents
Precision vs. Recall Tradeoff
Large chunks naturally have higher recall (each chunk covers more content), but precision may decrease (more irrelevant content within each chunk).
Small chunks (512 tokens):
โ High precision โ each chunk is topic-focused
โ Low recall โ answer may be split across adjacent chunks
โ Requires more precise retrieval
Large chunks (5,000 tokens):
โ Lower precision โ chunks may contain irrelevant content
โ High recall โ answer is more likely to be within a selected chunk
โ High fault tolerance โ can find answers even with imperfect retrieval
Balancing Strategies
Several methods to reduce precision loss with large chunks:
1. Multi-level Indexing
Maintain both section-level and paragraph-level embeddings simultaneously. Use section-level retrieval to narrow the scope first, then use paragraph-level for ranking.
2. Summary Embeddings
Instead of embedding the entire section text, first use an LLM to generate a section summary, then embed the summary. Summaries are more condensed with higher semantic density, resulting in better embedding quality.
3. Multi-vector Representation
Generate multiple embeddings for a single section (e.g., section summary + first sentence of each paragraph within the section). A hit on any of them counts as that section being relevant.
4. Post-retrieval Reranking
After retrieving large chunks, use a cross-encoder or LLM to perform secondary ranking on paragraphs within the chunk, prioritizing the most relevant portions.
Comparison with Traditional RAG
Here's a detailed comparison between LongRAG and traditional RAG across multiple dimensions:
| Dimension | Traditional RAG | LongRAG |
|---|---|---|
| Chunk size | 256โ512 tokens | 3,000โ100,000 tokens |
| Vectors in index | Many (hundreds per document) | Few (single digits to dozens per document) |
| Retrieval precision | High (each chunk is topic-focused) | Medium (mixed topics within chunks) |
| Context coherence | Low (fragmented) | High (complete semantic units) |
| Token consumption/query | Low (~2Kโ5K tokens) | High (~15Kโ50K tokens) |
| Inference latency | Lower | Higher (more tokens to process) |
| Retrieval latency | Higher (large index) | Lower (small index) |
| Multi-hop reasoning | Weak (requires multiple retrievals) | Strong (large context naturally covers it) |
| Dependence on retrieval quality | Extremely high | Moderate |
| LLM requirements | Any model | Requires long-context model |
| Cost per query | Lower | Higher |
| Build complexity | High (needs fine-grained chunking + reranking) | Low (coarse-grained splitting suffices) |
Concrete Performance Data
Let's be precise about what the paper actually reports, because this set of numbers gets cited wrong a lot.
Retrieval side (NQ / HotpotQA, Wikipedia):
LongRAG groups related documents into 4K-token retrieval units, cutting the NQ corpus from 22M units to 600K. The paper's own words: answer recall@1 improves from 52% (DPR) to 71%. For HotpotQA, the corpus drops from 5M to 500K units and recall@2 goes from 47% to 72%.
Note that these are recall@1 / recall@2, not "top-5 versus top-4 hit rates" โ the two are frequently conflated. The claim is "strong retrieval performance with only a few (fewer than 8) top units," not "large chunks beat small chunks at the same top-k."
Generation side:
Without any training, LongRAG reaches 62.7 EM on NQ and 64.3 EM on HotpotQA, which the paper describes as on par with fully fine-tuned SoTA models. The main table's reader was GPT-4o at the time (the paper compares six readers). On non-Wikipedia datasets, Qasper F1 goes from 22.5% to 25.9% and MultiFieldQA-en from 51.2% to 57.5%.
The key driver is improved recall: small-chunk retrieval frequently misses the fragment containing the answer, while large units are more likely to encompass it.
Cost Analysis
No price table here โ per-token API pricing moves too fast for any hardcoded number to survive. For real figures, check the official pricing page.
What is stable is the ratio:
Traditional RAG per query:
Retrieval: 5 chunks ร 512 tokens โ 2,600 input tokens
LongRAG per query:
Retrieval: 3 sections ร 6,000 tokens โ 18,000 input tokens
Roughly 7x the input tokens.
Since the overwhelming majority of RAG query cost is driven by input tokens (output is usually only a few hundred), that ratio is roughly the cost ratio. Two things shrink the gap in practice: prompt caching (when the same sections are hit repeatedly, cache reads are far cheaper than full-price input) and batch discounts. If your traffic pattern can exploit either, measure before concluding.
Whether the tradeoff is worthwhile depends on the use case. For high-value queries in legal, medical, or financial domains, paying several times more for more accurate and complete answers is usually reasonable. For low-value, high-frequency everyday Q&A, traditional RAG is clearly more economical.
Implementation Guide
Here's a complete LongRAG retrieval pipeline implementation in TypeScript.
Section-Level Splitter
interface Section {
id: string;
documentId: string;
title: string;
content: string;
tokenCount: number;
embedding: number[];
metadata: { level: number; position: number };
}
/**
* Split documents by section structure, not fixed token counts.
* Sections exceeding maxTokens are recursively split at the next heading level;
* adjacent sections that are too small are automatically merged.
*/
function splitBySection(
document: { id: string; content: string },
maxTokens = 8000,
minTokens = 500,
): Section[] {
const lines = document.content.split('\n');
const sections: Section[] = [];
let currentTitle = '';
let currentContent: string[] = [];
let level = 1;
let idx = 0;
function flush() {
if (!currentContent.length) return;
const content = currentContent.join('\n');
const tokens = estimateTokens(content);
sections.push({
id: `${document.id}_s${idx}`,
documentId: document.id,
title: currentTitle,
content,
tokenCount: tokens,
embedding: [],
metadata: { level, position: idx++ },
});
currentContent = [];
}
for (const line of lines) {
const m = line.match(/^(#{1,6})\s+(.+)/);
if (m) {
flush();
level = m[1].length;
currentTitle = m[2].trim();
currentContent = [line];
} else {
currentContent.push(line);
}
}
flush();
// Merge adjacent sections that are too small
return sections.reduce<Section[]>((merged, section) => {
const prev = merged[merged.length - 1];
if (prev && prev.tokenCount < minTokens) {
prev.content += '\n\n' + section.content;
prev.tokenCount += section.tokenCount;
} else {
merged.push({ ...section });
}
return merged;
}, []);
}
function estimateTokens(text: string): number {
const zh = (text.match(/[ไธ-้ฟฟ]/g) || []).length;
const en = text.replace(/[ไธ-้ฟฟ]/g, '').split(/\s+/).length;
return Math.ceil(zh * 1.5 + en * 1.3);
}
LongRAG Retrieval Pipeline
import { cosineSimilarity, generateEmbedding } from './utils';
interface RetrievalResult {
sections: Section[];
totalTokens: number;
query: string;
}
interface LongRAGConfig {
topK: number; // How many sections to retrieve
maxContextTokens: number; // Maximum context tokens
minRelevanceScore: number; // Minimum relevance score
}
/**
* Core retrieval logic for LongRAG.
* Key differences from traditional RAG:
* 1. Retrieval units are sections (thousands of tokens), not small chunks (hundreds of tokens)
* 2. top-k is smaller (3-5), because each result is already large
* 3. Has token budget control to avoid exceeding the LLM's context window
*/
async function retrieveSections(
query: string,
index: Section[], // Each section already has an embedding
config: LongRAGConfig = { topK: 5, maxContextTokens: 50000, minRelevanceScore: 0.3 },
): Promise<RetrievalResult> {
const queryEmbedding = await generateEmbedding(query);
// Compute similarity โ filter โ sort
const scored = index
.map((section) => ({
section,
score: cosineSimilarity(queryEmbedding, section.embedding),
}))
.filter((item) => item.score >= config.minRelevanceScore)
.sort((a, b) => b.score - a.score);
// Select top-k sections within token budget
const selected: Section[] = [];
let totalTokens = 0;
for (const item of scored) {
if (selected.length >= config.topK) break;
if (totalTokens + item.section.tokenCount > config.maxContextTokens) continue;
selected.push(item.section);
totalTokens += item.section.tokenCount;
}
// Sort by original position (preserve reading order)
selected.sort((a, b) => {
if (a.documentId !== b.documentId) return a.documentId.localeCompare(b.documentId);
return a.metadata.position - b.metadata.position;
});
return { sections: selected, totalTokens, query };
}
Complete Pipeline: Retrieval + Generation
/**
* Full LongRAG pipeline: split โ index โ retrieve โ generate answer
*/
async function longRAGPipeline(
documents: { id: string; content: string }[],
query: string,
) {
// Step 1: Section splitting
const allSections = documents.flatMap((doc) => splitBySection(doc, 8000, 500));
// Step 2: Retrieval (assuming sections already have embeddings)
const retrieval = await retrieveSections(query, allSections, {
topK: 4, maxContextTokens: 40000, minRelevanceScore: 0.25,
});
// Step 3: Assemble context and generate answer
const context = retrieval.sections
.map((s, i) => `=== Source ${i + 1}: ${s.title} ===\n${s.content}`)
.join('\n---\n\n');
const answer = await callLLM({
// Inject the model ID from config โ don't hardcode it; it will expire
model: process.env.LONG_CONTEXT_MODEL!,
system: `Answer the question based on the reference materials. Cite source numbers, and clearly state when information is insufficient.`,
messages: [{ role: 'user', content: `Reference materials:\n${context}\n\nQuestion: ${query}` }],
maxTokens: 1000,
});
return {
answer,
sources: retrieval.sections.map((s) => ({ documentId: s.documentId, title: s.title })),
tokensUsed: retrieval.totalTokens,
};
}
Use Cases
LongRAG isn't a silver bullet. Here are the scenarios where it particularly excels and where it doesn't.
Particularly Well-Suited
1. Lengthy Legal Documents
Legal texts are characterized by extensive cross-references between clauses, where the meaning of one clause depends on the context of others. Traditional RAG breaks all these cross-references when it splits contracts into small chunks. LongRAG preserves entire sections or even entire contracts, letting the LLM see the complete relationships between clauses.
User asks: "Can the owner terminate the contract if the contractor delays delivery?"
Traditional RAG might only find:
"If delivery is delayed by more than 30 days, the owner has the right to terminate the contract."
LongRAG would find the entire "Termination Clause" section, including:
- Definition of delivery delay
- 30-day grace period
- Force majeure exceptions
- Written notice requirement before termination
- Post-termination settlement procedures
2. Academic Papers
A paper's methodology, experimental design, and results analysis are typically spread across different sections but are closely related. LongRAG can retrieve the entire "Method + Experiments" segment at once, letting the LLM understand the causal relationship between methods and results.
3. Technical Manuals and API Documentation
Technical concepts typically require complete context to understand. An API endpoint's behavior may depend on authentication settings, rate limit policies, error code definitions, and other information scattered across different sections. Large chunks make it easier to retrieve all this dispersed information at once.
4. Multi-Hop Reasoning Queries
When an answer requires synthesizing information from multiple paragraphs, LongRAG has a natural advantage:
Question: "Under what circumstances can the company not pay year-end bonuses?"
Requires synthesizing:
- Year-end bonus calculation rules (Chapter 4)
- Employee evaluation criteria (Chapter 7)
- Special exception clauses (Chapter 12)
LongRAG is more likely to retrieve all three chapters.
5. Scenarios Where Context Coherence Matters More Than Precision
Customer service knowledge bases, product FAQs, policy manuals โ in these scenarios, providing users with a complete, coherent answer matters more than precisely citing a specific paragraph. LongRAG's large context enables the LLM to generate smoother, more complete responses.
Less Suitable
1. Precise Fact Lookups in Very Large Corpora
If your corpus contains millions of documents and users are simply looking for a specific number or date, traditional RAG's small chunks + precise retrieval is more efficient. LongRAG would consume a large amount of unnecessary tokens in this scenario.
2. Low-Latency Requirements
LongRAG feeds roughly 7โ8x more tokens to the LLM than traditional RAG (the configuration above works out to 7x; it moves with your chunk sizes), with inference latency increasing proportionally. For scenarios requiring millisecond-level responses (such as real-time search suggestions), this may be unacceptable.
3. Cost-Sensitive High-Frequency Queries
Consuming 15Kโ50K input tokens per query, with tens of thousands of queries per day, makes token costs very significant.
Limitations and Challenges
1. Requires Long-Context LLMs
LongRAG's prerequisite is that the LLM can handle a large number of input tokens. If your model only has an 8Kโ16K context window, LongRAG's large chunks simply won't fit.
Today most commercial APIs and newer open-weight models claim 100K+ support, so this constraint is far looser than it was in 2024. But claiming support is not the same as using it well โ see the NoLiMa results above, and the Databricks study across 20 models, which found only a handful of the most recent models maintain consistent accuracy above 64K. The real limit isn't whether the tokens fit; it's whether the model still finds them.
2. Token Costs Scale Linearly
More input tokens directly means higher API costs, and the scaling is linear. Again, no unit prices (they expire) โ just orders of magnitude: at 10,000 queries per day, traditional RAG at ~3K input tokens per query versus LongRAG at ~25K is 30M versus 250M tokens per day. Multiply by whatever your current rate is and the annualized gap is large.
At high traffic volumes this is impossible to ignore, and it is exactly why routing approaches like Self-Route (below) โ keeping the cheap path for easy questions โ earn their keep.
3. Increased Inference Latency
LLM inference time is roughly proportional to input tokens. Processing 25K tokens takes approximately 5โ8x longer than 3K tokens. In user-experience-sensitive scenarios (such as chatbots), this latency may be unacceptable.
Mitigation strategies:
- Use streaming responses so the time-to-first-token remains unchanged
- Cache results for popular queries
- Fall back to traditional RAG for simple queries that don't need long context
4. Embedding Quality Degrades with Text Length
Existing embedding models produce lower-quality semantic representations when processing long text. A 5,000-token passage may cover multiple topics, and a single embedding vector struggles to capture all of them simultaneously. (Maximum input length varies widely between embedding models โ check the official docs before choosing one, and don't assume it can swallow a whole section.)
Mitigation strategies:
- Use summary embeddings (the approach described earlier)
- Use ColBERT-style multi-vector representations
- Use multiple embeddings to represent a single section
5. The "Lost in the Middle" Problem
Research shows that LLMs pay weaker attention to middle portions of long contexts (relative to the beginning and end). If critical information happens to be in the middle of a long context, the LLM may overlook it.
Mitigation strategies:
- Place the most relevant sections at the beginning and end of the context
- Explicitly remind the LLM in the prompt to attend to all sections
- Limit total context length โ don't stuff content in without limits
6. Lack of Standardized Evaluation Benchmarks
Traditional RAG has mature evaluation frameworks (Precision@K, Recall@K, MRR, etc.), but evaluating LongRAG is more difficult. Because retrieval unit sizes differ, directly comparing Precision@K is unfair. There are currently no standardized evaluation benchmarks for LongRAG scenarios.
What Later Research Says
LongRAG is a June 2024 technical report. In the year-plus since, "long context vs RAG" has been tested repeatedly and the conclusions do not agree โ so here they are side by side, because reading only the LongRAG abstract leaves you far too optimistic.
The "long context wins" side: Google's Self-Route study (EMNLP 2024 industry track) benchmarked RAG against long-context LLMs across several public datasets and concluded that "when resourced sufficiently, LC consistently outperforms RAG in terms of average performance" โ while explicitly noting that RAG's much lower cost remains a distinct advantage. Hence Self-Route: let the model decide per query whether to take the RAG path or the long-context path, cutting compute cost substantially while staying close to long-context quality. That is exactly the hybrid strategy this article ends on, now with experimental backing.
The opposing side, aimed squarely at LongRAG's core assumption: In Defense of RAG in the Era of Long-Context Language Models (Sept 2024) argues that extremely long context "suffers from a diminished focus on relevant information and leads to potential degradation in answer quality." Their OP-RAG (order-preserve RAG, keeping retrieved chunks in their original document order) finds that answer quality first rises and then declines as the number of retrieved chunks grows, forming an inverted U-shaped curve โ there are sweet spots where OP-RAG beats a long-context LLM using far fewer tokens. That directly contradicts "more context is better."
Another empirical data point: Databricks' Long Context RAG Performance of LLMs ran RAG workflows across 20 open-source and commercial models, sweeping total context from 2K to 128K (and 2M where possible). The finding: retrieving more documents does improve performance, but only a handful of the most recent state-of-the-art LLMs maintain consistent accuracy above 64K, and long context brings its own distinct failure modes.
And a problem with the benchmarks themselves: NoLiMa (ICML 2025), mentioned earlier, shows NIAH-style tests are inflated by literal overlap; strip the lexical cues and long-context degradation is much worse than assumed.
How to read these conflicting results? My reading:
- "Large chunks improve recall" holds up โ LongRAG's retrieval numbers (recall@1 from 52% to 71%) are real.
- "Therefore stuff the large chunks in and let the model pick" is conditionally true: conditional on your model genuinely not degrading at that length, and on the answer not requiring non-literal semantic association to locate.
- "Long context makes RAG obsolete" is the least defensible version. OP-RAG's inverted U and Self-Route's cost conclusion both point the same way: retrieval is still worth doing; the optimal granularity and top-k just have to be found by measurement, not by formula.
A naming trap worth flagging: two different 2024 papers are both called LongRAG. This article covers Jiang et al., arXiv:2406.15319 (long retriever + long reader). The other is Zhao et al., arXiv:2410.18050 (EMNLP 2024 Main), a dual-perspective RAG system for long-context QA with a completely different architecture. When you see LongRAG cited, check which one is meant.
Summary
LongRAG's core insight is simple: When the LLM's comprehension ability is strong enough and the context window is large enough, you don't need to put all the pressure on retrieval.
Traditional RAG's design was reasonable in the era of small context windows: precise splitting, precise retrieval, giving the LLM only the most essential information. But the cost of this strategy is information fragmentation and over-reliance on retrieval precision.
LongRAG redistributes this pressure:
- Retrieval: Shifts from "precisely find the most relevant small fragments" to "roughly find relevant large segments"
- Comprehension: Shifts from "piece together answers from fragmented context" to "understand and answer within complete context"
This isn't meant to replace traditional RAG, but rather provides another effective design choice now that long-context models are widespread. Choose the most suitable strategy based on your document characteristics, query types, cost budget, and latency requirements.
The most pragmatic approach is likely a hybrid strategy: use traditional RAG for simple queries to save tokens, and switch to LongRAG for complex queries to improve quality. A single query classifier can make this happen โ and Self-Route's results are essentially an endorsement of exactly that: let the model decide which path to take, and cost drops sharply while quality stays close to pure long context.
Further Reading
- LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs โ The original LongRAG paper
- Chunking Strategies: How Splitting Methods Determine Whether RAG Can Find Answers โ Detailed comparison of traditional chunking strategies
- Contextual Retrieval: Giving Each Chunk Its Own Context โ Anthropic's alternative approach to solving fragmentation
- Cross-Encoder Reranking: Using Precision Ranking Models to Compensate for Coarse Retrieval โ Remediation strategy when retrieval isn't precise enough
Changelog
- 2026-08-19: Fact-checked against primary sources and refreshed; perishable details handed back to official docs. Added to the "RAG Techniques Compendium" series.
References
- LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs โ Jiang et al. (2024), the original LongRAG paper proposing the complete framework of large-chunk retrieval combined with long-context models
- LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering โ Zhao et al. (EMNLP 2024), a different paper with the same name; easy to cite by mistake
- Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach โ Li et al. (EMNLP 2024), Self-Route: long context wins on average but costs more; route between the two
- In Defense of RAG in the Era of Long-Context Language Models โ Yu et al. (2024), OP-RAG; answer quality follows an inverted U as chunk count grows, arguing against unbounded context
- Long Context RAG Performance of Large Language Models โ Leng et al. (NeurIPS 2024 workshop), long-context RAG measured across 20 models, with failure modes
- NoLiMa: Long-Context Evaluation Beyond Literal Matching โ Modarressi et al. (ICML 2025), long-context performance collapses once literal overlap is removed
- GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models โ Li et al. (2024, EMNLP), graph-structure agent system for processing long documents, surpassing GPT-4-128K with a 4K window
- Retrieval-Augmented Generation for Large Language Models: A Survey โ Gao et al. (2023), comprehensive analysis of RAG evolution and design tradeoffs across generations
- Searching for Best Practices in Retrieval-Augmented Generation โ Wang et al. (2024), systematic experiments on the relationship between chunking strategies and RAG performance
- Anthropic โ Effective Context Engineering for AI Agents โ Anthropic engineering blog on compaction and compression strategies for long-context management
- Multi-Head RAG: Solving Multi-Aspect Problems with LLMs โ Besta et al. (2024), innovative approach using multi-head attention as retrieval keys, complementary to LongRAG
- Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG โ Singh et al. (2025), analysis of Agentic RAG applications in long-document scenarios
Loading...