RAG Chunking Strategy & Vector Index Calculator
Estimate how many vectors a corpus produces at a given chunk size and overlap, what those vectors cost in index RAM, and what one pass of embedding costs at the published rate for the model you pick. Every figure below is arithmetic on the inputs, with the assumptions stated at the bottom of the page.
Corpus and model parameters
RAG index capacity and memory allocation Calculated from the inputs
What these inputs imply
- Corpus chunked into 114,890 vector embeddings at 512 tokens per chunk.
- Switching to 8-bit scalar quantisation would cut stored vector bytes roughly fourfold.
- Estimated vector index RAM: 0.99 GB.
How these numbers are calculated
- Chunk count is the corpus token total divided by the effective stride, which is chunk size multiplied by one minus the overlap ratio. Overlap re-embeds text, so it raises the vector count rather than leaving it unchanged.
- Vector bytes are dimensions multiplied by 4 for unquantised float32, or by 1 for 8-bit scalar quantisation.
- Index RAM applies the 1.5x multiplier Qdrant documents for in-memory collections, where the extra 50% covers index structures, point versions and temporary segments during optimisation. A flat index carries no graph, so no multiplier is applied.
- Embedding cost follows the published OpenAI list rate for the model selected above: $0.02 per million tokens for text-embedding-3-small and $0.13 for text-embedding-3-large. BGE-M3 and Nomic Embed v1.5 are open-weight and shown at zero because the calculator assumes you run them on your own hardware; a hosted endpoint for either would carry its own rate. This is a one-time indexing cost rather than a monthly charge, and re-indexing after a model change repeats it in full.
These are sizing estimates, not measurements. Payload storage, replication factor and filtering indexes all add to the real footprint, and no retrieval-quality figure is estimated here because recall depends on your corpus and questions, not on these inputs.
Primary sources: Qdrant capacity planning, Qdrant quantization, OpenAI embeddings guide, OpenAI API pricing.
Choosing the inputs
The two inputs that move every other number are chunk size and embedding width. Both are explained in the guides below.
- RAG chunking strategy: picking chunk size and overlap — how to choose the two values at the top of this form.
- RAG pipeline architecture: components and build order — where sizing sits relative to everything else.
- Qdrant vs Milvus vs Pinecone — which stores expose quantisation and index parameters at all.
- RAG retrieval debugging — what to check when the index is sized correctly and results are still wrong.