Chunking is how you prepare content for Retrieval-Augmented generation. A long page or PDF is cut into smaller passages, each of which is turned into an embeddings vector and stored, so a search can return the specific section that answers a question.
Chunk size is a trade-off. Very small chunks lose context, so a retrieved sentence may make no sense alone. Very large chunks blur several topics together and fill up the context window. Splitting along natural boundaries such as headings and paragraphs, with a small overlap between neighbors, usually beats cutting at a fixed character count.
Keep metadata with each chunk, such as the page title, section heading and URL, so results can be filtered and cited. Test retrieval on real questions, because chunking choices show up directly in answer quality.
Related: Vector Database, Semantic Search, Knowledge Base.