Video summary
Document Loaders in LangChain | Generative AI using LangChain | Video 10 | CampusX
Main summary
Key takeaways
Main Ideas & Lessons
-
Why Document Loaders Matter in LangChain RAG
- RAG (Retrieval-Augmented Generation) needs an external knowledge base (e.g., PDFs, web pages, CSVs, databases) so the LLM can answer questions it otherwise couldn’t.
- A key part of building RAG apps in LangChain is converting data from many sources into a standard internal format the rest of the pipeline can use.
-
What Document Loaders Are (Core Concept)
- Document loaders in LangChain:
- fetch data from different sources (text files, PDFs, web pages, CSVs, directories, etc.)
- convert it into standardized
Documentobjects
- A
Documentobject includes:page_content: the extracted text/contentmetadata: details like source, creation/modification info, page number, row info, etc.
- The loader output is typically:
- a Python list of
Documentobjects (often one item per page/row, depending on loader type)
- a Python list of
- Document loaders in LangChain:
-
Loader Types Covered (Most Used)
-
Text Loader
- Loads plain
.txtfiles intoDocumentobjects. - Useful for: logs, code snippets, transcripts, etc.
- Loads plain
-
PDF Loader (
PyPDFLoader)- Loads PDFs and processes them page-by-page.
- Output: one
Documentper PDF page- Example: a 23-page PDF → 23 documents
- Limitation: best for text-based PDFs; may struggle with scanned images or complex layouts.
-
Directory Loader
- Loads multiple files (e.g., multiple PDFs) from a folder.
- Uses:
- a directory path
- a glob/pattern to select files (e.g.,
*.pdf) - an inner loader class (e.g.,
PyPDFLoaderfor PDFs)
- Produces a combined list of
Documentobjects across all files/pages.
-
Web Base Loader
- Loads and extracts text content from a web page URL.
- Internally relies on:
- requests to fetch HTML
- BeautifulSoup to parse HTML and extract readable text
- Works best for mostly static pages; struggles with JS-heavy pages
- (Selenium URL Loader is mentioned as an alternative.)
- Usually returns one document per URL (but can accept multiple URLs).
-
CSV Loader
- Loads CSV files and creates one
Documentper row. - Each document includes:
- row content represented as text (column names + values)
- metadata such as the source and row/index details
- Loads CSV files and creates one
-
-
Key Method Choice:
load()vslazy_load()(Important Methodology)-
load()= eager loading- Loads all documents at once into memory.
- Returns a list of
Documentobjects. - Best when document count is small and RAM use is acceptable.
-
lazy_load()= lazy loading- Returns a generator of
Documentobjects. - Fetches/processes one document at a time, not all at once.
- Better for large numbers of files/pages/rows and when you want streaming/low memory.
- Returns a generator of
-
-
How This Fits into a RAG Workflow
- After loading documents into
Documentobjects, the rest of the pipeline typically uses those outputs for:- chunking (text splitters)
- embedding (vector databases)
- retrieval (retrievers)
- generation (LLM answer creation)
- After loading documents into
-
Custom Loaders When Needed
- If a data source doesn’t have an existing LangChain document loader, you can create a custom document loader by:
- defining a class
- inheriting from the base loader class
- implementing
load()and/orlazy_load()logic
- If a data source doesn’t have an existing LangChain document loader, you can create a custom document loader by:
Instructional / Step-by-Step Portions (As Presented)
A) General Document Loader Pattern in LangChain
- Import the loader from
langchain_community.document_loaders - Create a loader instance with required parameters (examples below)
- Call:
loader.load()to get a list ofDocumentobjects (eager)- or
loader.lazy_load()to get a generator ofDocumentobjects (lazy)
- Use each returned
Document:doc.page_contentdoc.metadata
B) TextLoader Usage (Conceptual Steps)
- Import
TextLoader - Initialize it with:
- path to the
.txtfile - optionally encoding (e.g., UTF-8)
- path to the
- Call
load()to get a list of documents (often 1 document for a single file) - Extract:
- first document →
docs[0].page_content - metadata →
docs[0].metadata
- first document →
C) PyPDFLoader Usage (Conceptual Steps)
- Import
PyPDFLoader - Initialize it with:
- PDF file path
- Call
load():- returns a list where each page becomes one document
- Extract per-page:
docs[i].page_contentdocs[i].metadata(includes page number and other PDF metadata)
D) DirectoryLoader Usage (Conceptual Steps)
- Import
DirectoryLoaderplus a “document loader class” (e.g.,PyPDFLoader) - Initialize it with:
path: folder containing documentsglob: file pattern (e.g.,*.pdf)loader_cls: which loader to use for each file (e.g.,PyPDFLoader)
- Call
load()to get combined documents from all matching files
E) WebBaseLoader Usage (Conceptual Steps)
- Import
WebBaseLoader - Initialize it with:
- a URL (or list of URLs)
- Call
load()to get documents:- generally 1 document per URL
- Use returned document text to ask questions via an LLM chain
F) CSVLoader Usage (Conceptual Steps)
- Import
CSVLoader - Initialize with:
- CSV file path
- Call
load():- returns documents where each CSV row becomes one document
- Extract:
docs[i].page_content(row formatted as text with column/value pairs)docs[i].metadata(source and row-related metadata)
G) Choosing Between load() and lazy_load() (Rule of Thumb)
- Use
load()when:- document/file count is small
- you need everything in memory at once
- Use
lazy_load()when:- document/file count is large
- you want streaming or to avoid high RAM usage
Speakers / Sources Featured
- Speaker: Nitish (host/presenter of the channel)
- Primary referenced source/library:
- LangChain
langchain_community.document_loaders
- Referenced external components/libraries:
- ChatGPT (as an example of a chatbot limitation prompting RAG)
- PyPDF library (used internally by
PyPDFLoader) - Requests and BeautifulSoup (used internally by
WebBaseLoader) - Selenium URL Loader (mentioned as an alternative for JS-heavy pages)
- Python (lists, generators, and the
loadvslazy_loadconcept)