Video summary
Project 36 : Resume Categorization Using Machine Learning
Main summary
Key takeaways
Resume Categorization Using Machine Learning (Python)
This video builds an end-to-end resume categorization application using Python + NLP + machine learning, along with a Streamlit web UI. The workflow lets you upload resumes (PDF), classify them into job categories, save results, and download predictions as CSV.
1) Application (Streamlit) demo: how categorization works
- Users upload one or multiple resumes via the UI (expects PDF resumes).
- The app:
- Reads each PDF
- Extracts text
- Cleans the text (removes URL/email/special characters + stopwords)
- It then predicts the resume’s category (examples shown: Data Science, Java Developer, Business Analyst/Analyst, DevOps engineer, etc.).
Outputs and workflow
- Creates output folders per category inside a chosen output directory (e.g.,
categorized resume/Java developer,.../data science). - Optionally downloads predictions as a CSV containing:
- file name
- category
- Supports updating automatically for newly uploaded files and allows deleting/resubmitting resumes within the app workflow.
- Allows bulk selection with a size limit noted in the UI (up to 200MB).
2) Dataset and model training approach
- Training data comes from Kaggle resume datasets with a CSV structure:
- Columns:
categoryandresume
- Columns:
- Datasets referenced:
resume_dataset.csvupdated_resume_dataset.csv(used in the tutorial)
Framing the task
- The problem is set up as a supervised multi-class classification task.
- There are many categories (examples include Data Science, HR Advocate, Arts, Web Designing, Mechanical Engineer, Sales, Health and Fitness, Civil Engineer, Java Developer, Business Analyst, etc.).
- Dataset size mentioned: ~962 rows, 2 columns (category + resume).
3) Core NLP preprocessing (text cleaning)
Using regex plus NLP tooling (e.g., NLTK stopwords):
- Libraries include:
renltkstopwords, etc.
- The cleaning process removes:
- URLs
- emails
- special characters (via regex)
- stopwords (NLTK English stopwords)
A clean() function is defined and applied across the dataset, for example:
df['resume'].apply(lambda x: clean(x))
The cleaned text is then used for vectorization and classification.
4) Feature extraction and label encoding
- Label encoding converts categorical labels (e.g., “data science”, “Java developer”) into numeric IDs using
LabelEncoder. - TF-IDF vectorization converts the cleaned resume text into numeric feature vectors using
TfidfVectorizer(TF-IDF).
5) Train/test split
- Data is split using
train_test_split:- 80% training
- 20% testing
random_state=42
6) Model training + comparison (multiple classifiers)
The tutorial trains several models and compares test accuracy:
- K-Nearest Neighbors (KNN)
- Logistic Regression
- Random Forest
- Support Vector Classifier (SVC)
- Multinomial Naive Bayes
- A One-vs-Rest approach wrapping an estimator for multi-class handling (one-vs-rest logic is evaluated and debugged after an “estimator” usage issue)
Approximate observed accuracies
- KNN: around 0.98
- Logistic Regression: around 0.99
- Random Forest: around 0.98–0.99
- SVC: around 0.99
- Multinomial NB: around 0.97
- One-vs-rest: also evaluated and compared
7) Evaluation and prediction mapping back to category names
- The selected model predicts a sample resume (based on extracted skills/text).
- Predictions initially return numeric IDs.
- A
category_map(ID → label) maps numeric outputs back to category names (otherwise defaults to “unknown”).
8) Persisting artifacts (model + vectorizer)
Both the trained model and the TF-IDF vectorizer are saved using pickle:
model.pkltf-idf.pkl
The Streamlit app later loads these artifacts to run predictions on uploaded resumes.
9) Streamlit app implementation details
- Run command:
streamlit run app.py
Key imports mentioned:
pickle(to load model + TF-IDF)PyPDF2(to extract text from uploaded PDFs)re(regex cleaning patterns)streamlit as st
UI elements
- File upload:
st.file_uploader(..., type="pdf", accept_multiple_files=True)
- Output directory:
- selection via
st.text_input
- selection via
- Button to trigger categorization
Per-file processing (for each uploaded PDF)
- Extract text (first page noted as preferred/simplified)
- Clean resume text
- TF-IDF transform
- Predict category
- Write outputs:
- files into corresponding category folders
- results into a DataFrame for CSV download
10) Handling non-PDF sources (docs → PDF conversion)
The tutorial briefly addresses Kaggle dataset format issues:
- If resumes are in DOCS format and need conversion to PDF, it provides a utility that:
- iterates a directory
- checks file extensions
- converts using a “convert” method from a docs-to-PDF module
Main speakers / sources
- Speaker: The video narrator/author (unnamed in subtitles), presenting a “multiverse of 100+ data science project series”.
- Sources referenced:
- Kaggle resume datasets (
resume_dataset.csv,updated_resume_dataset.csv) - Libraries such as NLTK, scikit-learn, PyPDF2, and Streamlit.
- Kaggle resume datasets (