Video summary
The SEO Audit Process I'd Use in 2026
Main summary
Key takeaways
Main ideas / lessons (SEO audit “battle plan”)
- Agencies often charge a lot for SEO audits, but you can follow a structured, step-by-step process to create your own data-driven plan.
- The audit emphasizes starting with crawl/indexing issues, then moving through internal linking, performance, and technical/content hygiene.
- It repeatedly blends:
- Quantitative data from crawling and analytics tools, and
- Qualitative/intent-based judgment when diagnosing content and opportunities.
- Later steps expand from classic SEO fixes into topic dominance and AI/LLM retrieval optimization (getting mentioned/covered across many platforms and improving citations AI uses).
Methodology: step-by-step SEO audit process (detailed)
1) Find crawling and indexing opportunities
Create a Google Sheet
- Add a first tab named “crawl”.
- (A template is mentioned as existing in “Gotcha SEO Academy,” but the approach demonstrates building from scratch.)
Run a crawl with Screaming Frog
- Strongly recommended for the audit (speaker has used it for ~a decade).
Configure Screaming Frog (quickest, useful crawl)
- Crawl configuration
- Uncheck/disable options that add heavy load (e.g., resource links).
- Keep remaining settings aligned with the speaker’s “simple” configuration.
- Extraction
- Use standard settings.
- Structured data auditing is optional and site-dependent (more useful for local/ecommerce).
- Duplicates / analysis enhancements
- Enable near duplicates
- Enable spell & grammar check
- Enable embedding-related options:
- semantic similarity
- low relevance
Set up API access (minimum set recommended = 3)
- Google Analytics 4
- Google Search Console
- PageSpeed Insights
- (Optionally integrate AI platforms later; starting without them is acceptable.)
Run the crawl
- Export/import results into the Google Sheet.
- Keep the crawl open for later visualizations and steps.
2) Find internal linking opportunities
- From the crawl export in the Google Sheet:
- Freeze the header row
- Add filters
- Focus on indexable content
- Prioritize columns for internal linking:
- Crawl depth
- Flag pages deeper than 3 clicks
- Why: deeper pages are harder to crawl/index → can’t rank if not indexed.
- Definition: crawl depth = number of clicks to reach a URL.
- Unique in links (internal links count)
- Flag pages with fewer than ~5 internal links (example threshold used)
- Crawl depth
Interpret patterns
- Many deeply buried pages + low internal links suggests:
- missing internal-link “injection” on relevant pages, and/or
- weak topic support (content not strongly connected to the site’s theme)
3) Find page loading speed opportunities
- Use the Performance Score column from the crawl results.
- Mark pages with Performance Score < 80.
Why it matters
Even if it’s not a direct ranking factor:
- Bad UX increases bounce / “pogo-sticking” risk.
- AI/LLM crawlers have very limited time (speaker references < ~3 seconds).
- Better speed improves crawling, indexing, and UX/conversions.
Practical diagnosis
- Use Lighthouse (Chrome tooling/extension).
- Test ~5–10 pages.
- If desktop looks fine, test mobile (mobile is often the real issue).
Optimization expectation
- Fixing one page can improve performance sitewide.
4) Find 404s and broken links
- Review status codes in the crawl.
Not all 404s are bad
- A 404 is acceptable if the page is intentionally removed and should disappear from the index.
“Bad” 404s
- 404 pages that still have positive value signals.
Action criteria used
- Filter to 404 pages
- Check whether they have KPIs such as:
- traffic
- Search Console impressions
- clicks
- engagement/event metrics
If a 404 has positive KPIs
- Decide whether to redirect it or rework/restore content to recapture demand.
Impact concept
- Leaving valuable pages at 404 can cause large impression losses.
5) Find thin content
- Use word count as the initial flag.
- Flag pages with < 500 words.
Important clarification
- Flagging ≠ automatically adding junk content.
- Put pages into investigate buckets for:
- improvement
- deletion
- restructuring
6) Find duplicate content
Definition
- Same exact content appears on more than one page.
Detection approach
- Use Screaming Frog’s near-duplicate/duplicate match data.
- If configured properly, it may show few/no near duplicates (example: none found on their site).
- Verification method:
- Run a Siteliner crawl for duplicate detection.
- If match percentage < 30%, it’s likely not a big concern.
- Sort by match percentage and check for:
- exact duplicates
- highly similar (“near-duplicate”) pages
Potential action
- If similarity is high but not exact, make pages more unique.
7) Find keyword cannibalization
Definition
- Two pages compete for the same keywords with the same intent.
Clarification
- Pages can target similar keywords (singular/plural, same seed topic) without cannibalizing if intent differs.
Intent-based example
- A commercial “sales” page and an investigative “buyer guide” page can work together rather than compete—depending on keyword variants and intent.
Simple identification approach
- In the crawl:
- filter/search by title text contains a seed term
- check whether multiple pages target the same modifiers/intent and thus truly compete
Rule of thumb
- You can share seed keywords if intent modifiers make pages genuinely different.
- Still recommended: keep one core topic per page.
8) Find irrelevant content
Goal
- Keep topic focus and “subject matter expertise” tight.
- Avoid drifting into topics that don’t belong to the site theme (harder to build strong topic clusters).
Detection method
- Search/filter by title tags that do not contain the core theme term (example used: “SEO”).
Evaluate candidates
- Even if such pages rank and get traffic, consider:
- redirecting
- deleting
- reworking to become relevant
If irrelevant pages have positive KPIs
- Prioritize preserving value while improving relevance.
9) Find weak clusters
Cluster definition
- A group of pages around one topic, with support assets linking back to a core page.
Process
- Identify both strong and weak clusters.
- If a topic underperforms, examine how strong the support assets are.
Example logic
- One cluster might have ~20 supporting pages (moderately strong).
- Another pillar + extensive supporting reviews can reach ~75–100 assets (very strong).
Key takeaway
- Often improves SEO to strengthen one cluster vs applying many scattered fixes.
10) Find content quality opportunities (qualitative diagnosis using data)
- Use organic performance metrics to select pages to inspect.
- Filter by low Google Search Console impressions.
- Highlight the top 10–15 weakest pages by impressions.
Define “low quality / not worth keeping”
General criteria:
- no traffic
- no impressions
- no backlinks
- If older pages still have none of these signals: consider deletion or a strategy reset.
If the page has modest signals but is low-quality
- Consider “beefing up” with more relevant information (e.g., bios, interviews, transcripts), if appropriate.
11) Find on-page SEO opportunities
Bare-minimum checklist for organic pages
Ensure the keyword appears in:
- title tag
- meta description
- URL slug
- H1
- first sentence
NLP/related-topic coverage
- Use NLP concepts to ensure coverage of relevant subtopics.
12) Find content optimization opportunities
- For pages needing improvement (often those ranking around ~6 in the example):
Workflow
- Run the keyword through Rankability (content optimizer).
- Use “import URL content” + run optimizer.
Use the output to:
- identify unused topics
- add missing subtopics covered by competitors/top results
Timing concept
- Revisit/refresh assets as they age (competitors and topical coverage change).
Expected impact
- Content rework can improve ranking by ~2–4 spots (as claimed).
13) Find low-hanging fruits (keyword footprint sweet spot)
- Use keyword position data to find pages near top results.
- Focus on pages ranking approximately positions 2–15.
Process
- Sort by average position and filter by relevant cluster/keyword set.
- Verify in incognito/private mode to avoid personalization bias.
Prioritize fixes (usually)
- Rankability optimization (relevance)
- internal linking
- loading speed
- If still not improving: move to backlinks as the next tier
14) Find clustering opportunities (position 50+ established demand signals)
- Define a bucket: keywords/pages ranking 50 and beyond.
- Ignore newly published pages; the approach targets pages with demand signals.
- In GSC:
- sort/filter so you see cases with impressions but insufficient relevance/optimization.
Example insight
- If there’s demand for “technical SEO training” but you only have academy content (no dedicated page):
- build a dedicated landing page.
How to build
- Send keyword to Rankability
- Extract NLP keywords and create an outline
- Decide AI vs manual based on competition:
- lower competition → more AI
- higher competition (e.g., SEO industry) → more manual tailoring
Core principle
- After indexing, relevance is emphasized as the most important factor.
15) Find topic domination opportunities (cover a topic across many “surfaces” + AI retrieval)
Definition
- Cover a topic on your site and also across other platforms that Google and AI systems index/retrieve from.
Goal
- Influence both:
- traditional first-page rankings
- AI/ChatGPT-style retrieval outputs
Process
- Use Search Console to find a keyword you already do well for.
- Check which assets appear (book page, blog, etc.).
- Identify missing coverage and improve the best candidate asset.
Next steps
- Check YouTube:
- if missing, create video content (YouTube matters because Google owns it)
- Examine which external sources rank (e.g., Reddit, industry publications)
- Track recurring subreddits and set alerts for mentions to respond/promote
- Pitch inclusion on lists when feasible
16) Find critical citation opportunities (reverse-engineer AI citations)
Concept
- AI answers often use RAG (retrieval) from sources it can cite.
Process
- For the target topic/keyword:
- run the query on multiple AI platforms:
- ChatGPT
- Perplexity
- Claude
- Grock
- check each platform’s citations
- identify gaps (where your brand/site isn’t cited)
- outreach to sources missing your citation
- run the query on multiple AI platforms:
Why multi-platform
- Each system retrieves slightly differently, producing unique citation gaps.
17) Find competitor’s top performing pages (use proof, then replicate framework)
Rationale
- AI-generated ideas are less reliable than pages with measurable link attraction.
Process
- Use Ahrefs or Semrush
- Go to “best by links” for a competitor
- Identify page frameworks correlated with high link acquisition
Noted pattern
- Often statistics-driven pages attract links.
Strategy
- Replicate the proven framework, but keep it unique.
- Consider transferring successful templates from other industries to differentiate.
18) Find misinformation about your brand
Goal
- Detect inaccurate or inconsistent brand info that can harm trust, retrieval, or recommendations.
Quick method
- Use a GPT/plugin approach (built by the speaker) to probe what the model knows from training data only (no web search).
- Input the brand name.
Cautions
- Not perfect; models can hallucinate if the brand isn’t well represented.
Action principle
- If results show weak/incorrect knowledge, improve brand clarity and accuracy.
Speakers / sources featured
- Speaker: Nathan Gotch
- Tools/platforms mentioned:
- Screaming Frog
- Google Sheets
- Gotcha SEO Academy (templates mentioned)
- Google Analytics 4 (GA4)
- Google Search Console (GSC)
- PageSpeed Insights
- Lighthouse (Chrome/extension)
- Siteliner
- Rankability (content optimizer)
- Rankability / “detailed Chrome extension” (intent/optimization workflow)
- Ahrefs / Semrush (competitor “best by links” research)
- ChatGPT
- Perplexity
- Claude
- Grock
- YouTube
- Search Engine Journal (example: pitching lists)
- A “free GPT”/probe tool built by the speaker (brand misinformation probing; name not clearly specified)