Quick overview
This workflow crawls a website from its sitemap, extracts on-page content and links, uses Anthropic Claude to identify keywords and entities, stores everything as an SEO knowledge graph in Postgres, then posts crawl health stats to a Slack channel.
How it works
- Triggers manually or on a weekly schedule to start a crawl using the configured seed URL, sitemap URL, page limit, and Slack channel.
- Fetches sitemap.xml, parses and deduplicates page URLs (optionally same-domain only), and batches them for processing up to the max page count.
- For each page URL, downloads the HTML and extracts the title, meta description, H1s, body text, and outbound links.
- Sends the page content to Anthropic Claude to return structured JSON for primary/secondary keywords, named entities, and topic categories.
- Writes the page node plus keyword, entity, and link relationships into Postgres tables (seo_pages, seo_keywords, seo_entities, seo_page_keywords, seo_page_entities, seo_page_links).
- After all pages are processed, queries Postgres for digital-twin stats (pages, unique keywords, entities, internal links, orphan pages) and posts the summary to Slack.
Setup
- Create a Postgres database and add an n8n Postgres credential, then create the required tables and constraints used by the queries (seo_pages, seo_keywords, seo_entities, seo_page_keywords, seo_page_entities, seo_page_links, including the ON CONFLICT keys).
- Add an HTTP Header Auth credential for the Anthropic API (x-api-key) and ensure the selected model name in the config matches your Anthropic access.
- Add a Slack credential, set or create the target channel, and update seedUrl, sitemapUrl, maxPages, sameDomainOnly, anthropicModel, and slackChannel in the crawl configuration.