Custom extraction for SEO: 8 things to pull from every page

Use CSS selectors and regex to extract prices, authors, dates, analytics tags and schema fields from every page, then flag pages with custom rules.

Every crawler records the standard SEO fields: title, meta description, canonical, status code. But the questions you actually need to answer are often specific to the site. Which product pages are out of stock but still indexable? Which articles have no author? Which templates dropped the analytics tag?

Custom extraction answers those by pulling any value out of every page’s HTML while you crawl. This guide covers how it works and eight extractions worth setting up.

How does custom extraction work?

You define an extractor: a name, and either a CSS selector or a regular expression.

Type Reads Returns Best for
CSS selector The parsed page The element’s text, its inner HTML, or an attribute Visible elements and meta tags
Regex The HTML source The match, or just the capture group if there is one Values inside scripts or comments

By default you get the first match; turn on All matches to collect every one. The values show up as columns in the URLs table and in exports, so you can filter, sort and spot gaps.

In Crawlens, extractors live under Extraction & rules. Use Test on a URL to check an extractor against a live page before crawling; changes apply to the next crawl.

8 extractions worth setting up

What Type Expression Output
1. Analytics / tag manager ID Regex GTM-[A-Z0-9]+ Match
2. GA4 measurement ID Regex G-[A-Z0-9]{6,} Match
3. Author CSS meta[name="author"] Attribute content
4. Publish date CSS meta[property="article:published_time"] Attribute content
5. Price CSS .price (your theme’s selector) Text
6. Stock status CSS [itemprop="availability"] or your theme’s badge Attribute href or text
7. Breadcrumb trail CSS nav.breadcrumb a with All matches Text
8. Schema @type values Regex "@type"\s*:\s*"([^"]+)" with All matches Capture group

Adjust selectors to your theme: open a page in the browser, inspect the element and copy a stable class or attribute. A few notes:

Turn extractions into checks with custom rules

Extraction tells you the values; custom rules flag the pages where they’re wrong. A Crawlens custom rule has:

Some examples:

Rule Scope Conditions
Pages missing the GTM container Indexable Extracted GTM ID is empty
Titles without the brand Indexable Title does not contain YourBrand
Out-of-stock products still indexable Indexable Extracted Stock contains OutOfStock
Articles without an author Indexable URL contains /blog/ and extracted Author is empty
Slow, deep pages HTML Response time is greater than 1000 and click depth is greater than 3

Before saving, Preview shows how many URLs the rule would flag on the latest crawl. Custom rules then run in every audit like the built-in checks, appear in the issues list and reports, and can be re-applied to an existing crawl with Re-run audit.

Tips

Checklist

Frequently asked questions

What is custom extraction in SEO crawling?

It's a crawler feature that collects specific values from every page's HTML while it crawls, using CSS selectors or regular expressions. The values appear as extra columns next to the usual SEO data.

Should I use a CSS selector or a regex?

Use a CSS selector for visible elements and attributes, like a price in span.price or a meta tag's content. Use a regex for things inside scripts or the raw source, like an analytics ID in a gtag call.

How do I extract only part of a match with regex?

Use a capture group. In Crawlens, if the regex has a group in parentheses, only that part is returned; otherwise the whole match is.

Can I extract every match on a page, not just the first?

Yes. Turn on the option for all matches to collect every value, for example every author on a multi-author page. Crawlens keeps up to 100 values per extractor per page.

What are custom rules?

Custom rules are your own audit checks. They flag pages where built-in fields or extracted values meet conditions you set, and they run in every audit alongside the built-in checks.

· Founder, Crawlens

Dien builds Crawlens, a desktop crawler for technical SEO audits, and writes about the checks it runs: crawling, indexing, JavaScript rendering and Search Console data.

Audit your own site with Crawlens

Crawl, run 70 checks, and see them next to Search Console and Core Web Vitals data.

Download free for Windows