Every crawler records the standard SEO fields: title, meta description, canonical, status code. But the questions you actually need to answer are often specific to the site. Which product pages are out of stock but still indexable? Which articles have no author? Which templates dropped the analytics tag?
Custom extraction answers those by pulling any value out of every page’s HTML while you crawl. This guide covers how it works and eight extractions worth setting up.
How does custom extraction work?
You define an extractor: a name, and either a CSS selector or a regular expression.
| Type | Reads | Returns | Best for |
|---|---|---|---|
| CSS selector | The parsed page | The element’s text, its inner HTML, or an attribute | Visible elements and meta tags |
| Regex | The HTML source | The match, or just the capture group if there is one | Values inside scripts or comments |
By default you get the first match; turn on All matches to collect every one. The values show up as columns in the URLs table and in exports, so you can filter, sort and spot gaps.
In Crawlens, extractors live under Extraction & rules. Use Test on a URL to check an extractor against a live page before crawling; changes apply to the next crawl.
8 extractions worth setting up
| What | Type | Expression | Output |
|---|---|---|---|
| 1. Analytics / tag manager ID | Regex | GTM-[A-Z0-9]+ |
Match |
| 2. GA4 measurement ID | Regex | G-[A-Z0-9]{6,} |
Match |
| 3. Author | CSS | meta[name="author"] |
Attribute content |
| 4. Publish date | CSS | meta[property="article:published_time"] |
Attribute content |
| 5. Price | CSS | .price (your theme’s selector) |
Text |
| 6. Stock status | CSS | [itemprop="availability"] or your theme’s badge |
Attribute href or text |
| 7. Breadcrumb trail | CSS | nav.breadcrumb a with All matches |
Text |
8. Schema @type values |
Regex | "@type"\s*:\s*"([^"]+)" with All matches |
Capture group |
Adjust selectors to your theme: open a page in the browser, inspect the element and copy a stable class or attribute. A few notes:
- Tracking IDs live inside scripts, so use regex. Comparing the extracted ID across pages quickly shows templates that are missing the tag or use the wrong container.
- Dates and authors support content audits: find stale articles, posts with no author for E-E-A-T, or dates that never update.
- Prices and stock help e-commerce audits: indexable pages for products that are out of stock, or variants missing a price.
- Schema types show which templates output which structured data, alongside the built-in structured data checks.
Turn extractions into checks with custom rules
Extraction tells you the values; custom rules flag the pages where they’re wrong. A Crawlens custom rule has:
- a title, severity (critical, warning or notice) and optional why it matters / how to fix text
- a scope: indexable pages, all HTML pages, or every URL
- conditions on page fields (URL, title, meta description, H1, canonical, content type, indexability reason, status code, word count, response time, click depth, inlinks) or on any extracted value
- operators like contains, is, matches regex, is empty, is greater than, and whether all or any conditions must match
Some examples:
| Rule | Scope | Conditions |
|---|---|---|
| Pages missing the GTM container | Indexable | Extracted GTM ID is empty |
| Titles without the brand | Indexable | Title does not contain YourBrand |
| Out-of-stock products still indexable | Indexable | Extracted Stock contains OutOfStock |
| Articles without an author | Indexable | URL contains /blog/ and extracted Author is empty |
| Slow, deep pages | HTML | Response time is greater than 1000 and click depth is greater than 3 |
Before saving, Preview shows how many URLs the rule would flag on the latest crawl. Custom rules then run in every audit like the built-in checks, appear in the issues list and reports, and can be re-applied to an existing crawl with Re-run audit.
Tips
- Name extractors clearly. They become column names and rule fields.
- Prefer stable selectors (IDs,
itemprop, data attributes) over layout classes that change with a redesign. - Use capture groups in regex to return just the value, not the surrounding code.
- Start with one template. Test on a URL, crawl a section, then roll out site-wide.
- Ask your AI assistant. Extracted values are available to AI tools over MCP, so you can ask questions like “which product pages have no price?” See connecting AI tools.
Checklist
- Extractors for the site-specific fields you audit (tracking, authors, dates, prices, schema)
- Each extractor tested on a live URL before crawling
- Custom rules for the gaps that matter, previewed before saving
- Clear titles and fix instructions on custom rules, so issues explain themselves
- Re-run the audit after changing rules
Frequently asked questions
What is custom extraction in SEO crawling?
It's a crawler feature that collects specific values from every page's HTML while it crawls, using CSS selectors or regular expressions. The values appear as extra columns next to the usual SEO data.
Should I use a CSS selector or a regex?
Use a CSS selector for visible elements and attributes, like a price in span.price or a meta tag's content. Use a regex for things inside scripts or the raw source, like an analytics ID in a gtag call.
How do I extract only part of a match with regex?
Use a capture group. In Crawlens, if the regex has a group in parentheses, only that part is returned; otherwise the whole match is.
Can I extract every match on a page, not just the first?
Yes. Turn on the option for all matches to collect every value, for example every author on a multi-author page. Crawlens keeps up to 100 values per extractor per page.
What are custom rules?
Custom rules are your own audit checks. They flag pages where built-in fields or extracted values meet conditions you set, and they run in every audit alongside the built-in checks.