CommunityCodierung & Entwicklunggithub.com

Vectle/skills

This skill covers debugging steps when robots.txt blocks a pricing page crawl. Use it when a crawler starts skipping URLs or when you are auditing a new source. It is not for ignoring robots.txt; the fix is finding which rule matches and sourcing the data from an allowed channel like the sitemap, an official API, or the site owner.

Was ist skills?

skills is a Claude Code agent skill that this skill covers debugging steps when robots.txt blocks a pricing page crawl. Use it when a crawler starts skipping URLs or when you are auditing a new source. It is not for ignoring robots.txt; the fix is finding which rule matches and sourcing the data from an allowed channel like the sitemap, an official API, or the site owner.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/Vectle/skills/tree/HEAD/skills/robots-txt-blocking-pricing-page-crawl-758e2a1729cdf8d5

Installed? Explore more Codierung & Entwicklung skills: steipete/bluebubbles, steipete/eightctl, steipete/blucli · View all 6 →

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

robots.txt is blocking the pricing page crawl

TL;DR

When robots.txt disallows the pricing path, that is the site owner declining crawler access, and the correct response is to respect it rather than route around it. The debug part is quick: fetch robots.txt, find which rule matches your path, and check whether your crawler is accidentally hitting disallowed URLs it does not need. For the pricing data itself, use the site's official API, a licensed data vendor, the sitemap's allowed pages, or ask the owner for a sanctioned feed.

The error

(no HTTP error; the crawler refuses to fetch)
robots.txt disallows: /pricing/*
scraper skipped 42 URLs as disallowed

When this helps

  • a crawler suddenly skips URLs it used to fetch
  • auditing a new source before a briefing agent depends on it
  • pricing pages vanish from a crawl while blog pages still work
  • deciding whether a competitor pricing source is crawlable at all

When it doesn't

  • you want to ignore robots.txt and fetch anyway; this skill will not help with that
  • the path is allowed but returns 403 anyway, that is a bot-protection problem, not robots.txt
  • you need data behind a login wall

Works with

Any crawler: python urllib.robotparser, scrapy, or curl for manual checks. robots.txt semantics are stable across versions.

Steps

1. Read the file your crawler is obeying

curl -s "https://YOUR-site/robots.txt"

Expected: The raw rules. Look for the Disallow lines under the User-agent block that matches your crawler, or under the wildcard block.

2. Find which rule matches your pricing URLs

curl -s "https://YOUR-site/robots.txt" | grep -i -A2 -B2 "pricing"

Expected: The specific Disallow line covering your path. If a narrower rule allows part of the section and a broader one blocks it, the most specific match wins per the standard.

3. Check your crawler is not wandering into disallowed paths by accident

from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.set_url("https://YOUR-site/robots.txt")
rp.read()
for u in ["https://YOUR-site/pricing", "https://YOUR-site/pricing/enterprise", "https://YOUR-site/blog/pricing-guide"]:
    print(u, "allowed:" , rp.can_fetch("IntelBriefingBot", u))

Expected: A per-URL allowed verdict. Blog or docs pages about pricing are often allowed even when the pricing app itself is blocked; fetch only what is allowed.

4. Source the pricing data from an allowed channel

curl -s "https://YOUR-site/sitemap.xml" | head -20

Expected: The sitemap's crawlable URLs. Pricing pages that matter for intel often have public summaries, press releases, or a vendor API; if the data truly only exists behind a disallowed path, ask the site owner for a feed.

Other ways people phrase this

crawler blocked by robots.txt disallow rule

General form of the same situation. The file is the site owner's stated preference; treat a Disallow as a no.

pricing page not in crawl, robots.txt check

Pricing sections are the most commonly disallowed paths on SaaS sites. Check for allowed alternates like public pricing summaries or press releases first.

scraper respects robots.txt but misses data

That is the tool working correctly. The gap is filled with sanctioned sources, not with ignoring the file.

Why it happens

robots.txt is the site owner's machine-readable statement of where crawlers are welcome. Pricing paths get disallowed because they are high-value, frequently changing, and expensive to serve to bots. A crawler that obeys the file is doing its job; the data gap is a sourcing problem, not a crawler bug.

Edge cases

  • robots.txt is advisory, not access control; obeying it is a norms and ToS matter, and ignoring it is how crawlers get IP-blocked.
  • Some sites block all bots but allow specific partner user agents; do not spoof another bot's user agent to qualify.
  • A missing robots.txt means everything is technically fetchable, but terms of service and rate etiquette still apply.
  • Sitemap URLs can include paths robots.txt disallows; the Disallow still governs crawling even when the sitemap lists the page.

Provenance

Resolved from the public thread: https://vectle.com/posts/pst_SazLC5udAbtYND83hj1H1Q


Source: Vectle. Version: skv_psAetA1Q4R83R1qW5T3CVg. Original contributor attribution is preserved in the frontmatter. Content remains subject to Vectle’s terms and its contributors’ rights; this mirror does not grant a new content license.

Individual skills in this repo

This repo contains 2 individual skills — each has its own dedicated page.

Verwandte Skills