Lead Signal Scout
Turn a product URL into a shortlist where every entry sits on a page the agent opened, read and dated, and belongs to someone who could actually buy. Unverified, stale and unqualified prospects never reach the shortlist.
Six gates are mandatory. Skipping one invalidates the run; say so rather than shipping the output.
- Intake (step 1) - do not research until the brief can reject a weak match.
- Verification (step 4) - do not write up a prospect whose page has not been fetched.
- Screening (step 4) - do not abandon a query that surfaced results without opening one of them and writing down why.
- Recency (step 5) - do not shortlist a signal that is undated or past the cutoff.
- Qualification (step 6) - do not shortlist a prospect with no evidenced problem of its own, no ability to buy, no usable contact route, or no identified company in a company-led mode.
- Audit (step 8) - do not present results until
validate_run.pyreturns PASS.
Write every prose field against references/writing-rules.md. A report that reads as generated gets ignored, whatever the research behind it.
Workflow
1. Intake
Read references/intake.md. Build the product brief, then ask once, in a single batched message, for the four things a landing page cannot tell you: geography, price band and buying motion, exclusions, and any existing-customer or past-win data. Proceed with labelled assumptions if the user does not answer.
Record icp.price_band_source as user or assumed. This matters: a band you invented may not remove anyone. With assumed, do not score commercial_fit below 3 on price grounds. A number nobody confirmed has no business deciding which real companies get dropped.
Set max_age_days here. The default is 120. Use 60 when the signals will be hiring posts or vendor searches, where a four-month-old request is almost always filled.
2. Set mode and depth
Mode decides what is optimised for. Depth decides how many. They are independent.
| Mode | Optimises for | Weights (pain / fit / commercial / timing / reach / evidence) |
|---|---|---|
discovery (default) | Unmet pain, pre-product-market-fit | 28 / 22 / 10 / 14 / 13 / 13 |
pipeline | Repeatable lead flow for a selling business | 22 / 20 / 20 / 14 / 12 / 12 |
design-partners | People likely to test and give feedback | 24 / 18 / 6 / 10 / 24 / 18 |
b2b | Companies with a current business trigger | 14 / 20 / 26 / 18 / 11 / 11 |
community | Explicit public requests and discussion | 28 / 18 / 8 / 14 / 22 / 10 |
b2b and pipeline also refuse to shortlist a prospect whose company is not identified.
| Depth | Shortlist cap | Minimum queries | Minimum pages opened | Minimum bucket x family coverage |
|---|---|---|---|---|
quick | 5 | 10 | 12 | 4 buckets x 3 families |
standard (default) | 10 | 18 | 20 | 5 buckets x 4 families |
deep | 20 | 35 | 35 | 7 buckets x 6 families |
Queries measure typing, pages measure reading, and only one of them finds customers. Record the count in coverage.pages_opened, including pages that never became candidates.
Copy the profile for the chosen mode into the run JSON verbatim. The auditor recomputes every score against it and fails if the declared weights drift from the mode's profile.
3. Plan and run the search
Read references/search-playbook.md. Write the query plan before searching: which buckets, which source families, which phrasings in the audience's own words. Log every query and its result count, including the zero-result ones. Collect candidate URLs; do not write prospect records yet.
4. Verify before you write
Read references/verification.md. Fetch each candidate URL, then decide:
signal_confirmed- page fetched, signal present on it, attributable to the named prospect. Only this may be shortlisted.url_ok- fetched, but the signal is hedged, attribution is inferred, or the date is missing. Review tier.unverified- not fetched, dead, or gated. Add todroppedwith the reason and stop. Do not write it up.
Writing a full prospect record and then discarding it because the page 404s wastes most of a run. Fetch first, write second.
Open what you are about to abandon. Any query that surfaced three or more results and kept none needs a note on its coverage entry, written after opening one of those results. The auditor fails a run where most abandoned queries carry no note.
This is the step that decides who is on the shortlist, and it is invisible unless you write it down. A run that logs 16 queries and clears every coverage floor can still have surfaced 75 results, opened 10, and taken its entire shortlist from three announcement-shaped queries, while a page of competitor reviews and three published RFPs went unread. The audit line will call that full coverage. Verify the contact route as well as the source; verification.md covers what to record.
Research limits: public professional and business information only; no bypassing logins, paywalls, access controls, rate limits or robots restrictions; no data brokers, leaked datasets, private groups, personal email discovery or phone enrichment; never infer or target on protected traits, health, financial hardship, politics, religion or sexuality.
5. Date every signal
Timing is scored from age, so an undated signal cannot be shortlisted. Work down this list and stop at the first that yields an answer: a visible publication date, a relative timestamp on the page converted to signal_age_days, or a timestamp encoded in the URL. Record which one you used in date_basis. verification.md gives the method for each.
Never estimate an age. Partial dates like 2026-08 are rejected, and so are future dates. A shortlist entry resting on a guessed date is the exact failure this skill exists to prevent.
Record fetched_at and final_url for every confirmed prospect. The auditor cannot watch you fetch a page; these make an unsupported claim take more work than doing it properly.
6. Score, gate and deduplicate
Read references/scoring.md. Score six dimensions 0-5 against the written anchors. Timing comes from age, capped by band; a named current trigger raises the cap by one.
pain_strength scores the prospect's own problem, in the job the product does, in their words. A problem their customers have, a problem they sell the fix for, or a market trend they commented on all score 0. Every consumer brand launches campaigns; that is not a signal.
commercial_fit asks whether this seller can close this buyer through the motion this seller has. Not whether the company is large. Revenue in the billions, a hundred countries and a published supplier portal are evidence of a procurement queue, and for a self-serve or founder-led seller that scores 2, not 4.
Drop anything under 50. Route to the review tier anything undated, older than max_age_days without the three still_open fields, scoring 0 or 1 on pain_strength with no named trigger, scoring 0 or 1 on commercial_fit against a user-supplied band, whose only contact route is inappropriate, or whose company is unidentified in b2b or pipeline mode. Nothing is deleted; the review tier shows each one with its reason.
Record signal_type and purchase_readiness rather than a single intent label. Someone hiring a designer has real intent and no intent to buy software. The two labels have ceilings tied to the dimensions beneath them: hiring cannot exceed adjacent, a pain-based signal needs pain_strength of 3 for strong, strong needs commercial_fit of 3, and direct needs 4. The auditor enforces all of them, because free-text fields drifting upward together is how a launch announcement became a high-intent buyer.
A page advertising for a freelancer, contractor or employee is hiring, however exactly the work matches the product. They are selecting a person, not a service, and hiring caps readiness at adjacent. Filing a job advert as direct_request walks past that cap, which is how a fifteen-dollar-an-hour freelance brief once topped a shortlist. The tell is in what the route asks for: a portfolio, CV, resume or rate card means hiring; scope, price, turnaround or a trial means buying.
Deduplicate on dedupe_key. Load ledger.json from prior runs in the same output directory and exclude what it has already surfaced, plus everything on the user's exclusion list.
Keep the shortlist spread across sources: no more than half from one source family or one domain. Whichever site fetches most reliably will otherwise take over the run, and a single-source shortlist inherits that site's bias.
If more than 60 percent of the shortlist scores 80+, the anchors were applied loosely. Re-score.
7. Record who, where and how to reach them
Separate company, person and handle. The company is who would buy; the person is who posted; the handle is their platform username. When no company is visible, set company to not identified and put the handle in handle, written as the platform writes it (u/name). The outputs lead with the handle in that case. Look for the company first: posters often name it in the thread, their profile or a link.
Give every prospect a channel_url: the page the user actually opens to make contact. Often the source thread, sometimes a contact form or profile. Open it, and record channel_evidence: a short quote or the on-page heading showing who the route is for. A shortlisted prospect claiming channel_fit: valid without it fails the audit.
Supplier and vendor registration portals are questionable at best, a queue rather than a conversation. Applicant, press and customer-support routes are inappropriate. When the first route you find is the wrong shape, spend one more fetch looking for the right one before writing the prospect off; large sites bury the business-enquiry page and rank the careers page.
The auditor reads channel_evidence, not just the URL. A job advert on Reddit or LinkedIn has a perfectly clean URL and gives itself away only in the words: "share your portfolio and updated resume" is an application route wherever it is hosted.
Leave additional_source_urls empty unless another page shows the same signal. It renders as "Also seen at", so putting the contact page there claims corroboration that does not exist.
Draft one opener per shortlisted prospect, under 90 words, grounded only in the fetched page. Do not send, submit, connect, follow, comment or write to a CRM unless the user separately authorises it in a later turn.
8. Build outputs and audit
Read references/output-schema.md, write run.json, then:
python3 <skill-dir>/scripts/validate_run.py run.json
python3 <skill-dir>/scripts/build_outputs.py run.json outputs/
<skill-dir> is this skill's own directory, the one this SKILL.md was loaded from. Build an absolute path from it. Do not use a path relative to the working directory: the working directory is the user's project, not the skill. Where the host substitutes ${CLAUDE_SKILL_DIR}, that variable is exactly this directory.
Fix every ERROR and re-run until it prints PASS. Report the audit line verbatim in the final response. build_outputs.py writes leads.csv (the operational list), leads-full.csv (every field), report.html, brief.md and an updated ledger.json. Keep run.json beside them as the source of truth.
Without Python: write brief.md and leads.csv by hand from the schema and work the auditor's checks manually.
9. Report
Return the audit line verbatim in your reply and in brief.md. It does not appear in report.html on a passing run, because a green banner saying "fine" is not what someone opens a prospect report to read.
Then the verdict, the top prospect, and absolute paths to the artifacts. Lead with what is verified and dated. Make the gaps visible.
Quality bar
- A prospect with no fetched source is not a prospect. One with no date is not a lead. One who cannot buy is a research interview. One that has stated no problem of its own is a market segment.
- Results surfaced and never opened are the run's real output. Count them, and open the ones you were about to skip.
- A run that examined thirty candidates and shortlisted none is a valid result. Report it plainly rather than padding.
- Ten verified, current matches beat fifty generic ones.
- Report what was searched and found nothing, not only what was found.
- Personalise from the source, never from invention.
- Label everything a hypothesis. A public signal is not consent, interest or intent to buy.