Files
Corey Haines f86637eace feat: add prospecting skill + truelist integration (#308)
* feat: add prospecting skill + truelist integration

New skill: skills/prospecting/
- SKILL.md (251 lines, well under 500 limit): branch picker for SaaS / B2B /
  Local SMB, shared 5-phase framework (ICP -> discovery -> qualify -> score ->
  output), compliance guardrails, tool selection quick-picks, output formats
- references/saas-prospecting.md: tech stack signals, funding/hiring triggers,
  SaaS-specific sources and qualification
- references/b2b-prospecting.md: industry/firmographic signals, trigger events,
  decision-maker mapping, B2B-specific sources
- references/local-prospecting.md: 4-tier website status classification,
  browser-assisted research workflow (generalized from the local-client-
  prospector pattern), proximity scoring
- references/data-sources.md: deep dives on Apollo, Clay, ZoomInfo, Clearbit,
  Hunter, Snov, Truelist, LinkedIn Sales Nav, BuiltWith, Crunchbase, RB2B,
  with sequencing recommendations across the three branches
- references/compliance.md: CAN-SPAM, GDPR, CASL, platform ToS (LinkedIn,
  Google Maps, Apollo/ZI/Clearbit), anti-patterns, audit checklist
- evals/evals.json: 6 evals (2 SaaS, 2 B2B, 1 Local SMB, 1 deliverability)

New integration:
- tools/integrations/truelist.md: email deliverability validation
  (Deliverable / Risky / Undeliverable / Unknown classification)

Registry + marketplace wiring:
- tools/REGISTRY.md: truelist row + new Email Verification category section
- .claude-plugin/marketplace.json: bumped to 2.1.0, prospecting added to
  plugin description
- VERSIONS.md: prospecting 1.0.0 + 2.1.0 changelog entry
- README.md: skill table re-synced, prospecting added to ASCII flow under
  Sales & GTM column

All 41 skills pass validation. sync-skills.js is idempotent.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(prospecting): add GitHub stargazers/forks/watchers as discovery channel

Net-new in this commit:
- tools/clis/github-prospects.js: zero-dep Node CLI with commands
  stargazers / forks / watchers / user / rate-limit. Pagination via Link header,
  optional --enrich for full profile data, --with-email / --with-company /
  --with-blog filters, --format csv|json output, --dry-run preview. Uses
  GITHUB_TOKEN for 5000/hr rate limit (vs 60/hr unauthenticated).
- tools/integrations/github.md: integration guide covering auth, rate limits,
  endpoints, workflows for SaaS prospecting, compliance notes (public API, not
  scraping), CLI reference.

Skill updates:
- skills/prospecting/SKILL.md: added GitHub to the tool selection quick picks
  and to the tool integrations table.
- skills/prospecting/references/saas-prospecting.md: added GitHub to Tier 3
  buying signals plus a dedicated "GitHub prospecting pattern (when audience
  is developers)" subsection with end-to-end workflow.
- skills/prospecting/references/data-sources.md: added GitHub deep-dive
  section between RB2B and Free fallbacks.

Registry:
- tools/REGISTRY.md: github row in Tool Index, new Developer Intent / GitHub
  category section.

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(prospecting): apply review suggestions

CLI hardening + optimization:
- github-prospects.js: encodeURIComponent on username path interpolation
  (defense in depth; GitHub usernames are restricted enough that this is safe
  in practice, but good hygiene).
- github-prospects.js: refactored enrichUsers to filter inline and support
  --target N early termination. Previously, --with-email on a 1000-star repo
  would enrich all 1000 users before filtering down to the ~50 that match.
  Now you can pass --target 25 to stop as soon as 25 matches are found,
  saving API quota on restrictive filters.
- github.md: documented the new --target flag.

Reverse cross-references (so prospecting is discoverable from sibling skills):
- cold-email: added prospecting as the natural upstream skill
- customer-research: added "Translating customer research into an ICP for
  outbound" hand-off to prospecting
- competitor-profiling: distinguished from prospecting ("this skill does deep
  research on specific accounts; prospecting builds the initial list")

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(truelist): align integration doc with actual OpenAPI spec

Source of truth: Truelist-Labs/truelist-openapi (OpenAPI 3.1).

The earlier integration doc had inferred (and wrong) endpoint paths, request
shapes, and status enum values. Corrected against the published spec:

Base URL: https://api.truelist.io
Endpoints (real):
- POST /api/v1/verify_inline?email=... (sync single, email is query param)
- POST /api/v1/verify (async bulk, body: {emails: [...]})
- GET /me (account info)

Real email_state enum:
- ok, email_invalid, risky, unknown, accept_all
(not the inferred "Deliverable / Risky / Undeliverable / Unknown")

Real email_sub_state enum:
- email_ok, is_disposable, is_role, unknown_error, failed_smtp_check

Also corrected:
- Truelist has an official MCP server (Truelist-Labs/truelist-mcp) — was
  marked as MCP unavailable
- Truelist has 7 official SDKs (Node, Python, Ruby, PHP, Go, Java, .NET) +
  framework integrations (Django, Laravel, Next.js, Rails, React, Svelte,
  Vue, WordPress) — was marked as SDK unavailable
- Native integrations with Mailchimp, Klaviyo, HubSpot, Zapier, Make, n8n,
  Clay, Salesforce, ActiveCampaign, Brevo, ConvertKit, Drip, BigCommerce,
  Go High Level — was unlisted
- Rate limits: 10 req/s per endpoint (was unspecified)

Files updated:
- tools/integrations/truelist.md: full rewrite against spec
- tools/REGISTRY.md: MCP and SDK columns now show ✓ for truelist; classifier
  note in the Email Verification section reflects real enum values
- skills/prospecting/evals/evals.json: eval #6 expected_output and assertions
  use real email_state values and mention the MCP server
- skills/prospecting/references/data-sources.md: Truelist deep-dive uses real
  endpoint paths, real enum values, and lists the MCP/SDK ecosystem

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(prospecting): add Firecrawl + Browserbase for single-target site research

Both tools are programmatic scrapers, but their use in prospecting is
strictly bounded: extract content from individual public business sites
(the prospect's own website URL), never from the platforms hosting them
(Google Maps, LinkedIn, Yelp, Apollo, etc.). This matches the line drawn
by the original local-client-prospector reference skill and our own
compliance section.

New integration docs:
- tools/integrations/firecrawl.md: REST + MCP + SDKs (Node/Python/Go/Rust);
  scrape / map / crawl / extract / search endpoints; explicit "when NOT to
  use" section listing the prohibited platforms.
- tools/integrations/browserbase.md: real Chromium via Playwright/Puppeteer
  or Stagehand (AI-friendly natural-language extraction); session
  recordings; useful when rendering or interaction is required.

Prospecting skill updates:
- SKILL.md: added Firecrawl + Browserbase to tool selection quick picks
  and tool integrations table.
- references/data-sources.md: new "Firecrawl / Browserbase (single-target
  site research)" section between RB2B and Free fallbacks. Includes the
  compliance line inline so the framing isn't lost.
- references/local-prospecting.md: optional "programmatic verification"
  paragraph in the browser research workflow — once you have a candidate's
  URL from manual Maps discovery, you can hit it programmatically.
- references/compliance.md: anti-pattern #1 now explicitly clarifies that
  Firecrawl/Browserbase are fine for the prospect's own website but not
  for the platforms hosting prospects.

Registry:
- tools/REGISTRY.md: firecrawl + browserbase rows in Tool Index, new "Site
  Scraping (single-target only)" category section with the compliance
  framing in the agent recommendation.

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-26 11:20:26 -07:00

8.0 KiB
Raw Permalink Blame History

Local SMB Prospecting Reference

For when the user sells to local small businesses — shops, gyms, restaurants, salons, clinics, professional services, contractors, real estate, fitness studios, dental practices.

Adapted from and generalized beyond the local-client-prospector pattern (browser-assisted discovery + website status classification + proximity scoring).


ICP Signals That Matter (Local SMB branch)

Operational signals

  • Active business — Google Business Profile updated, recent reviews, recent hours updates
  • Recent activity — open right now, regular hours posted, recent photos uploaded by owner
  • Customer engagement — owner responding to reviews, posts on social, active calendar (for service businesses)

Online presence signals (the core SMB qualification axis)

The reference local-client-prospector skill uses website status as the primary qualification — port this directly. Four classifications:

Status Definition Typical outcome
No site found No credible standalone website after cross-checked search Hot prospect for web/marketing service
Social only Facebook, Instagram, WhatsApp, Linktree, booking portal, marketplace page only — no standalone site Hot prospect for web/marketing service
Weak site Standalone site exists but outdated, broken, very thin, non-mobile-friendly, or missing clear contact/conversion flow Warm prospect for refresh / rebuild service
Has site Credible, modern standalone site exists Low prospect unless other signals apply (e.g., poor SEO, weak conversion design)

Proximity signals

  • Distance from the user's location or service area
  • Density — clusters of similar businesses in one area = neighborhood targeting opportunity
  • Travel time — useful when in-person discovery, install, or service delivery is required

Decay signals

  • Closed permanently (Google Maps banner)
  • Reviews paused or business listing reported as closed
  • Last activity (review, post) >12 months ago

Discovery Sources (Local SMB branch)

Primary

  • Google Maps (browser, manual) — search "category near [location]" and walk the visible results. Cross-check details. Don't bulk-extract.
  • Yelp — secondary verification; complementary categories
  • Bing Local / Apple Maps — different coverage on smaller businesses
  • Facebook Pages search — many SMBs are Facebook-only

Cross-verification

  • Business's own website (if any)
  • Industry directories (e.g., Healthgrades for medical, OpenTable for restaurants, Avvo for legal)
  • Local Chamber of Commerce listings
  • State business registries for incorporation status
  • Search results for "[business name] [city]" to discover non-Maps presence

Browser Research Workflow

  1. Open a browser and search Google Maps for the category near base_location
  2. Build a candidate list from visible local results, search results, and public directories
  3. For each candidate, inspect public sources to fill required fields
  4. Search the exact business name plus city/town to check whether a standalone website exists
  5. Classify website status per the table above
  6. Mark confidence: High (2+ sources), Medium (1 source + consistent evidence), Low (incomplete/ambiguous)

When the user explicitly asks for subagents AND subagents are available, split candidates into non-overlapping batches and ask each subagent to verify only website/social/contact status. Don't use subagents for the primary search if it slows progress.

Optional: programmatic verification with Firecrawl or Browserbase

Once you have a candidate's website URL (found via manual Maps/Yelp discovery), you can speed up website-status classification by hitting the URL programmatically:

  • Firecrawl for simple "is this site live, modern, mobile-friendly, conversion-flow-equipped" reads — returns clean markdown you can inspect
  • Browserbase when the candidate site requires JS rendering, has a cookie consent dialog, or you need session state

Strict line: use these on the individual business's URL. Don't point them at Google Maps, Yelp, or any platform whose ToS prohibits bulk extraction — discovery stays manual.

See data-sources.md for setup details.


Qualification Checklist (Local SMB branch)

  • Business is active (recent reviews or activity in last 6 months)
  • Category matches user's service offering
  • Distance / proximity within target radius
  • Website status classified
  • Phone or contact channel verified
  • At least one cross-source confirms business operates at the listed address
  • Not a duplicate / chain location / out-of-scope category
  • Not closed permanently

Lead Scoring (Local SMB)

Use this simple rubric (matches local-client-prospector pattern):

Score Criteria
Hot No site found OR social-only + phone present + active business + within target radius
Warm Weak site, poor online presentation, or marketplace/booking-page only
Cold Good website already present OR low confidence
Skip Closed, duplicate, outside radius, irrelevant category, or not a business prospect

Output Columns (Local SMB branch)

Chat table (≤15 rows):

| Score | Business | Category | Area | Distance | Website status | Website/Social | Phone | Why it's a prospect | Confidence |

CSV:

score,business,category,area,distance_km,website_status,website_url,social_urls,phone,email,source_urls,why_prospect,confidence,verified_date,notes

Rules:

  • Keep "Why it's a prospect" short and actionable
  • Use Not found instead of leaving blank fields
  • Include source links sparingly, not all of them
  • After the table, add Best first outreach targets with the top 3 leads and one practical reason each
  • If confidence is low, state exactly what remains uncertain

Top Outreach Targets Selection (Local SMB)

Prioritize for the top 3 hot leads:

  1. No site / social only + phone present = clearest service opportunity
  2. High review count = active, established business with real customers
  3. Owner-responded reviews = engaged owner = more likely to evaluate a vendor
  4. Industry alignment with your service specialty beats generic category match

Each top target rationale should be one sentence naming the gap and the signal: "No standalone website (cross-checked); 80+ Google reviews with owner replies; 2 km from target area."


Compliance Notes (Local SMB-specific)

The local branch is the most scraping-sensitive of the three motions. Specifically:

  • Google Maps Terms of Service prohibit bulk extraction. Treat browser visits as research, not as data acquisition.
  • Don't store full Google Maps Place IDs in your CRM — the ToS limits storage of Maps data.
  • Public business contact channels only: published phone, contact form, info@ email. Don't reach individual employees through their personal channels.
  • Owner/operator name when published on the business's own site is OK to use. If you only got it from LinkedIn, mark the source.

Common Mistakes (Local SMB)

  1. Bulk-scraping Google Maps — fastest way to violate ToS and lose the research channel.
  2. Treating Google Maps data as truth — listings go stale. Cross-check hours, status, and reviews.
  3. Skipping the website status cross-check — finding "no site" on Maps doesn't mean no site exists; do an exact-name web search before classifying.
  4. Targeting only the largest businesses — they're already covered by other providers. The 25 employee SMBs are the under-served opportunity.
  5. Generic outreach to all hot leads — local SMBs respond better to outreach that names their specific gap ("I noticed your menu isn't visible on mobile") than generic pitches.
  6. Ignoring chains and franchises as Skip — sometimes the franchisee is the buyer and they have local marketing authority. Verify before skipping.