Files
Corey Haines f86637eace feat: add prospecting skill + truelist integration (#308)
* feat: add prospecting skill + truelist integration

New skill: skills/prospecting/
- SKILL.md (251 lines, well under 500 limit): branch picker for SaaS / B2B /
  Local SMB, shared 5-phase framework (ICP -> discovery -> qualify -> score ->
  output), compliance guardrails, tool selection quick-picks, output formats
- references/saas-prospecting.md: tech stack signals, funding/hiring triggers,
  SaaS-specific sources and qualification
- references/b2b-prospecting.md: industry/firmographic signals, trigger events,
  decision-maker mapping, B2B-specific sources
- references/local-prospecting.md: 4-tier website status classification,
  browser-assisted research workflow (generalized from the local-client-
  prospector pattern), proximity scoring
- references/data-sources.md: deep dives on Apollo, Clay, ZoomInfo, Clearbit,
  Hunter, Snov, Truelist, LinkedIn Sales Nav, BuiltWith, Crunchbase, RB2B,
  with sequencing recommendations across the three branches
- references/compliance.md: CAN-SPAM, GDPR, CASL, platform ToS (LinkedIn,
  Google Maps, Apollo/ZI/Clearbit), anti-patterns, audit checklist
- evals/evals.json: 6 evals (2 SaaS, 2 B2B, 1 Local SMB, 1 deliverability)

New integration:
- tools/integrations/truelist.md: email deliverability validation
  (Deliverable / Risky / Undeliverable / Unknown classification)

Registry + marketplace wiring:
- tools/REGISTRY.md: truelist row + new Email Verification category section
- .claude-plugin/marketplace.json: bumped to 2.1.0, prospecting added to
  plugin description
- VERSIONS.md: prospecting 1.0.0 + 2.1.0 changelog entry
- README.md: skill table re-synced, prospecting added to ASCII flow under
  Sales & GTM column

All 41 skills pass validation. sync-skills.js is idempotent.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(prospecting): add GitHub stargazers/forks/watchers as discovery channel

Net-new in this commit:
- tools/clis/github-prospects.js: zero-dep Node CLI with commands
  stargazers / forks / watchers / user / rate-limit. Pagination via Link header,
  optional --enrich for full profile data, --with-email / --with-company /
  --with-blog filters, --format csv|json output, --dry-run preview. Uses
  GITHUB_TOKEN for 5000/hr rate limit (vs 60/hr unauthenticated).
- tools/integrations/github.md: integration guide covering auth, rate limits,
  endpoints, workflows for SaaS prospecting, compliance notes (public API, not
  scraping), CLI reference.

Skill updates:
- skills/prospecting/SKILL.md: added GitHub to the tool selection quick picks
  and to the tool integrations table.
- skills/prospecting/references/saas-prospecting.md: added GitHub to Tier 3
  buying signals plus a dedicated "GitHub prospecting pattern (when audience
  is developers)" subsection with end-to-end workflow.
- skills/prospecting/references/data-sources.md: added GitHub deep-dive
  section between RB2B and Free fallbacks.

Registry:
- tools/REGISTRY.md: github row in Tool Index, new Developer Intent / GitHub
  category section.

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(prospecting): apply review suggestions

CLI hardening + optimization:
- github-prospects.js: encodeURIComponent on username path interpolation
  (defense in depth; GitHub usernames are restricted enough that this is safe
  in practice, but good hygiene).
- github-prospects.js: refactored enrichUsers to filter inline and support
  --target N early termination. Previously, --with-email on a 1000-star repo
  would enrich all 1000 users before filtering down to the ~50 that match.
  Now you can pass --target 25 to stop as soon as 25 matches are found,
  saving API quota on restrictive filters.
- github.md: documented the new --target flag.

Reverse cross-references (so prospecting is discoverable from sibling skills):
- cold-email: added prospecting as the natural upstream skill
- customer-research: added "Translating customer research into an ICP for
  outbound" hand-off to prospecting
- competitor-profiling: distinguished from prospecting ("this skill does deep
  research on specific accounts; prospecting builds the initial list")

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(truelist): align integration doc with actual OpenAPI spec

Source of truth: Truelist-Labs/truelist-openapi (OpenAPI 3.1).

The earlier integration doc had inferred (and wrong) endpoint paths, request
shapes, and status enum values. Corrected against the published spec:

Base URL: https://api.truelist.io
Endpoints (real):
- POST /api/v1/verify_inline?email=... (sync single, email is query param)
- POST /api/v1/verify (async bulk, body: {emails: [...]})
- GET /me (account info)

Real email_state enum:
- ok, email_invalid, risky, unknown, accept_all
(not the inferred "Deliverable / Risky / Undeliverable / Unknown")

Real email_sub_state enum:
- email_ok, is_disposable, is_role, unknown_error, failed_smtp_check

Also corrected:
- Truelist has an official MCP server (Truelist-Labs/truelist-mcp) — was
  marked as MCP unavailable
- Truelist has 7 official SDKs (Node, Python, Ruby, PHP, Go, Java, .NET) +
  framework integrations (Django, Laravel, Next.js, Rails, React, Svelte,
  Vue, WordPress) — was marked as SDK unavailable
- Native integrations with Mailchimp, Klaviyo, HubSpot, Zapier, Make, n8n,
  Clay, Salesforce, ActiveCampaign, Brevo, ConvertKit, Drip, BigCommerce,
  Go High Level — was unlisted
- Rate limits: 10 req/s per endpoint (was unspecified)

Files updated:
- tools/integrations/truelist.md: full rewrite against spec
- tools/REGISTRY.md: MCP and SDK columns now show ✓ for truelist; classifier
  note in the Email Verification section reflects real enum values
- skills/prospecting/evals/evals.json: eval #6 expected_output and assertions
  use real email_state values and mention the MCP server
- skills/prospecting/references/data-sources.md: Truelist deep-dive uses real
  endpoint paths, real enum values, and lists the MCP/SDK ecosystem

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(prospecting): add Firecrawl + Browserbase for single-target site research

Both tools are programmatic scrapers, but their use in prospecting is
strictly bounded: extract content from individual public business sites
(the prospect's own website URL), never from the platforms hosting them
(Google Maps, LinkedIn, Yelp, Apollo, etc.). This matches the line drawn
by the original local-client-prospector reference skill and our own
compliance section.

New integration docs:
- tools/integrations/firecrawl.md: REST + MCP + SDKs (Node/Python/Go/Rust);
  scrape / map / crawl / extract / search endpoints; explicit "when NOT to
  use" section listing the prohibited platforms.
- tools/integrations/browserbase.md: real Chromium via Playwright/Puppeteer
  or Stagehand (AI-friendly natural-language extraction); session
  recordings; useful when rendering or interaction is required.

Prospecting skill updates:
- SKILL.md: added Firecrawl + Browserbase to tool selection quick picks
  and tool integrations table.
- references/data-sources.md: new "Firecrawl / Browserbase (single-target
  site research)" section between RB2B and Free fallbacks. Includes the
  compliance line inline so the framing isn't lost.
- references/local-prospecting.md: optional "programmatic verification"
  paragraph in the browser research workflow — once you have a candidate's
  URL from manual Maps discovery, you can hit it programmatically.
- references/compliance.md: anti-pattern #1 now explicitly clarifies that
  Firecrawl/Browserbase are fine for the prospect's own website but not
  for the platforms hosting prospects.

Registry:
- tools/REGISTRY.md: firecrawl + browserbase rows in Tool Index, new "Site
  Scraping (single-target only)" category section with the compliance
  framing in the agent recommendation.

All 41 skills still pass validation. sync-skills.js still no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-26 11:20:26 -07:00

258 lines
9.4 KiB
JavaScript
Executable File

#!/usr/bin/env node
const TOKEN = process.env.GITHUB_TOKEN
const BASE_URL = 'https://api.github.com'
const USER_AGENT = 'marketingskills-prospects-cli'
function parseArgs(args) {
const result = { _: [] }
for (let i = 0; i < args.length; i++) {
const arg = args[i]
if (arg.startsWith('--')) {
const key = arg.slice(2)
const next = args[i + 1]
if (next && !next.startsWith('--')) {
result[key] = next
i++
} else {
result[key] = true
}
} else {
result._.push(arg)
}
}
return result
}
const args = parseArgs(process.argv.slice(2))
async function api(path, opts = {}) {
const url = path.startsWith('http') ? path : `${BASE_URL}${path}`
const headers = {
'Accept': 'application/vnd.github+json',
'X-GitHub-Api-Version': '2022-11-28',
'User-Agent': USER_AGENT,
}
if (TOKEN) headers['Authorization'] = `Bearer ${TOKEN}`
if (args['dry-run']) {
return {
_dry_run: true,
method: 'GET',
url,
headers: { ...headers, Authorization: TOKEN ? 'Bearer ***' : undefined },
}
}
const res = await fetch(url, { headers })
const rateLimitRemaining = res.headers.get('x-ratelimit-remaining')
const rateLimitReset = res.headers.get('x-ratelimit-reset')
if (res.status === 401 || res.status === 403) {
const body = await res.text()
return {
error: `HTTP ${res.status}`,
hint: TOKEN
? 'Token rejected — check GITHUB_TOKEN scopes (public_repo is enough for public data).'
: 'Set GITHUB_TOKEN env var to raise rate limit from 60/hr to 5000/hr.',
rate_limit_remaining: rateLimitRemaining,
rate_limit_reset_unix: rateLimitReset,
body,
}
}
if (!res.ok) {
const body = await res.text()
return { error: `HTTP ${res.status}`, body }
}
const data = await res.json()
const linkHeader = res.headers.get('link') || ''
const nextMatch = linkHeader.match(/<([^>]+)>;\s*rel="next"/)
return {
data,
next: nextMatch ? nextMatch[1] : null,
rate_limit_remaining: rateLimitRemaining,
}
}
async function paginate(path, { limit, perPage = 100 } = {}) {
const initial = path.includes('?') ? `${path}&per_page=${perPage}` : `${path}?per_page=${perPage}`
const all = []
let next = initial
let lastRate = null
while (next) {
const result = await api(next)
if (result._dry_run) return result
if (result.error) return result
lastRate = result.rate_limit_remaining
all.push(...result.data)
if (limit && all.length >= limit) {
return { data: all.slice(0, limit), rate_limit_remaining: lastRate, truncated: true }
}
next = result.next
}
return { data: all, rate_limit_remaining: lastRate, truncated: false }
}
async function getUser(login) {
const result = await api(`/users/${encodeURIComponent(login)}`)
if (result._dry_run || result.error) return result
return result.data
}
function matchesFilter(user, opts) {
if (!user) return false
if (opts['with-email'] && !user.email) return false
if (opts['with-company'] && !user.company) return false
if (opts['with-blog'] && !user.blog) return false
if (opts['type'] && user.type !== opts.type) return false
return true
}
async function enrichUsers(users, opts = {}, { concurrency = 5, targetCount } = {}) {
const matched = []
for (let i = 0; i < users.length; i += concurrency) {
const batch = users.slice(i, i + concurrency)
const profiles = await Promise.all(batch.map(u => getUser(u.login)))
for (const profile of profiles) {
if (!profile || profile.error) continue
if (matchesFilter(profile, opts)) matched.push(profile)
}
if (targetCount && matched.length >= targetCount) {
return matched.slice(0, targetCount)
}
}
return matched
}
function toCSV(users) {
const cols = ['login', 'name', 'company', 'email', 'blog', 'location', 'bio', 'twitter_username', 'public_repos', 'followers', 'created_at', 'html_url']
const escape = (v) => {
if (v === null || v === undefined) return ''
const s = String(v).replace(/\r?\n/g, ' ')
if (s.includes(',') || s.includes('"')) return `"${s.replace(/"/g, '""')}"`
return s
}
const lines = [cols.join(',')]
for (const u of users) {
lines.push(cols.map(c => escape(u[c])).join(','))
}
return lines.join('\n')
}
function parseRepo(input) {
if (!input) return null
const trimmed = input.replace(/^https?:\/\/github\.com\//, '').replace(/\.git$/, '').replace(/\/$/, '')
const parts = trimmed.split('/')
if (parts.length < 2) return null
return { owner: parts[0], repo: parts[1] }
}
async function main() {
const [command, ...rest] = args._
let result
switch (command) {
case 'stargazers': {
const repo = parseRepo(rest[0])
if (!repo) { result = { error: 'Usage: stargazers <owner/repo> [--limit N] [--target N] [--enrich] [--with-email] [--with-company] [--with-blog] [--format csv|json]' }; break }
const limit = args.limit ? parseInt(args.limit, 10) : undefined
const target = args.target ? parseInt(args.target, 10) : undefined
const page = await paginate(`/repos/${repo.owner}/${repo.repo}/stargazers`, { limit })
if (page._dry_run || page.error) { result = page; break }
let users = page.data
if (args.enrich || args['with-email'] || args['with-company'] || args['with-blog'] || args.type) {
users = await enrichUsers(users, args, { targetCount: target })
}
if (args.format === 'csv') {
console.log(toCSV(users))
return
}
result = { count: users.length, rate_limit_remaining: page.rate_limit_remaining, truncated: page.truncated, users }
break
}
case 'forks': {
const repo = parseRepo(rest[0])
if (!repo) { result = { error: 'Usage: forks <owner/repo> [--limit N] [--target N] [--enrich] [--with-email] [--with-company] [--with-blog] [--format csv|json]' }; break }
const limit = args.limit ? parseInt(args.limit, 10) : undefined
const target = args.target ? parseInt(args.target, 10) : undefined
const page = await paginate(`/repos/${repo.owner}/${repo.repo}/forks`, { limit })
if (page._dry_run || page.error) { result = page; break }
const forkOwners = page.data.map(f => f.owner)
let users = forkOwners
if (args.enrich || args['with-email'] || args['with-company'] || args['with-blog'] || args.type) {
users = await enrichUsers(forkOwners, args, { targetCount: target })
}
if (args.format === 'csv') {
console.log(toCSV(users))
return
}
result = { count: users.length, rate_limit_remaining: page.rate_limit_remaining, truncated: page.truncated, users }
break
}
case 'watchers': {
const repo = parseRepo(rest[0])
if (!repo) { result = { error: 'Usage: watchers <owner/repo> [--limit N] [--target N] [--enrich] [--with-email] [--with-company] [--with-blog] [--format csv|json]' }; break }
const limit = args.limit ? parseInt(args.limit, 10) : undefined
const target = args.target ? parseInt(args.target, 10) : undefined
const page = await paginate(`/repos/${repo.owner}/${repo.repo}/subscribers`, { limit })
if (page._dry_run || page.error) { result = page; break }
let users = page.data
if (args.enrich || args['with-email'] || args['with-company'] || args['with-blog'] || args.type) {
users = await enrichUsers(users, args, { targetCount: target })
}
if (args.format === 'csv') {
console.log(toCSV(users))
return
}
result = { count: users.length, rate_limit_remaining: page.rate_limit_remaining, truncated: page.truncated, users }
break
}
case 'user': {
const login = rest[0]
if (!login) { result = { error: 'Usage: user <username>' }; break }
result = await getUser(login)
break
}
case 'rate-limit': {
const res = await api('/rate_limit')
result = res._dry_run || res.error ? res : res.data
break
}
default:
result = {
error: 'Unknown command',
usage: {
stargazers: 'stargazers <owner/repo> [--limit N] [--target N] [--enrich] [--with-email] [--with-company] [--with-blog] [--type User|Organization] [--format csv|json]',
forks: 'forks <owner/repo> [--limit N] [--target N] [--enrich] [--with-email] [--with-company] [--with-blog] [--type User|Organization] [--format csv|json]',
watchers: 'watchers <owner/repo> [--limit N] [--target N] [--enrich] [--with-email] [--with-company] [--with-blog] [--format csv|json]',
user: 'user <username>',
'rate-limit': 'rate-limit',
},
notes: [
'Set GITHUB_TOKEN env var for 5000 req/hr (vs 60/hr unauthenticated).',
'Token needs only public_repo scope for public data; no scope is required to list public stargazers/forks.',
'--enrich fetches each users full profile (1 extra request per user). Use with --limit on large repos.',
'--with-email / --with-company / --with-blog imply --enrich.',
'--target N stops enrichment as soon as N users match the filters (saves API quota on restrictive filters).',
'--format csv outputs prospecting-ready CSV; default JSON.',
'Pair with Apollo, Clay, Hunter, or Truelist to fill in missing emails.',
],
}
}
console.log(JSON.stringify(result, null, 2))
}
main().catch(err => {
console.error(JSON.stringify({ error: err.message }))
process.exit(1)
})