From c08ae5ff3d9fe46c47bfc93cf3b300c373078902 Mon Sep 17 00:00:00 2001 From: Corey Haines <34802794+coreyhaines31@users.noreply.github.com> Date: Mon, 18 May 2026 22:24:09 -0700 Subject: [PATCH 1/3] fix: refresh image + video model lists to current releases (May 2026) (#311) Video (skills/video/SKILL.md): - Sora 2 promoted from "limited availability" caveat to a first-class model - Kling updated to 2.5/3.0 from Kuaishou - Added Seedance (ByteDance) as a fast, low-cost batch option - Added Hailuo / MiniMax for character-consistency use case - Pika updated to 2.x - Added Hunyuan Video / Wan 2 (open-weight, self-hosted) - New "Quick picks" guide for cinematic / batch / character-consistent / self-hosted / storyboard / image-to-video use cases - Trigger phrases updated to include Sora, Seedance, Hailuo, MiniMax, Hunyuan, Wan Image (skills/image/SKILL.md): - Gemini Image labelled with Nano Banana / Nano Banana Pro family names - Flux entry annotates Pro 1.1, Kontext, Dev, Schnell variants (Kontext for in-image editing) - Ideogram bumped to 3.0 - GPT Image entry renamed to "ChatGPT Images 2.0 / GPT Image" reflecting the newer family - Midjourney bumped to v7 - Added Recraft V3 for vector + brand illustration use case - Stable Diffusion 3.5 / SDXL specifics noted - Updated "When to Use Which" decision tree to match - Trigger phrases updated to include Flux Kontext, ChatGPT Images, Nano Banana, Recraft All 40 skills still pass validation. Co-authored-by: Claude Opus 4.7 --- skills/image/SKILL.md | 37 +++++++++++++++++++++---------------- skills/video/SKILL.md | 24 +++++++++++++++++------- 2 files changed, 38 insertions(+), 23 deletions(-) diff --git a/skills/image/SKILL.md b/skills/image/SKILL.md index 99cadeb..2fbcbca 100644 --- a/skills/image/SKILL.md +++ b/skills/image/SKILL.md @@ -1,6 +1,6 @@ --- name: image -description: "When the user wants to create, generate, edit, or optimize images for marketing — blog heroes, social graphics, product mockups, profile banners, listing visuals, or brand assets. Also use when the user mentions 'AI image generation,' 'generate an image,' 'create a graphic,' 'product mockup,' 'hero image,' 'social media graphic,' 'banner image,' 'cover photo,' 'profile banner,' 'listing screenshot,' 'Flux,' 'Midjourney,' 'DALL-E,' 'GPT Image,' 'Ideogram,' 'Gemini image,' 'Canva,' 'Figma,' 'image optimization,' 'compress images,' 'WebP,' or 'OG image.' Use this for general-purpose marketing image creation and optimization. For paid ad image creative and platform-specific ad specs, see ad-creative. For video production, see video." +description: "When the user wants to create, generate, edit, or optimize images for marketing — blog heroes, social graphics, product mockups, profile banners, listing visuals, or brand assets. Also use when the user mentions 'AI image generation,' 'generate an image,' 'create a graphic,' 'product mockup,' 'hero image,' 'social media graphic,' 'banner image,' 'cover photo,' 'profile banner,' 'listing screenshot,' 'Flux,' 'Flux Kontext,' 'Midjourney,' 'DALL-E,' 'GPT Image,' 'ChatGPT Images,' 'Ideogram,' 'Gemini image,' 'Nano Banana,' 'Recraft,' 'Stable Diffusion,' 'Canva,' 'Figma,' 'image optimization,' 'compress images,' 'WebP,' or 'OG image.' Use this for general-purpose marketing image creation and optimization. For paid ad image creative and platform-specific ad specs, see ad-creative. For video production, see video." metadata: version: 2.0.0 --- @@ -55,36 +55,41 @@ Generate original images from text prompts. The fastest way to create unique mar | Model | Best For | Text in Images | API | Cost | |-------|----------|:-:|-----|------| -| **Gemini Image** (Google) | All-around, editing, text rendering | Good | [Gemini API](https://ai.google.dev/gemini-api/docs/image-generation) | Check [pricing](https://ai.google.dev/gemini-api/docs/pricing) | -| **Flux** (Black Forest Labs) | Photorealism, brand consistency, batch | Limited | [BFL API](https://docs.bfl.ai/), Replicate, fal.ai | Check [pricing](https://docs.bfl.ai/quick_start/pricing) | -| **Ideogram** | Typography, branded graphics | Best | [Ideogram API](https://developer.ideogram.ai/) | Check [pricing](https://about.ideogram.ai/api-pricing) | -| **GPT Image** (OpenAI) | General purpose, ChatGPT integration | Good | [OpenAI API](https://platform.openai.com/docs/guides/image-generation) | Check [pricing](https://platform.openai.com/docs/pricing) | -| **Midjourney** | Artistic, high-aesthetic | Poor | No official API | Subscription-based | -| **Stable Diffusion** | Self-hosted, customizable | Varies | Open source | Free (GPU costs) | +| **Gemini Image** (Google, "Nano Banana" / Nano Banana Pro) | All-around, editing, multi-image reference, text rendering | Good | [Gemini API](https://ai.google.dev/gemini-api/docs/image-generation) | Check [pricing](https://ai.google.dev/gemini-api/docs/pricing) | +| **Flux** (Black Forest Labs — Pro 1.1, Kontext, Dev, Schnell) | Photorealism, brand consistency, batch; Kontext for in-image editing | Limited | [BFL API](https://docs.bfl.ai/), Replicate, fal.ai | Check [pricing](https://docs.bfl.ai/quick_start/pricing) | +| **Ideogram 3.0** | Typography, branded graphics, accurate text rendering | Best | [Ideogram API](https://developer.ideogram.ai/) | Check [pricing](https://about.ideogram.ai/api-pricing) | +| **ChatGPT Images 2.0 / GPT Image** (OpenAI) | General purpose, ChatGPT integration, native editing | Good | [OpenAI API](https://platform.openai.com/docs/guides/image-generation) | Check [pricing](https://platform.openai.com/docs/pricing) | +| **Midjourney v7** | Artistic, high-aesthetic, art-directed visuals | Improved | No official API; Discord + Web | Subscription-based | +| **Recraft V3** | Vector + brand-consistent illustrations, design assets | Strong | [Recraft API](https://www.recraft.ai/docs) | Per-credit | +| **Stable Diffusion 3.5 / SDXL** | Self-hosted, customizable, fine-tunable | Varies | Open source | Free (GPU costs) | -**Note:** DALL-E 3 is deprecated. OpenAI's current image models are the GPT Image family (`gpt-image-1`, etc.). +**Note:** DALL-E 3 is fully deprecated. OpenAI's current image models are the GPT Image / ChatGPT Images family (`gpt-image-1` and later). ### When to Use Which ``` Need text/headlines in the image? -├── Yes → Ideogram (best), Gemini (good), GPT Image (decent) +├── Yes → Ideogram 3.0 (best), Gemini (good), GPT Image / ChatGPT Images (decent) └── No ↓ -Need product/brand consistency across images? -├── Yes → Flux (multi-image reference) +Need product/brand consistency across many images? +├── Yes → Flux (multi-image reference), Gemini Nano Banana Pro, Recraft V3 └── No ↓ -Need to edit an existing image? -├── Yes → Gemini (native editing), Flux Flex +Need to edit an existing image (in-place)? +├── Yes → Gemini (native editing), Flux Kontext, ChatGPT Images └── No ↓ -Need highest visual quality? -├── Yes → Flux Pro, Midjourney +Need vector / illustrative brand assets? +├── Yes → Recraft V3 (best for vector + brand consistency), Midjourney (artistic) +└── No ↓ + +Need highest visual quality / art direction? +├── Yes → Flux Pro 1.1, Midjourney v7 └── No ↓ Need volume at low cost? -└── Flux Klein, Gemini Flash +└── Flux Schnell, Gemini Flash, Stable Diffusion (self-hosted) ``` ### Prompting Basics diff --git a/skills/video/SKILL.md b/skills/video/SKILL.md index 06dcdfa..4f6917b 100644 --- a/skills/video/SKILL.md +++ b/skills/video/SKILL.md @@ -1,6 +1,6 @@ --- name: video -description: "When the user wants to create, generate, or produce video content using AI tools or programmatic frameworks. Also use when the user mentions 'video production,' 'AI video,' 'Remotion,' 'Hyperframes,' 'HeyGen,' 'Synthesia,' 'Veo,' 'Runway,' 'Kling,' 'Pika,' 'video generation,' 'AI avatar,' 'talking head video,' 'programmatic video,' 'video template,' 'explainer video,' 'product demo video,' 'video pipeline,' or 'make me a video.' Use this for video creation, generation, and production workflows. For video content strategy and what to post, see social. For paid video ad creative, see ad-creative." +description: "When the user wants to create, generate, or produce video content using AI tools or programmatic frameworks. Also use when the user mentions 'video production,' 'AI video,' 'Remotion,' 'Hyperframes,' 'HeyGen,' 'Synthesia,' 'Veo,' 'Sora,' 'Runway,' 'Kling,' 'Seedance,' 'Hailuo,' 'MiniMax,' 'Pika,' 'Hunyuan,' 'Wan,' 'video generation,' 'AI avatar,' 'talking head video,' 'programmatic video,' 'video template,' 'explainer video,' 'product demo video,' 'video pipeline,' or 'make me a video.' Use this for video creation, generation, and production workflows. For video content strategy and what to post, see social. For paid video ad creative, see ad-creative." metadata: version: 2.0.0 --- @@ -41,7 +41,7 @@ Pick the right tool for the job: | Approach | Best For | Tools | When to Use | |----------|----------|-------|-------------| | **Programmatic** | Templated, data-driven, batch video | Remotion, Hyperframes | Product updates, personalized videos, recurring content | -| **AI Generation** | Original footage from text/image prompts | Veo, Runway, Kling, Pika | B-roll, hero shots, creative visuals you can't film | +| **AI Generation** | Original footage from text/image prompts | Veo 3, Sora 2, Runway, Kling, Seedance | B-roll, hero shots, creative visuals you can't film | | **AI Avatars** | Talking-head presenter without filming | HeyGen, Synthesia | Explainers, tutorials, multilingual content | | **Editing/Repurposing** | Cutting long-form into short clips | Descript, Opus Clip, CapCut | Podcast/webinar → social clips | @@ -130,12 +130,22 @@ Generate original footage from text or image prompts. Use for B-roll, hero visua | Model | Resolution | Max Duration | Best For | Cost | |-------|-----------|-------------|----------|------| -| **Veo 3** (Google) | Up to 1080p (4K varies) | Variable | Highest quality, synced audio | API-based | -| **Runway Gen-4** | Up to 4K | ~10 sec/gen | Motion control, temporal consistency | $12-76/mo | -| **Kling 3.0** | Up to 1080p | Up to 2 min | Volume production, lowest cost | $0.029/sec | -| **Pika** | 1080p | Short clips | Fast generation, effects | Per-credit | +| **Veo 3** (Google) | Up to 1080p (4K varies) | Variable | Top overall quality, synced audio | API-based | +| **Sora 2** (OpenAI) | Up to 1080p | Up to ~20 sec | Cinematic + synced audio, ChatGPT/API integration | API + ChatGPT | +| **Runway Gen-4** | Up to 4K | ~10 sec/gen | Motion control, temporal consistency, edit-style workflows | $12-76/mo | +| **Kling 2.5/3.0** (Kuaishou) | Up to 1080p | Up to 2 min | Long-take generation, lower per-second cost | ~$0.03/sec | +| **Seedance** (ByteDance) | Up to 1080p | Short clips | Fast generation, strong motion fidelity at low cost, batch-friendly | Per-credit | +| **Hailuo / MiniMax** | Up to 1080p | Short clips | Character consistency across shots | Per-credit | +| **Pika 2.x** | 1080p | Short clips | Quick effects, image-to-video, lower bar to entry | Per-credit | +| **Hunyuan Video / Wan 2** | 720p–1080p | Variable | Open-source self-hosted; full control, no API fees | Free (GPU) | -**Sora (OpenAI)** has had limited availability and reliability issues. Check current status before recommending. +**Quick picks**: +- **Highest quality + audio**: Veo 3 or Sora 2 +- **Batch / volume / cost**: Kling, Seedance +- **Character consistency across multiple shots**: Hailuo +- **Self-hosted, brand-controlled**: Hunyuan Video or Wan 2 (open weights) +- **Storyboard → video workflow**: Runway, LTX Studio +- **Image-to-video from a still you already have**: Kling, Pika, Runway ### Prompting for Video Models From 8b0a54eef93d521a5ffabcee4aaef782b2638de8 Mon Sep 17 00:00:00 2001 From: Corey Haines <34802794+coreyhaines31@users.noreply.github.com> Date: Mon, 18 May 2026 22:24:17 -0700 Subject: [PATCH 2/3] fix(ai-seo): align with Google's official AI features optimization guide (#313) Source: https://developers.google.com/search/docs/fundamentals/ai-optimization-guide The skill was framed as "optimize for AI systems," which mis-states Google's explicit position: AI Overviews and AI Mode are powered by core Search ranking, so SEO best practices ARE the optimization strategy. Don't write separate content "for AI." The structural patterns the skill recommends (FAQ schema, comparison tables, 40-60 word answer blocks, llms.txt, /pricing.md) still help non-Google AI engines (ChatGPT, Claude, Perplexity, Copilot) materially. The skill now calls out the platform split explicitly rather than implying universal benefit. Major changes: - New section "Google's Official Stance vs. Multi-Platform Reality": Frames Google's "don't optimize for AI" position alongside the reality that other AI engines reward extractable structure. Quotes Google's own phrasing. - New section "Query Fan-Out (Google AI Search)": Explains Google's documented behavior of generating concurrent related queries. Reframes content strategy from per-query targeting to topical cluster coverage. - New section "Agentic Experiences": Browser agents accessing sites via visual rendering, DOM inspection, and accessibility tree. Semantic HTML, JS-free meaningful render, clean a11y tree. Mentions the emerging Universal Commerce Protocol (UCP). - New section "What NOT to Do": Explicit Google guidance: don't write for AI, don't chunk content for AI, don't scale variations (scaled content abuse spam policy), don't pursue inauthentic mentions, don't block AI crawlers if you want citation. - Machine-Readable Files section reframed: Now opens with Google's stance ("not required") + the reason to include them anyway (non-Google AI engines). - Schema markup note reframed: Notes Google's position that structured data is "not required for generative AI search" but still recommended. - Search Console expectations added: Sets correct expectation that no AI-specific Search Console reporting exists; standard SEO metrics + third-party tools are the measurement options. - Moved "AI SEO for Different Content Types" to references/content-types.md: Kept SKILL.md under 500 lines. Added a local-business/ecom subsection (Merchant Center + Google Business Profile + Business Agent) per Google's emphasis. All 40 skills still pass validation. Co-authored-by: Claude Opus 4.7 --- skills/ai-seo/SKILL.md | 128 ++++++++++++++-------- skills/ai-seo/references/content-types.md | 71 ++++++++++++ 2 files changed, 156 insertions(+), 43 deletions(-) create mode 100644 skills/ai-seo/references/content-types.md diff --git a/skills/ai-seo/SKILL.md b/skills/ai-seo/SKILL.md index 2f60334..49bd12c 100644 --- a/skills/ai-seo/SKILL.md +++ b/skills/ai-seo/SKILL.md @@ -66,6 +66,45 @@ In traditional search, you need to rank on page 1. In AI search, a well-structur - Optimized content gets cited 3x more often than non-optimized - Statistics and citations boost visibility by 40%+ across queries +### Google's Official Stance vs. Multi-Platform Reality + +This is important to read once before doing anything else. + +**Google's position** ([AI features optimization guide](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide)): +> "The best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems." + +Google explicitly says: +- **No special markup or files are required** for AI Overviews or AI Mode +- **Don't chunk content for AI** — write for people, organize with normal headings and paragraphs +- **Don't write separate content for AI** — that risks "scaled content abuse" spam policy +- **Helpful, reliable, people-first content** wins — same E-E-A-T standards as regular Search +- **No AI-specific Search Console reporting** — use standard SEO metrics + +**Other AI engines (ChatGPT, Claude, Perplexity, Copilot) behave differently:** +- They actively reward extractable structure — passages, FAQs, comparison tables, definition blocks +- They parse `llms.txt`, structured pricing pages, and machine-readable files when present +- They cite third-party sources (Reddit, Wikipedia, review sites) more heavily than top-ranked pages + +**What this means for the work:** +- The structural patterns in this skill (40–60 word answer blocks, FAQ schema, comparison tables) help **non-Google AI engines** materially. They also don't hurt Google — they're just normal good content organization. +- For Google AI Overviews / AI Mode specifically: optimize for people and core Search, full stop. Strong E-E-A-T, original information, semantic HTML, clean indexability. +- For ChatGPT/Claude/Perplexity: layer on the extractable structure + llms.txt + machine-readable files. + +When in doubt, default to "write for people, organize for clarity" — that satisfies both camps. + +### Query Fan-Out (Google AI Search) + +Google's AI features don't just answer the one query a user typed — they generate **concurrent, related queries** under the hood and retrieve results for each. + +Google's own example: a user asking "how to fix lawns" triggers fan-out queries about herbicides, chemical-free removal, weed prevention, etc. The AI synthesizes across all of them. + +**Implications:** +- Single-page-per-keyword targeting is less effective. Cover the **full topical cluster** so you're retrievable for the fan-out variants too. +- Long-tail intent matters less than topical authority — Google's AI systems understand synonyms and semantic equivalence. +- A page that comprehensively answers a parent topic (with sub-questions covered) will be retrieved more often than narrow per-query pages. + +**Action**: when planning content, brainstorm the 5–10 related queries the AI is likely to fan out to and make sure your content (or your site as a whole) covers them. + --- ## AI Visibility Audit @@ -228,6 +267,10 @@ AI systems don't just cite your website — they cite where you appear. ### Machine-Readable Files for AI Agents +> **Google's stance**: not required for AI Overviews or AI Mode. Their guide explicitly says you don't need new markup, AI files, or markdown to appear in generative AI search. +> +> **Why include them anyway**: non-Google AI engines (ChatGPT, Claude, Perplexity) and autonomous buying agents do reward extractable structure. The files below help with those engines without harming Google. + AI agents aren't just answering questions — they're becoming buyers. When an AI agent evaluates tools on behalf of a user, it needs structured, parseable information. If your pricing is locked in a JavaScript-rendered page or a "contact sales" wall, agents will skip you and recommend competitors whose information they can actually read. Add these machine-readable files to your site root: @@ -284,7 +327,32 @@ Structured data helps AI systems understand your content. Key schemas: | Reviews | `Review`, `AggregateRating` | Trust signals | | Organization | `Organization` | Entity recognition | -Content with proper schema shows 30-40% higher AI visibility. For implementation, use the **schema** skill. +Content with proper schema shows 30-40% higher AI visibility on non-Google AI engines. **Google's note**: structured data is "not required for generative AI search" but is recommended for overall SEO strategy. For implementation, use the **schema** skill. + +--- + +## Agentic Experiences + +Beyond AI search engines summarizing content, autonomous agents are starting to access sites directly — clicking, reading, comparing, even buying on behalf of users. Google's guide flags this as an emerging category to plan for. + +**How agents access your site:** +- **Visual rendering** — they screenshot/read the page like a user would +- **DOM inspection** — they parse the page's HTML structure +- **Accessibility tree** — they rely on the same semantic information assistive tech uses (labels, roles, landmarks, headings) + +**What to do:** +- **Render meaningful content without heavy JS gymnastics** — if the page is blank until 4 frameworks finish loading, agents see blank +- **Semantic HTML** — use `
`, `