{
  "markdown": "# 🎭 Generative Media Skills for AI Agents\n\n[![Powered by MuAPI](https://img.shields.io/badge/Powered%20by-MuAPI-6366f1?style=flat-square&logo=data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdCb3g9IjAgMCAyNCAyNCI+PHBhdGggZmlsbD0id2hpdGUiIGQ9Ik0xMiAyQzYuNDggMiAyIDYuNDggMiAxMnM0LjQ4IDEwIDEwIDEwIDEwLTQuNDggMTAtMTBTMTcuNTIgMiAxMiAyem0tMSAxNHYtNGgtMnYtMmg0djZoLTJ6bTAtOFY2aDJ2MmgtMnoiLz48L3N2Zz4=)](https://muapi.ai?utm_source=github&utm_medium=badge&utm_campaign=generative-media-skills)\n\n\n**The Ultimate Multimodal Toolset for Claude Code, Cursor, Gemini CLI, and OpenCode.**\nA high-performance, schema-driven architecture for AI agents to generate, edit, and display professional-grade images, videos, and audio — powered by the [muapi-cli](https://github.com/SamurAIGPT/muapi-cli).\n\n\n[🚀 Get Started](#-quick-start) | [🎬 Recipe Pack](#-recipe-pack) | [🎨 Expert Library](#-expert-library) | [⚙️ Core Primitives](#-core-primitives) | [🤖 MCP Server](#-mcp-server) | [📖 Reference](#-schema-reference)\n\n---\n\n<p align=\"center\"><a href=\"https://www.youtube.com/watch?v=SOXsxqnQGlc\"><img src=\"https://i.ytimg.com/vi/SOXsxqnQGlc/maxresdefault.jpg\" width=\"720\"></a></p>\n<p align=\"center\"><a href=\"https://www.youtube.com/watch?v=SOXsxqnQGlc\"><b>▶ Watch: Best AI Video Generator (API) in 2026 (Quality, Price, Uncensored, Editing)</b></a></p>\n\n## Related Projects\n\n- [minimax-music-3-api](https://github.com/SamurAIGPT/minimax-music-3-api) — Python SDK for MiniMax Music 3.0 text-to-music generation on Muapi.\n- [awesome-minimax-music-3-prompts](https://github.com/Anil-matcha/awesome-minimax-music-3-prompts) — Curated song prompts and lyrics-formatting guide for MiniMax Music 3.0.\n- [MiniMax-H3-API](https://github.com/Anil-matcha/MiniMax-H3-API) — Python SDK for MiniMax H3 video-generation workflows on Muapi.\n- [awesome-minimax-h3-prompts](https://github.com/Anil-matcha/awesome-minimax-h3-prompts) — Prompt gallery and runnable examples for the MiniMax H3 skills.\n- [Wan-3.0-API](https://github.com/Anil-matcha/Wan-3.0-API) — Python SDK and MCP server for Wan 3.0-compatible video-generation workflows.\n- [Wan-3.0-Prime-API](https://github.com/Anil-matcha/Wan-3.0-Prime-API) — Python SDK and MCP server for the higher-fidelity Wan 3.0 Prime tier.\n- [Open-Generative-AI](https://github.com/Anil-matcha/Open-Generative-AI) — Free self-hosted AI media studio — GUI alternative to these skills for the same model set\n- [Awesome-GPT-Image-2-API-Prompts](https://github.com/Anil-matcha/Awesome-GPT-Image-2-API-Prompts) — Curated GPT-Image-2 prompts to use with these skills\n- [Awesome-Gemini-Omni-API-Prompts](https://github.com/Anil-matcha/Awesome-Gemini-Omni-API-Prompts) — Curated Gemini Omni prompts for video generation\n- [Gemini-Omni-1.1-Flash-API](https://github.com/Anil-matcha/Gemini-Omni-1.1-Flash-API) — Python SDK and MCP server for Google's newly announced Gemini Omni 1.1 Flash update\n- [AI-Voice-Agent](https://github.com/Anil-matcha/AI-Voice-Agent) — Self-hosted AI voice agent for real-time voice conversations, sales calls, and customer support\n- [awesome-ai-image-models](https://github.com/Anil-matcha/awesome-ai-image-models) — compare AI image models by API, price & quality\n- [flux-3-video-api](https://github.com/SamurAIGPT/flux-3-video-api) — Python wrapper focused on FLUX 3 Text-to-Video and Image-to-Video\n- [ai-creator-academy](https://github.com/Anil-matcha/ai-creator-academy) — free curriculum teaching creators to monetize generative AI, built on these same skills\n- [Flux-3-Dev-API](https://github.com/Anil-matcha/Flux-3-Dev-API) — Python wrapper for Black Forest Labs' FLUX 3 (Dev variant) — text-to-image, image-to-image, text-to-video, image-to-video\n- [Grok-Imagine-Image-2-API](https://github.com/Anil-matcha/Grok-Imagine-Image-2-API) — Python SDK and MCP server for Grok Imagine Image 2.0 generation and editing through MuAPI\n- [midjourney-api](https://github.com/Anil-matcha/midjourney-api) — Python SDK for Midjourney V7, V8, and Niji image generation through MuAPI\n- [suno-api](https://github.com/Anil-matcha/suno-api) — Python SDK for Suno music, audio, and voice workflows through MuAPI\n- [awesome-flux-3-api-prompts](https://github.com/Anil-matcha/awesome-flux-3-api-prompts) — FLUX 3 API guide, prompts, and parameters\n- [seedance-2.5-mcp](https://github.com/Anil-matcha/seedance-2.5-mcp) — MCP server for generating Seedance 2.5 Preview videos through MuAPI.\n- [seedance-2-mcp](https://github.com/Anil-matcha/seedance-2-mcp) — MCP server for generating Seedance 2 videos through MuAPI.\n- [Text-to-Speech-API](https://github.com/Anil-matcha/Text-to-Speech-API) — narration and dialogue API examples for media workflows.\n- [Speech-to-Text-API](https://github.com/Anil-matcha/Speech-to-Text-API) — transcription and audio-understanding API examples.\n- [Voice-Cloning-API](https://github.com/Anil-matcha/Voice-Cloning-API) — consent-aware speaking and singing voice workflows.\n- [Image-Enhancement-API](https://github.com/Anil-matcha/Image-Enhancement-API) — image enhancement examples for creative pipelines.\n- [Video-Utilities-API](https://github.com/Anil-matcha/Video-Utilities-API) — video upscaling and sound-generation utility examples.\n- [AI-3D-Model-API](https://github.com/Anil-matcha/AI-3D-Model-API) — 3D asset generation comparison and examples.\n\n## ✨ Key Features\n\n- **🤖 Agent-Native Design** — CLI-powered scripts with structured JSON outputs, semantic exit codes, and `--jq` filtering for seamless agentic pipelines.\n- **🧠 Expert Knowledge Layer** — Domain-specific skills that bake in professional cinematography, atomic design, and branding logic.\n- **⚡ CLI-Powered Core** — All primitives delegate to [`muapi-cli`](https://www.npmjs.com/package/muapi-cli) — no curl, no JSON parsing, no boilerplate.\n- **🖼️ Direct Media Display** — Use the `--view` flag to automatically download and open generated media in your system viewer.\n- **📁 Local File Support** — Auto-upload images, videos, faces, and audio from your local machine to the CDN for processing.\n- **🌈 100+ AI Models** — One-click access to **Midjourney v7, Flux Kontext, Seedance 2.0, Kling 3.0, Veo3**, and more.\n- **🔌 MCP Server** — Run `muapi mcp serve` to expose all 19 tools directly to Claude Desktop, Cursor, or any MCP-compatible agent.\n\n---\n\n## 🏗️ Scalable Architecture\n\nThis repository uses a **Core/Library** split to ensure efficiency and high-signal discovery for LLMs:\n\n### ⚙️ Core Primitives (`/core`)\nThin wrappers around [`muapi-cli`](https://github.com/SamurAIGPT/muapi-cli) for raw API access.\n- `core/media/` — File upload\n- `core/edit/` — Image editing (prompt-based)\n- `core/platform/` — Setup, auth & result polling\n\n### 📚 Expert Library (`/library`)\nHigh-value skills that translate creative intent into technical directives.\n- **Cinema Director** (`/library/motion/cinema-director/`) — Technical film direction & cinematography.\n- **Nano-Banana** (`/library/visual/nano-banana/`) — Reasoning-driven image generation (Gemini 3 Style).\n- **UI Designer** (`/library/visual/ui-design/`) — High-fidelity mobile/web mockups (Atomic Design).\n- **Logo Creator** (`/library/visual/logo-creator/`) — Minimalist vector branding (Geometric Primitives).\n- **Seedance 2 (Doubao Video)** (`/library/motion/seedance-2/`) — Director-level cinematic video generation with text-to-video, image-to-video, and video extension with native audio-video sync.\n- **AI Clipping** (`/library/edit/ai-clipping/`) — Long video → ranked vertical short clips in one managed API call. Server-side transcription, virality ranking, dedupe, and face-tracked auto-crop — no local Whisper or LLM.\n- **YouTube Shorts** (`/library/social/youtube-shorts/`) — Platform-aware preset over AI Clipping (Shorts / TikTok / Reels / Feed defaults).\n\nPlus **41 ready-to-run workflow recipes** organized by output type — see [🎬 Recipe Pack](#-recipe-pack) below.\n\n---\n\n## 🎬 Recipe Pack\n\nForty-one LLM-orchestrated workflow recipes that combine multiple `muapi-cli` calls into named end-to-end pipelines (e.g. *photo of person → 3D action figure*, *product photo → cinematic 10s ad*). Each skill is a SKILL.md the agent reads and follows; bring your own consuming agent (Claude Code, Cursor, MCP) — these are recipes, not bash wrappers.\n\n**Motion / Video (16)**\n\n| Skill | Description |\n|:---|:---|\n| [3D Logo Animation](library/motion/3d-logo-animation/) | Transform a 2D logo into a premium 3D version and animate it with professional cinematic effects |\n| [AI Fight Scene Generator](library/motion/ai-fight-scene/) | High-cut-density action / fight scene — 16-cell storyboard image drives Seedance 2.0 i2v for shot-by-shot choreography |\n| [Animal Vlogger Video](library/motion/animal-video-generator/) | Hilarious, ultra-realistic anthropomorphic-animal vlogger acting like a human in a real-world setting |\n| [Cartoon Dance Animation](library/motion/cartoon-dance-animation/) | Convert a photo into a Pixar-style 3D cartoon, then animate using a reference dance/motion video |\n| [Character Story Video](library/motion/character-story-video/) | Multi-part animated story video — establish a consistent character then animate sequential scenes |\n| [Drone-Style Video](library/motion/drone-style-video/) | Aerial drone-perspective footage — bird's-eye sweeps, orbit shots, and flyover sequences |\n| [Giant Product Showcase](library/motion/giant-product-showcase/) | Dramatic giant-scale product visual (building-sized object next to a person), optionally animated |\n| [Jewelry Product Video](library/motion/jewelry-product-video/) | Luxury jewelry ad with high-end commercial cinematography and detailed macro animation |\n| [Music Video](library/motion/music-video/) | Short music video from a song theme — keyframes, animation per beat, matching music track |\n| [One-Shot Video](library/motion/one-shot-video/) | Single continuous cinematic shot — no cuts, one seamless flowing scene |\n| [Cinematic Product Ad](library/motion/product-ad-cinematic/) | Cinematic 5–10s product ad from a product photo + brand brief |\n| [Product Showcase Video](library/motion/product-showcase-video/) | Dynamic product showcase with explosive ingredient arrangement + realistic motion animation |\n| [Product Video Ad Maker](library/motion/product-video-ad-maker/) | High-end cinematic product video ad starting from a simple product photo |\n| [Talking Baby Video](library/motion/talking-baby-video/) | Viral-style talking-baby video with custom costumes and scripts |\n| [UGC Lifestyle Try-On](library/motion/ugc-lifestyle-try-on/) | UGC-style lifestyle photos & video of a person using your product — authentic, social-native |\n| [UGC Video Factory](library/motion/ugc-video-factory/) | Person photo + product photo + script → 10s vertical 9:16 UGC video ad with native dialogue (Nano-Banana Pro Edit → Seedance 2.0 VIP i2v) |\n\n**Social (5)**\n\n| Skill | Description |\n|:---|:---|\n| [Instagram Post](library/social/instagram-post/) | Polished on-brand Instagram post — hero image + caption + hashtags |\n| [Product Campaign Pack](library/social/product-campaign/) | Full multi-channel campaign — hero visuals, social assets, short ad video, platform crops |\n| [RedNote Cover](library/social/rednote-cover/) | Xiaohongshu (小红书) cover image — vibrant lifestyle aesthetic with typography overlay |\n| [Social Media Pack](library/social/social-pack/) | Re-render a hero image into Instagram / TikTok / Shorts / X aspect ratios |\n| [UGC Ads Workflow](library/social/ugc-ads-workflow/) | UGC video ad pipeline — combine selfie + product image, write script, animate |\n\n**Visual / Images & Design (21)**\n\n| Skill | Description |\n|:---|:---|\n| [Action Figure Generator](library/visual/action-figure-generator/) | Convert a photo of a person into a custom 3D action figure with collectible toy packaging |\n| [Ad Creative Set](library/visual/ad-creative/) | High-converting ad set — hero image, copy variations, platform crops for Meta / Google / LinkedIn |\n| [Amazon Product Listing Pack](library/visual/amazon-product-listing/) | Full Amazon listing image set — hero, lifestyle, infographic, comparison/detail closeups |\n| [Blog Header](library/visual/blog-header/) | Professional 1200×628 blog header image with optional title composition guidance |\n| [Brand Kit](library/visual/brand-kit/) | Cohesive brand visual kit — logo concept, color palette, typography pairings |\n| [Brochure Designer](library/visual/brochures/) | Multi-page brochure — cover, inner spread, back — for business, real estate, events, launches |\n| [Couple Grid Creator](library/visual/couple-grid-creator/) | Stylized 6-box grid of a couple in romantic poses, each pose framed inside cardboard packaging |\n| [Brand Design Guide](library/visual/design-guide/) | Comprehensive design guide — palette, typography, UI components, visual identity rules |\n| [Fashion Try-On](library/visual/fashion-try-on/) | Virtually try outfits by combining a person's photo + clothing item, optional fashion model video |\n| [Floor Plan Rendering](library/visual/floor-plan-rendering/) | Design a 2D floor plan and convert into a realistic 3D architectural rendering |\n| [Interior Design](library/visual/interior-design/) | Pro interior design visualizations — redesign rooms, generate concepts, visualize furniture styles |\n| [Interior Design Visualizer](library/visual/interior-design-visualizer/) | Generate an empty room and fill it with stylish furniture / decor; or redesign an existing room |\n| [Keyboard Art Maker](library/visual/keyboard-art-maker/) | Artistic top-down photos of keyboard keycaps arranged to spell custom messages |\n| [Logo + Branding Package](library/visual/logo-branding/) | Logo + full branding package — variations (dark/light/icon), palette, mockups |\n| [Logo Generator](library/visual/logo-generator/) | Quick single-shot polished logo — fast, clean vector aesthetic with accurate brand-name text |\n| [Multi-Angle Reshoot](library/visual/multi-angle-reshoot/) | Re-render a subject from dramatic camera angles (fish-eye, bird's-eye, low, macro) — identity preserved |\n| [Multi-Angle Shots](library/visual/multi-angle-shots/) | Full multi-angle product shot set — front, side, back, top-down, 45° |\n| [Selfie with Celebrities](library/visual/selfie-with-celebrities/) | Realistic behind-the-scenes selfie of the user with a celebrity; optional cinematic long-take |\n| [Storyboard Generator](library/visual/storyboard/) | Generate N keyframes for a short story or scene sequence (image only, no video) |\n| [URL to Design](library/visual/url-to-design/) | Analyze a website URL and generate a redesigned, improved UI with modern aesthetics |\n| [YouTube Thumbnail](library/visual/youtube-thumbnail/) | High-CTR YouTube thumbnail — striking imagery, bold text placement, emotional face/subject |\n\nEach recipe declares its `inputs` and a `Steps` body. Pass the inputs and let your agent execute the steps via `muapi` CLI calls (or raw API for endpoints that don't yet have a CLI alias — see the per-skill *Notes for the Executing Agent* footer).\n\n---\n\n## 🚀 Quick Start\n\n### 1. Install the muapi CLI\n\nThe core scripts require [`muapi-cli`](https://www.npmjs.com/package/muapi-cli). Install it once:\n\n```bash\n# via npm (recommended — no Python required)\nnpm install -g muapi-cli\n\n# via pip\npip install muapi-cli\n\n# or run without installing\nnpx muapi-cli --help\n```\n\n### 2. Configure Your API Key\n\n```bash\n# Interactive setup\nmuapi auth configure\n\n# Or pass directly\nmuapi auth configure --api-key \"YOUR_MUAPI_KEY\"\n\n# Get your key at https://muapi.ai/dashboard?utm_source=github&utm_medium=readme&utm_campaign=generative-media-skills\n```\n\n### 3. Install the Skills\n\n```bash\n# Install all skills to your AI agent\nnpx skills add SamurAIGPT/Generative-Media-Skills --all\n\n# Or install a specific skill\nnpx skills add SamurAIGPT/Generative-Media-Skills --skill muapi-media-generation\n\n# Install to specific agents\nnpx skills add SamurAIGPT/Generative-Media-Skills --all -a claude-code -a cursor\n```\n\n### 4. Generate Your First Image\n\n```bash\nmuapi image generate \"a cyberpunk city at night\" --model flux-dev\n\n# Download the result automatically\nmuapi image generate \"a sunset over mountains\" --model hidream-fast --download ./outputs\n\n# Extract just the URL (agent-friendly)\nmuapi image generate \"product on white bg\" --model flux-schnell --output-json --jq '.outputs[0]'\n```\n\n### 5. Run an Expert Skill\n\n```bash\n# Use Nano-Banana reasoning to generate a 2K masterpiece\nbash library/visual/nano-banana/scripts/generate-nano-art.sh \\\n  --file ./my-source-image.jpg \\\n  --subject \"a glass hummingbird\" \\\n  --style \"macro photography\" \\\n  --resolution \"2k\" \\\n  --view\n```\n\n### 6. Direct a Cinematic Scene\n\n```bash\ncd library/motion/cinema-director\n\n# Create a 10-second epic reveal\nbash scripts/generate-film.sh \\\n  --subject \"a cybernetic dragon over Tokyo\" \\\n  --intent \"epic\" \\\n  --model \"kling-v3.0-pro\" \\\n  --duration 10 \\\n  --view\n\n# Animate a reference image into video\nbash library/motion/seedance-2/scripts/generate-seedance.sh \\\n  --mode i2v \\\n  --file ./concept.jpg \\\n  --subject \"camera slowly pulls back to reveal the full landscape\" \\\n  --intent \"reveal\" \\\n  --view\n\n# Extend an existing video\nbash library/motion/seedance-2/scripts/generate-seedance.sh \\\n  --mode extend \\\n  --request-id \"YOUR_REQUEST_ID\" \\\n  --subject \"camera continues pulling back to reveal the vast city\" \\\n  --duration 10\n```\n\n---\n\n\n### OpenCode\n\n```bash\n# Clone the repo and set the MUAPI_API_KEY env var\ngit clone https://github.com/SamurAIGPT/Generative-Media-Skills\nexport MUAPI_API_KEY=your_key_here\n\n# Skills auto-load from .opencode/skills/ when you run opencode in this directory\nopencode\n```\n## 🤖 MCP Server\n\nRun muapi as a **Model Context Protocol server** so Claude Desktop, Cursor, or any MCP-compatible agent can call generation tools directly — no shell scripts needed.\n\n```bash\nmuapi mcp serve\n```\n\n**Claude Desktop config** (`~/Library/Application Support/Claude/claude_desktop_config.json`):\n\n```json\n{\n  \"mcpServers\": {\n    \"muapi\": {\n      \"command\": \"muapi\",\n      \"args\": [\"mcp\", \"serve\"],\n      \"env\": { \"MUAPI_API_KEY\": \"your-key-here\" }\n    }\n  }\n}\n```\n\nThis exposes **19 structured tools** with full JSON Schema input/output definitions:\n\n| Tool | Description |\n|------|-------------|\n| `muapi_image_generate` | Text-to-image (14 models) |\n| `muapi_image_edit` | Image-to-image editing (11 models) |\n| `muapi_video_generate` | Text-to-video (13 models) |\n| `muapi_video_from_image` | Image-to-video (16 models) |\n| `muapi_audio_create` | Music generation (Suno) |\n| `muapi_audio_from_text` | Sound effects (MMAudio) |\n| `muapi_enhance_upscale` | AI upscaling |\n| `muapi_enhance_bg_remove` | Background removal |\n| `muapi_enhance_face_swap` | Face swap image/video |\n| `muapi_enhance_ghibli` | Ghibli style transfer |\n| `muapi_edit_lipsync` | Lip sync to audio |\n| `muapi_edit_clipping` | AI highlight extraction |\n| `muapi_predict_result` | Poll prediction status |\n| `muapi_upload_file` | Upload local file → URL |\n| `muapi_keys_list` | List API keys |\n| `muapi_keys_create` | Create API key |\n| `muapi_keys_delete` | Delete API key |\n| `muapi_account_balance` | Get credit balance |\n| `muapi_account_topup` | Add credits (Stripe checkout) |\n\n---\n\n## ⚡ Agentic Pipeline Examples\n\n```bash\n# Submit async, capture request_id, poll when ready\nREQUEST_ID=$(muapi video generate \"a dog running on a beach\" \\\n  --model kling-master --no-wait --output-json --jq '.request_id' | tr -d '\"')\n\n# ... do other work ...\n\nmuapi predict wait \"$REQUEST_ID\" --download ./outputs\n\n# Pipe a prompt from another command\ngenerate_prompt | muapi image generate - --model flux-dev\n\n# Chain: upload → edit → download\nURL=$(muapi upload file ./photo.jpg --output-json --jq '.url' | tr -d '\"')\nmuapi image edit \"make it look like a painting\" --image \"$URL\" \\\n  --model flux-kontext-pro --download ./outputs\n```\n\n---\n\n## 📖 Schema Reference\n\nThis repository includes a streamlined `schema_data.json` that core scripts use at runtime to:\n- **Validate Model IDs**: Ensures the requested model exists.\n- **Resolve Endpoints**: Automatically maps model names to API endpoints.\n- **Check Parameters**: Validates supported `aspect_ratio`, `resolution`, and `duration` values.\n\nDiscover all available models via the CLI:\n\n```bash\nmuapi models list\nmuapi models list --category video --output-json\n```\n\n---\n\n## 🔧 Compatibility\n\nOptimized for the next generation of AI development environments:\n- **Claude Code** — Direct terminal execution via tools + MCP server mode.\n- **Gemini CLI / Cursor / Windsurf** — Seamless integration as local scripts.\n- **MCP** — Full Model Context Protocol server with typed input/output schemas.\n- **CI/CD** — `--output-json`, `--jq`, semantic exit codes for scripting.\n\n---\n\n## 📄 License\nMIT © 2026\n",
  "bytes": 20791,
  "sha": "6d03ba8386aa9578d7d79adbac64238c7a31047ee880bd9382d3ead52ffec6eb",
  "repo_slug": "samuraigpt/generative-media-skills",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/skl_samuraigpt_generative_media_skills_muapi_0be6623b/readme"
}