Hermes for Artists emblem Hermes for ArtistsDeconstruct Agents Lab
The Vault · Creative Hermes skills

Skills Vault

A curated catalog of real, installable Hermes Agent skills for art & creative work, each one linking to its source.

Image Generation & Editing

Image Generation & Editing

Turns Hermes into an image studio: local/cloud diffusion pipelines, node-based workflows, and photo-realistic generation with your own trained likeness. The core toolkit for artists who want production-grade stills without leaving the agent.

  • comfyui Generate images, video & audio with ComfyUI — install, manage nodes/models, run workflows via comfy-cli + REST/WS.
  • stable-diffusion Text-to-image with Stable Diffusion via Diffusers — img2img, inpainting, custom pipelines.
  • black-forest-labs/skills Official first-party FLUX model skills for image generation, maintained by the FLUX creators themselves.
  • PhotoGPT Generates realistic photos with personal-model support inside Hermes Agent.
  • codex-vision Use Codex CLI for image analysis and generation, including GPT-image workflows and identity-preserving reference-image generation.
  • agnes-image Generate and edit images through Agnes AI's Agnes-Image-2.0-Flash REST API, including text-to-image, image-to-image editing, and multi-image composition.
  • runninghub Run RunningHub.ai image, video, audio, music, 3D, text understanding, and custom ComfyUI workflow APIs from Hermes.
  • comfyui-image ComfyUI image-generation bundle for Flux, SDXL, SD3, img2img, ControlNet, inpainting, and upscaling workflows.
  • grok-imagine-video Generate, edit, and animate images and videos through xAI's Grok Imagine API from Hermes.
  • vertex-imagen Generate high-quality images with Google Vertex AI Imagen 3 and future Gemini image models, with aspect-ratio and local-output support.
  • RunComfy GPT Image 2 Generate and edit images with OpenAI GPT Image 2 on RunComfy, emphasizing embedded text, logos, typography, and instruction precision.
  • baoyu-image-gen API-based image generation across OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, GLM-Image, MiniMax, Jimeng, Seedream, Replicate, and Agnes.
  • baoyu-cover-image Generate article cover images with configurable type, palette, rendering style, text treatment, mood, and aspect ratio.
  • marketing image Create, generate, edit, and optimize marketing images such as blog heroes, social graphics, product mockups, profile banners, listing visuals, and brand assets.
Video & Motion

Video & Motion

From math-explainer animation to HTML-driven social video to full prose-to-storyboard pipelines — the skills artists and filmmakers reach for once a project moves from stills to moving image.

  • manim-video Manim CE animations: 3Blue1Brown-style math/algorithm videos.
  • hyperframes HTML-as-source video — animated titles, social overlays, captioned talking-heads, shader transitions, rendered to MP4/WebM.
  • kanban-video-orchestrator Multi-agent video pipeline on Hermes Kanban — scope a brief, build a team, route scenes, monitor to render.
  • ascii-video Convert video/audio to colored ASCII MP4/GIF.
  • Storyboard Turns prose into editable, multi-frame film storyboards — parses, critiques, and learns your directing style.
  • gif-search Search and download GIFs from Tenor for visual content and chat media.
  • youtube-content Fetch YouTube transcripts and transform them into chapters, summaries, threads, blog posts, and timestamped quote extracts.
  • agnes-video Generate cinematic videos with Agnes AI's Agnes-Video-V2.0 model, including text-to-video, image-to-video, multi-image video, and keyframe animation.
  • comfyui-video ComfyUI video-generation bundle for LTX, Seedance, Wan, AnimateDiff, camera motion, and character-consistency workflows.
  • RunComfy AI Video Generation Router skill for RunComfy text-to-video, image-to-video, and video-extension models including HappyHorse, Wan, Seedance, Kling, Veo, Hailuo, and Dreamina.
  • RunComfy AI Avatar Video Create AI avatar, talking-head, lip-sync, virtual-presenter, and audio-driven video through RunComfy model routing.
  • GSAP Core Official GSAP skill for JavaScript animation fundamentals: tweens, easing, duration, stagger, defaults, responsive animation, and reduced motion.
  • GSAP ScrollTrigger Official GSAP skill for scroll-linked animation, pinning, scrubbed timelines, parallax, and ScrollTrigger patterns.
  • Stitch remotion Generate walkthrough videos from Stitch design projects using Remotion with transitions, zoom effects, and text overlays.
  • marketing video Plan and produce marketing videos using AI generation models, AI avatars, programmatic frameworks, product demos, explainers, and social clips.
  • vlog-auto-edit Complete workflow for AI agents to automatically edit travel vlogs from raw footage. Analyzes footage, plans narrative, selects clips, and renders.
Creative Coding & Real-Time Visuals

Creative Coding & Real-Time Visuals

Code-driven, generative, and real-time visual work — sketches, shaders, 3D scenes, and live visual-instrument rigs — for artists who think in code rather than timelines.

  • p5js p5.js sketches: generative art, shaders, interactive and 3D work.
  • touchdesigner-mcp Control TouchDesigner via the twozero MCP — operators, parameters, Python, real-time visuals (36 tools).
  • pretext DOM-free text layout — ASCII art, kinetic typography, text-as-geometry games as single-file HTML.
  • blender-mcp Control Blender from Hermes via socket — create objects, materials, animations, run bpy Python.
  • huashu-design HTML-native high-fidelity prototypes, interactive demos, slide decks, animations, design variants, expert review, and MP4/GIF export workflows.
  • blob-tracker Audio-reactive video rendering skill that tracks blobs in video and uses audio features to modulate visual intensity.
Voice, Music & Sound

Voice, Music & Sound

Composition, sound generation, and the audio-analysis skills that support them — for artists building songs, scores, or agent voices rather than pictures.

  • songwriting-and-ai-music Songwriting craft paired with Suno AI music prompts.
  • heartmula Suno-like song generation from lyrics and tags.
  • audiocraft Meta AudioCraft: MusicGen text-to-music, AudioGen text-to-sound.
  • whisper OpenAI Whisper — multilingual speech-to-text and translation, useful for subtitling and voice-workflow tooling.
  • songsee Generate spectrograms and multi-panel audio feature visualizations including mel, chroma, MFCC, flux, loudness, and tempogram views.
  • RunComfy AI Music Generate and edit AI music through RunComfy, routing between ElevenLabs Music and ACE Step for vocal tracks, instrumentals, jingles, extensions, and repairs.
  • pockettts-agent-bridge Real-time, low-latency Text-to-Speech (TTS) for AI agents with voice cloning support running locally.
Diagrams, Illustration & Design Systems

Diagrams, Illustration & Design Systems

Structured visual communication — hand-drawn diagrams, infographics, design-token systems, and full design-system references for artists who also do UI/product/editorial design work.

  • excalidraw Hand-drawn Excalidraw JSON diagrams — architecture, flow, sequence.
  • architecture-diagram Dark-themed SVG architecture/cloud/infra diagrams as HTML.
  • concept-diagrams Flat, minimal light/dark SVG diagrams as HTML — physics, chemistry, anatomy, floor plans, lifecycles.
  • baoyu-infographic Infographics: 21 layouts x 21 styles.
  • baoyu-article-illustrator Article illustrations with consistent type, style, and palette.
  • design-md Author, validate, and export Google's DESIGN.md token spec files.
  • popular-web-designs 54 real design systems (Stripe, Linear, Vercel) as ready HTML/CSS references.
  • claude-design Design one-off HTML artifacts — landing pages, decks, prototypes.
  • sketch Throwaway HTML mockups — 2-3 design variants to compare quickly.
  • typeui-hermes Use design skills to generate better UI with Hermes.
  • nano-pdf Edit PDF text, typos, and titles from natural-language instructions using the nano-pdf CLI.
  • powerpoint Create, read, edit, analyze, and design .pptx decks, slides, templates, speaker notes, and presentation artifacts.
  • pptx-author Build PowerPoint decks headlessly with python-pptx, with conventions for model-backed pitch decks, IC memos, and earnings notes.
  • code-wiki Generate codebase wiki documentation with architecture, Mermaid flowcharts, class diagrams, and sequence diagrams.
  • Nous style guide Nous Research brand identity skill for infographics, diagrams, blog art, character cards, and social graphics using the official monochrome visual system.
  • Stitch generate-design Generate new UI screens from text prompts or images, edit existing screens, and create design variants through Stitch MCP.
  • UI UX Pro Max design-system Design-token architecture, component specifications, CSS variables, spacing and typography scales, and brand-compliant presentation generation.
  • UI UX Pro Max banner-design Design social, ad, website-hero, creative-asset, and print banners with multiple art directions and AI-generated visuals.
  • UI UX Pro Max slides Create strategic HTML presentations using Chart.js, design tokens, responsive layouts, copywriting formulas, and contextual slide strategies.
  • baoyu-slide-deck Generate professional slide deck images from content, with outlines, style instructions, and individual shareable slide images.
Comics, Pixel & Playful Art

Comics, Pixel & Playful Art

Lighter, faster, meme-and-nostalgia-driven visual formats — for artists making comics, retro game art, or internet-native jokes rather than fine-art output.

  • baoyu-comic Educational, biography, and tutorial knowledge-comics.
  • pixel-art Pixel art with era-accurate palettes (NES, Game Boy, PICO-8).
  • ascii-art ASCII art via pyfiglet, cowsay, boxes, and image-to-ascii conversion.
  • meme-generation Generate real meme .png images — pick a template, overlay text with Pillow.
  • baoyu-xhs-images Generate social-media infographic image-card series with 12 visual styles, 8 layouts, and cartoon-style cards.
Writing & Creative Ideation

Writing & Creative Ideation

The non-visual craft side — voice, tone, and structured idea-generation for artists whose medium is language as much as image.

  • humanizer Humanize text: strip AI-isms and add real voice.
  • creative-ideation Generate ideas using named methods drawn from creative practice.
  • research-paper-writing End-to-end ML/AI research paper pipeline for experiment design, analysis, drafting, review, revision, and conference submission.
Vision & Multimodal Understanding

Vision & Multimodal Understanding

The perception layer under a lot of creative automation: classifying, segmenting, and reading images so an agent can curate references, cut out subjects, or caption a body of work.

  • clip OpenAI CLIP — zero-shot image classification, image-text matching, cross-modal retrieval.
  • llava Visual instruction tuning and multi-turn image chat (CLIP + Vicuna/LLaMA).
  • segment-anything SAM: zero-shot image segmentation via points, boxes, and masks.
  • ocr-and-documents Extract text, tables, images, OCR, equations, and markdown from PDFs and scanned documents using pymupdf or marker-pdf.
Hardware & Physical Making

Hardware & Physical Making

For creatives whose output leaves the screen entirely — agent-assisted circuit design and fabrication-ready artifacts.

  • Hermes-volta Natural-language analog circuit design agent — PySpice simulation, KiCad-compatible artifacts, Gerbers, dashboard, Telegram delivery.