跪拜 Guibai
← Back to the summary

9 AI Skills That Automate the Entire Video Production Pipeline

Making videos really isn't something a human should have to do.

I used to think writing articles was hair-pulling enough, until I started messing with video myself. After finishing the script, I'd freeze up like crazy when recording in front of the camera; after finally finishing the recording, I had to listen frame by frame during editing to find every "uh" and "that"; then I had to search everywhere for copyright-free B-roll footage, background music, make subtitles...

After going through the whole process, half my life was gone.

To save my hair, I spent a week breaking down the tedious steps of making videos into 9 AI Skills that can run automatically.

To put it bluntly, it's about outsourcing mental pain, recording awkwardness, and aesthetic disasters to AI.

But before talking about the video pipeline, let me first plug a practical tool that has recently gained an excellent reputation in the community and lets you "lie flat" in the graphic content track — the [Redya Xiaohongshu Graphic Skill].

If you usually need to update Xiaohongshu graphics frequently in addition to making videos, this tool can save you most of the effort.

Handy Recommendation: Redya Xiaohongshu Graphic Skill (redya-xiaohongshu)

This is an exclusive Skill provided by Redya AI. You can seamlessly install it into mainstream Agent tools like WorkBuddy, Codex, Claude Code, or Cursor.

Redya AI is an AI graphic generation tool focused on Xiaohongshu creation, suitable for Xiaohongshu AI graphic generation, video-to-Xiaohongshu graphics, public account article-to-Xiaohongshu graphics, as well as product seeding and e-commerce product graphic creation. It can start from a topic or existing content, first generate an overall outline, then generate titles, post body text, page-by-page image outlines, and image generation prompts, and finally batch-generate covers and content pages. Regular graphics can be linked to personal or brand knowledge bases, and product graphics can be bound to product profiles and select different seeding Skills. Users can create on the web interface, or let Agents like Codex, Claude Code, WorkBuddy call Redya AI.

Its core capabilities are very solid: you just need to input a topic, and it can automatically generate Xiaohongshu-style viral titles, post body text, and detailed prompts for each page's image, and even continue to batch-generate a complete set of exquisite images.

Whether you are doing regular knowledge-sharing graphics or product-promotion graphics, it can handle it easily. Especially its product mode, which supports you uploading product materials to establish an exclusive "product profile". When creating later, directly associate the profile, and the generated graphics and images will be very precise.

You can get its detailed information through the GitHub repository below:

https://github.com/rnthking/redya-skills

To use this Skill, you need a personal API Key. You can log in directly to the Redya AI official website:

https://hy.ithinkai.cn

You can get it with one click on the "Redya Skills" page. This Key is your personal account credential; remember to keep it safe and don't leak it in public screenshots or code repositories.

After getting the API Key, you can send the following installation task directly to your Agent (such as WorkBuddy, Codex, Claude Code, or Cursor) and let it automatically complete the environment check and configuration writing:

Please help me install the 'Redya Xiaohongshu Graphic' capability to the current AI environment:

1. Download the content of https://hy.ithinkai.cn/openapi/skill.md and save it as ~/.claude/skills/redya-xiaohongshu/SKILL.md
2. Set environment variable REDYA_API_KEY=rk-yourAPIKey
3. Set environment variable REDYA_BASE_URL=https://hy.ithinkai.cn

After installation, I want to use it to generate Xiaohongshu graphic notes.

After installation, you can call it directly in the Agent.

One way is to "review the outline first, then generate images". You can let the Agent first generate the title, body text, and image outline. After you check and fine-tune it to your satisfaction, let it batch-generate images.

If you don't need step-by-step confirmation, you can also directly give it a topic and let the Agent do it in one step, generating the copy and the complete set of images all at once.

In addition to using it in the Agent command line, you can also log in directly to the Redya AI web interface. In the visual interface, you can more intuitively input topics, modify copy, and perform batch image generation and fine-tuning.

If you need to do product promotion or e-commerce, you can use its "Product Graphic Mode". First, enter product information and real images in the backend. When generating graphics, associate the corresponding product profile, so the AI can create by referencing real product details, avoiding the awkward situation where the visuals don't match the copy.

It's worth mentioning that the Agent terminal and the web interface are fully connected. They share the same Redya account and credits. The history you generate in the Agent can also be found in the "My Creations" section on the web interface, making it convenient for you to do subsequent editing and downloading at any time.

If you are used to calling directly in a terminal or development tool, installing redya-xiaohongshu is the most convenient choice; if you prefer intuitive visual operation, just use the Redya AI web interface directly.

Okay, back to video.

To avoid making everyone dizzy, I've divided these 9 video Skills into three groups based on the "satisfaction points" during creation.

Let's see how they help us do the dirty work.

Group 1: Mental Liberation (Topic Selection, Script, Subtitles)

Writing scripts and making subtitles are the hardest hit areas for brain cell consumption. This group of Skills is here to save our brains.

1. Viral Video Analyzer: space-video-topic

Before, analyzing viral videos had an extremely disgusting process: first find a mini-program to download without watermark, then import it into Feishu Miaoji to transcribe, and finally painstakingly analyze the structure yourself.

This Skill combines these steps into one.

Give it a Douyin, YouTube, or Bilibili link, and it will automatically:

  1. Download without watermark (for Douyin links, I wrote logic to simulate browser refresh signatures, calling yt-dlp which is extremely stable).
  2. Generate a verbatim transcript with timestamps.
  3. Analyze the topic structure, and incidentally help you reconstruct new titles and voiceover scripts.

I tested it with one of my previous videos, and the structure it output was very clear and directly usable.

2. De-AI Script: space-video-script

The biggest problem with AI-written scripts is that they don't "sound human". They often start with "In this ever-changing time and space...", which can suffocate you when read aloud.

The core logic of this Skill is to "speak human". It automatically reconstructs written language into spoken language suitable for on-camera expression. For example, it changes those grand parallel sentences into "Let's talk about some real talk today".

At the same time, it outputs a shot-by-shot script accurate to the second, telling you where to look and what visuals to show at each second.

3. Breathing Subtitles: space-video-subtitle

Many automatic subtitle tools rigidly break sentences based on punctuation marks, resulting in a line of text that is extremely long and tiring to look at.

This Skill uses "voiceover-style sentence breaking": it segments based on the speaker's breathing and semantic pauses. Moreover, it focuses on proofreading typos that drive perfectionists crazy, like the Chinese homophones "在/再" and "的/得/地", and finally outputs standard SRT/ASS formats.

Group 2: Social Anxiety Rescue (Master Control, Editing, Voiceover & Music)

The scariest part of recording videos is freezing up. Saying one word wrong in front of the camera means re-recording the entire segment. This group of Skills is specifically designed to cure various recording awkwardnesses.

4. Director Master Control: space-video

When you have many Skills, both you and the AI can get confused: which one should I use now?

So I configured a "Director" role. When you say to it "help me make a video", it won't write blindly, but first asks you: Voiceover or screencast? Horizontal or vertical? Any reference videos?

After clarifying the requirements, it will call the corresponding sub-Skills on its own; you don't need to memorize those complex commands.

5. Painless Voiceover Editing: space-video-edit

When recording voiceovers, we inevitably make mistakes like "He... he... hello everyone, uh, no, hello everyone".

This Skill's processing method is very clever: it first converts the video to text, identifying slips of the tongue, filler words, and silent segments. Then it adopts the principle of "delete the earlier, keep the later" — because when people make a mistake, they often repeat the complete sentence correctly afterward.

It will directly cut out the earlier mistakes and filler words, stitch the smooth clips together, and realign the subtitles along with it.

6. Smart Sound Engineer: space-video-audio

If you really don't want to reveal your voice, it will call the free edge-tts to convert text into narration.

The most practical part is its mixing logic: it will automatically search royalty-free music libraries (like Pixabay) for background music, and during synthesis, it will also do "ducking" processing — when the voice comes in, the background music automatically lowers; when the voice stops, the music comes back up.

Finally, it normalizes the loudness to -14 LUFS. The whole process does not re-encode the video, producing the final cut in seconds.

Group 3: Visual Outsourcing (Motion B-roll, Hand-drawn Whiteboard, Cover)

When editing videos yourself, the biggest headache is visual material. A face talking from start to finish, and the audience runs away early. Going to stock footage sites, even after paying for a membership, you can't find the right stuff.

7. Code Motion Director: space-video-broll

Since you can't find suitable material, just draw it live with code.

This Skill finds a visual metaphor based on your voiceover content (like using a river to represent traffic, or collapse to represent data simplification), then uses pure HTML/CSS to arrange a set of continuous motion graphics over 30 seconds, rendering them frame by frame into a deterministic MP4 video.

Although the style leans geeky, the advantage is absolute originality, with zero copyright risk.

8. Hand-drawn Whiteboard Diagram: space-video-broll-sketch

If you give it a public account article, it will break down the logic inside into 6 to 10 frames of hand-drawn whiteboard diagrams.

White background, black single lines, with a bright blue accent, it looks like an invisible hand is drawing stroke by stroke on the screen.

The backend connects to ByteDance's Volcano Engine Seedance and libtv's Agent channel, generating and stitching frame by frame, with a very premium effect.

9. Modular Cover: space-video-cover

The cover determines the click-through rate. This Skill takes a "templated" route: it fixes your main color, font, and layout, only changing the specific title and accent color for each issue.

This not only produces images quickly but also ensures a highly unified visual style for your homepage. It supports one-click output for various sizes like WeChat Channels, Bilibili, Xiaohongshu, etc.

How to Run Them?

If you have an Agent terminal like Qoder, Claude Code, or Codex at hand, just enter this command line to install with one click:

Help me install these Skills: github.com/SpaceZephyr/creator-buddy/tree/main/video-skills

After installation, the three pipelines I run most often are:

Except for video downloading and video model generation which require configuring corresponding Keys, most of the other Skills do not depend on third-party paid APIs, and daily free use is completely sufficient.

Honestly, don't just watch. The code is all on GitHub. Go ahead and set up the one part you usually find most painful (like editing voiceovers or finding material) and give it a try.