Kling AI video creation in-depth solution

🛒 Kling's AI in-depth application solution for video creators and content teams covers core functions such as Wensheng video, Tusheng video, video extension, motion brush, and picture consistency control, combined with Kling's advantages of high image quality and strong motion consistency.

Kling AI video creation in-depth solution

Solution overview

With Kling AI as the core, this solution is aimed at short video creators, advertising production teams, film and television storyboard artists, and e-commerce content operators to build an end-to-end AI video creation workflow from prompt word design, screen generation, motion control to post-delivery. The solution makes full use of Kling's differentiated capabilities in Wensheng Video, Tusheng Video, Motion Brush, Lip Sync, and multi-lens narrative, and cooperates with tools such as Midjourney, CapCut, and HeyGen to form complementary collaboration links.

Unsolved problem: This solution does not involve 3D modeling and rendering, real-shot material collection, professional-grade color grading and mixing, nor does it replace the fine editing process of Premiere Pro / DaVinci Resolve. The solution focuses on the video generation stage and the back-and-forth linkage that Kling AI can complete independently.

Target Users:

  • Individual Creator: Xiaohongshu/Douyin/Bilibili/YouTube independent video creator, hoping to use AI to accelerate content production
  • Advertising/Marketing Team: Planning and execution staff of brand TVC, information flow advertising, and social media short videos
  • Film and television pre-production team: storyboard artist and director of short plays/micro-movies, using AI for shot preview and script visualization
  • E-commerce content operation: content producers for product display, evaluation videos, and live broadcast slicing

Prerequisites:

  • Have registered a Kling AI account (domestic version "Keling AI" or overseas version KlingAI) to understand the points/subscription mechanism
  • Have basic prompt word writing skills and understand the basic terminology of lens language
  • Prepare reference pictures or video materials to be processed (if you need to control pictures or videos)
  • Clarify video delivery specifications (duration, resolution, style, usage platform)

Toolchain list

Tools Purpose Required Account Level Estimated Fees Alternatives
Kling AI Core video generation engine Free version/Gold member ¥66/month Billed by points consumption
Midjourney High-quality reference images/keyframe generation Basic version $10/month Monthly subscription DALL·E / FLUX
CapCut AI editing, subtitles, dubbing, packaging Free version/Pro version Basic free Cutting / Premiere Pro
HeyGen Digital human broadcasting and multilingual dubbing Free version/Creator $24/month Pay-as-you-go billing D-ID / Synthesia
Runway Video polishing and green screen/erasure Free quota/Pro version $15/month Pay-as-you-go billing CapCut
Pika Supplementary video generation and style transfer Free version/Pro version Pay-as-you-go billing Kling can also complete it itself

Tool collaboration logic: Midjourney makes high-quality reference frames (key frame seeds) → Kling makes core text/image videos + motion control → CapCut makes editing packaging and subtitles → HeyGen makes digital voiceover dubbing → Runway makes green screen erasure and minor modifications. The strengths of each tool are placed in the most suitable process.

Preparation

Before officially launching the plan, please confirm the following preparations item by item:

Account and environment

  • [ ] Register a Kling AI account (domestic https://klingai.com/ or overseas klingai.com) and become familiar with the Creative Studio interface
  • [ ] Confirm the subscription plan: Free file for personal testing, platinum membership or enterprise API is recommended for high-frequency output
  • [ ] Prepare reference image materials, the recommended resolution is 1024×1024 or above, 16:9 or 9:16 is preferred
  • [ ] Prepare prompt word template library: collect 20-50 high-conversion prompt words and create templates according to style/scenario classification

Delivery specifications confirmed

  • [ ] Determine the video output duration: Kling single segment 5s/10s, multi-segment splicing to achieve longer duration
  • [ ] Determine resolution: 720p (fast generation) or 1080p (HD delivery)
  • [ ] Determine the voice/dubbing method: pure subtitles, AI dubbing or HeyGen digital spoken voice

Team division of labor (if applicable)

  • [ ] One person is responsible for prompt word design and screen review
  • [ ] One person is responsible for editing, packaging and subtitles
  • [ ] Clear acceptance criteria: each video must undergo at least one manual review before being released

Step-by-step guide

Step 1: Requirements dismantling and prompt word engineering

⏱ Estimated time: 0.5-1 day (first project), shortened to 30 minutes after reuse 🎯 Goal: Convert video requirements into executable prompt word solutions ⚠️ Prerequisites: Video delivery specifications confirmed

Operation instructions

This is the starting point of the entire creative chain and the step that most determines the quality of the output. Kling's compliance with prompt words has been greatly improved in the 3.0 series, but prompt words still need to be organized in a structured way to obtain stable output.

Specific operations

  1. Disassemble the video elements: Extract from the requirements document - scene (indoor/outdoor/city/nature), subject (character/animal/object/scene), action (running/flying/rotating/deformation), lens (push/pull/shake/move/follow/surround), atmosphere (tone/light/season/mood)
  2. Hierarchical writing prompts: Write according to the six-paragraph structure of "subject + action + environment + lens + style + image quality", for example:

    A young woman in a flowing red dress walks through a rain-soaked Tokyo alley at night, neon reflections on wet pavement, cinematic dolly follow shot, warm amber lighting mixed with cool blue tones, 4K ultra HD, film grain

  3. Negative prompt word configuration: Explicitly excluded content - deformation, multiple fingers, flickering, screen tearing, etc.
  4. Create a prompt word library: Create an exclusive prompt word template library for each project to facilitate batch reuse and iteration

Expert View: The "lens control word" of the prompt word is Kling's differentiated strength compared to Runway and Pika. It is recommended to explicitly add lens movement commands such as dolly zoom, tracking shot, and aerial slow descent in the prompt words. Kling's Camera Control capability can restore these lens languages ​​relatively stably. If the prompt words are well written, the subsequent post-production workload can be reduced by more than 50%.

Verification method

  • [ ] Prompt words are previewed in the Creative Studio preview window
  • [ ] Prepare at least 3 alternative prompts with different styles/angles
  • [ ] Negative prompt words have been fully configured

Step 2: Reference image production and key frame generation (image pre-production video)

⏱ Estimated time: 1-2 hours 🎯 Goal: Generate high-quality reference frames as seed material for Tusheng videos ⚠️ Precondition: The prompt word template in step 1 has been completed

Operation instructions

Kling’s Tusheng video quality has long been in the first echelon. Therefore, the quality of the reference image directly determines the upper limit of the final video. Never just randomly find a low-resolution image from the Internet as input.

Specific operations

  1. Use Midjourney to generate key frames: Based on the visual scheme refined in step 1, generate high-definition reference frames in Midjourney. Recommended Midjourney 6.0 / 6.1, use --ar 16:9 --style raw --v 6.1 parameters to get a more cinematic output
  2. Reference image preprocessing:
    • Crop to target aspect ratio (16:9 for landscape, 9:16 for portrait)
    • Denoising + Sharpening (can be enhanced with CapCut or Photoshop’s AI)
    • Avoid areas with excessive detail or cluttered textures (Kling tends to flicker in these areas)
  3. Multi-Image Reference: For scenes that require high character consistency, prepare 2-3 reference pictures of the same character from different angles. Kling 3.0’s Multi-Image Reference can extract character features from multiple images.
  4. Synthetic reference picture backup directory: named according to scene/shot number to facilitate subsequent batch operations

Expert View: Tusheng videos are Kling’s absolute strength. Compared with purely literary videos, the output of graphic videos is much more predictable. If this is your first time using Kling, it is recommended to start with Tusheng Video to build up your confidence, and then transition to Vincent Video. After Midjourney produces a picture, you can directly drag it into Kling's picture-making video panel. This combination of "Midjourney as a seed + Kling as a driver for motion" is currently one of the links with the highest output quality in AI video creation work.

Verification method

  • [ ] The resolution of the reference image is no less than 1024×1024, and the composition is clear
  • [ ] Multi-angle reference pictures have been prepared for character consistency scenes
  • [ ] Reference images are archived and named according to project/shot specifications

Step 3: Core video generation (Wensheng video + Tusheng video)

⏱ Estimated time : 2-4 hours (including multiple rounds of trial and error) 🎯 Goal: Batch produce AI video clips that meet quality requirements ⚠️ Prerequisites: Prompt words and reference pictures are ready

Operation instructions

This is the core implementation link of the plan. Kling Creative Studio provides two types of entrances: Wensheng Video and Tusheng Video, with functions such as Motion Brush, Camera Control, and Lip Sync for fine control.

Specific operations

  1. Select generation mode:
    • Pure copywriting driver: Directly paste the prompt word in step 1 and select "Wensheng Video"
    • Picture driver (recommended): Upload the reference picture in step 2, select "Picture Video", and set the exercise intensity (1-10, the larger the value, the greater the range of motion)
    • Multi-Image Reference: When character/scene consistency is required, turn on Multi-Image Reference and upload 2-3 reference pictures.
  2. Parameter configuration:
    • Duration: 5s (quick production) or 10s (more complete narrative)
    • Gears: Standard (quick preview), Professional (high-quality output)
    • Resolution: 720p (Draft Review) or 1080p (Final Delivery)
    • Lens direction: Customizable push/pull/pan/shift parameters
  3. Motion Brush Fine Control:
    • "Brush" the desired motion area on the static reference image
    • Set the movement direction (left and right/up and down/rotate/zoom)
    • Suitable for: characters' hair flowing, water ripples, leaves swaying, clouds flowing
  4. Lip Sync dubbing and lip-syncing (if needed):
    • Upload portraits of people
    • Provide spoken audio or copy text
    • Kling automatically generates lip sync videos
    • Applicable to: product introduction, oral broadcast, virtual anchor scene
  5. Multiple rounds of iteration:
    • Record the parameter combination (cue word + exercise intensity + gear + seed value) after each generation
    • Adjust prompt words or exercise intensity for unsatisfactory clips and regenerate them
    • It is recommended to generate 3-5 variants of the same prompt word for later screening

Expert View: Motion Brush is an important tool that differentiates Kling from competing products. After being introduced in version 1.6, it allows AI video generation to evolve from "hitting luck" to "hitting where you point it." Typical usage is: first use Tusheng Video to make the background motionless or slightly moving, and then use Motion Brush to precisely control the movement direction of the foreground subject. Compared with Runway's Motion Brush, Kling's brush selection range is larger and the boundaries are more natural, which is especially suitable for natural landscapes and portrait scenes.

About Points Management: Daily points for free files are limited. It is recommended to confirm all prompt words and parameters and then generate them in batches to avoid wasting points through repeated trial and error. The production cost of a single segment for platinum members is about ¥3-5 (reference value), and the material cost of a 30-second film (6 segments × 5s) is about ¥18-30.

Verification method

  • [ ] No obvious tearing, flickering or distortion in each video
  • [ ] Action continuity meets delivery requirements
  • [ ] Lip Sync Lip and audio synchronization rate ≥ 90%
  • [ ] Parameter combinations have been recorded and can be reproduced

Step 4: Picture consistency control and multi-camera narrative

⏱ Estimated time: 1-2 hours 🎯 Goal: Ensure that characters, scenes, and styles remain consistent across multiple videos ⚠️ Precondition: The core fragment of step three has been generated

Operation instructions

This is the most overlooked aspect of AI video creation. No matter how high the quality of a single video is, if the character's appearance, scene tone, and lens style are not consistent between the two segments, the audience will immediately become upset.

Specific operations

  1. Fixed character reference picture: All shots involving the same character use the same character portrait as the input of Tusheng video
  2. Multi-Image Reference Lock Character: In Kling 3.0, upload the character's front face picture + full body picture + scene reference picture at the same time to lock the character's identity
  3. Unified scene style parameters: All related lenses use the same set of style parameters (Motion Strength, gear, aspect ratio)
  4. Shot Planning List: Before actual generation, first make a shot planning list (Shot List), marking the reference image, prompt word core word, duration, and movement direction of each segment.
Shot number Type Reference image Prompt word core Movement direction Duration
Shot-01 Panoramic explanation City night view reference picture Aerial photography, light flow Slow descent 5s
Shot-02 Mid-shot character Character's front view The heroine walks and looks back Left → right follow-up shot 5s
Shot-03 Close-up Character close-up Eyes, breeze blowing hair Nudge + Motion Brush hair 5s
Shot-04 Cut scene Empty shot Empty shot of street, raindrops Panning from left to right 5s
  1. Color temperature consistency: If you need to switch between different scenes, uniformly mark the color temperature tendency (such as warm golden hour tone or cool blue mood) in the prompt word to reduce the pressure of post-production color adjustment.

Expert view: Picture consistency is the threshold for AI video generation to move from "showing off skills" to "industrialization". Although Kling 3.0's Multi-Image Reference and character consistency capabilities are not 100% perfect, they are enough to support scenes such as brand stories and product videos within 30 seconds. If your project is longer than 60 seconds or involves multi-character dialogue, it is recommended to include a manual consistency check node after each shot is generated.

Verification method

  • [ ] The facial features of the same character in all shots are consistent (identifiable as the same person/object)
  • [ ] Scene tone/style transitions naturally between shots
  • [ ] The shot planning form has been completely filled in and the generation parameters have been recorded

Step 5: Post-editing, dubbing and packaging

⏱ Estimated time: 2-3 hours 🎯 Goal: Edit the clips produced by Kling into films, complete dubbing, subtitles and packaging ⚠️ Prerequisite: The clips of steps three and four have been reviewed

Operation instructions

The videos output by Kling are high-quality footage, not final films. Post-production tools are required to complete editing and splicing, dubbing recording, subtitle generation and packaging animation.

Specific operations

  1. Use CapCut for rough cutting and splicing:
    • Import the clips output by Kling in the order of the shot planning table CapCut
    • Adjust the duration and connection transition of each segment (recommend cross-dissolve/flash-to-white transition to hide small differences in AI generation)
    • Add background music (CapCut's built-in copyrighted music library or upload your own audio)
  2. AI subtitle generation:
    • One-click speech recognition using CapCut’s smart subtitles feature
    • Manually proofread proper nouns and sentence fragments
    • Set subtitle style (font/size/position/stroke)
  3. Dubbing plan selection:
    • Pure subtitles: no dubbing required, suitable for purely visual storytelling
    • AI Speech Synthesis: CapCut built-in TTS or use HeyGen's AI voice cloning to choose the tone that suits your style
    • Digital Voice Broadcast: If you need a character to commentate on the scene, use HeyGen to generate a digital voice broadcast clip and intersperse it with the Kling screen
    • Lip Sync dubbing: If there is a close-up of the character's face in the shot, it is suitable to use Kling's own Lip Sync capability
  4. Packaging animation:
    • Add title/subtitle/title and ending
    • Fine-tuning of color grading (CapCut grading panel or LUT)
    • Preview the entire film rhythm before outputting

Expert View: CapCut and Kling work very closely together. CapCut’s “image-to-text” function can directly import Kling videos as material. It is also recommended to make "micro corrections" to the AI-generated clips in CapCut - Kling's clips occasionally have tiny flickering frames, which can be solved by manually cutting off 1-2 frames on the timeline without having to regenerate the entire clip.

Verification method

  • [ ] The whole film is smooth and has no lag, and the transition is natural.
  • [ ] Subtitles and voice are synchronized, no typos
  • [ ] The packaging format and resolution meet the delivery requirements
  • [ ] Preview the entire film before exporting

Step 6: Quality Review and Delivery

⏱ Estimated time: 1 hour 🎯 Goal: Ensure that the quality of the final film reaches the standard and complete the delivery ⚠️ Prerequisite: The draft version of step five has been exported

Operation instructions

Quality review is the last gate for the program and the last line of defense to prevent the "weird effects" generated by AI from leaking to the audience.

Specific operations

  1. Frame-by-frame review: Check the key clips frame by frame to confirm that there are no following points:
    • Screen flickers or jitters (AI generated FAQ)
    • Sudden deformation of character's face
    • Object edges flicker/pixelate
    • Unreasonable action logic (such as objects suddenly disappearing/appearing)
  2. Audio Review:
    • Balanced background music volume and dubbing volume
    • The dubbing has a natural intonation (AI voices need to pay special attention to sentence segmentation)
    • No popping or ambient noise
  3. Compliance Check:
    • If it involves brand/character portraits, confirm that authorization has been obtained
    • Confirm that the video content does not violate the platform content policy
    • Confirm that the copyright of the materials (pictures/music) used is clear
  4. Multiple format export:
    • HD main version (1080p H.264)
    • Platform adapted version (Douyin 9:16, Bilibili 16:9, YouTube 16:9)
    • Material backup version (retain CapCut project files to facilitate subsequent modifications)

Verification method

  • [ ] The audit list has been confirmed item by item, and there are no remaining issues.
  • [ ] Complete deliverables (finished film + project files + material backup)
  • [ ] Customer/Team Confirm Acceptance

Step 7: Effect review and template precipitation

⏱ Estimated time: 0.5-1 hour 🎯 Goal: Precipitate effective prompt words and workflow parameters for this project to improve the efficiency of the next time ⚠️ Prerequisite: Step 6 has been delivered

Operation instructions

A core advantage of AI video creation is "reproducibility-iteration". Spending 30 minutes doing a review after each project can significantly reduce the marginal cost of the next project.

Specific operations

  1. Record effective prompt words: Add the prompt words with the highest quality in this project to the team prompt vocabulary library, and mark the specific parameters (model version, exercise intensity, gear, seed value)
  2. Parameters for failed analysis: Which combinations of prompt words produced scrapped films? Record it to avoid repeated traps
  3. Update reference library: Save verified high-quality reference pictures and Kling-matched pictures into the material library
  4. Update lens planning template: Optimize the column fields of the lens planning table based on this experience
  5. Quantified output efficiency: Compare the difference in working hours between pure manual production and AI-assisted production

Expert View: A team with more than 50 verified prompt words and more than 20 sets of stable parameter templates is 5-10 times more efficient in AI video creation than a "from scratch" team. Cue word management is the most underestimated lever in industrializing AI video.

Verification method

  • [ ] Prompt dictionary has been updated with at least 3 valid entries
  • [ ] The project review document has been archived
  • [ ] 1-2 areas identified that can be optimized in the next project

Expected results

Indicators Pure traditional method AI-assisted method (this plan)
Single 30-second short video production cycle 3-5 days (including shooting + post-production) 0.5-1 day (including AI generation + review)
Material production cost (30-second film) ¥5,000-20,000 (shooting + actors + post-production) ¥30-200 (AI subscription + points consumption)
Multi-version production capability Single version 3-5 style variations of the same script
Personnel skill requirements Photographer + editor + voice actor 1 person only needs to master the prompt words + editing
Repeatedly modify costs Reshoot or edit Modify prompt words and regenerate

Acceptance criteria

  • [ ] From demand confirmation to final delivery, a single 30-second video takes ≤ 1 working day
  • [ ] The picture has no obvious AI defects (flickering/distortion/tearing)
  • [ ] Customer or target platform review passed once
  • [ ] All generation parameters and prompt words have been archived and can be reproduced

Frequently Asked Questions and Troubleshooting

Q: Is Kling’s free quota enough? How to avoid wasting points? A: The free daily points can generate approximately 10-20 5-second basic videos, which are suitable for experience and testing. For high-frequency output, it is recommended to upgrade to platinum membership (¥266/month). The key to saving points: first use the prompt words and parameters to confirm the effect in the preview mode, and then use the official generation to consume points after confirmation.

Q: The video generated by Kling occasionally flickers. How to solve it? A: Flicker is commonly caused by the following three situations: (1) The exercise intensity is too high (it is recommended to gradually increase it from 3-5); (2) The texture of the reference picture is too complex (change to a simpler background); (3) The prompt word lacks words related to picture stability (try to add stable footage, no flicker). In addition, you can manually cut off 1-2 frames of the flash frame in CapCut.

Q: How to deal with the inconsistent appearance of characters in multiple videos? A: Make sure that all relevant shots use the same character's front face picture as the input reference picture of the Tusheng video; enable Multi-Image Reference in Kling 3.0 to upload multi-angle character pictures; fix the character description (such as "the same young woman with long black hair and blue eyes") in the prompt word to avoid using different descriptions in different shots.

Q: Does Kling support commercial use? How is copyright defined? A: Videos generated by platinum and above packages are supported for commercial use, but must comply with Kuaishou/Kling AI's "Service Agreement" and "Generated Content Usage Specifications". It is recommended not to include third-party copyrighted elements (such as Disney characters, well-known brand logos), and to keep generation records for audit purposes.

Q: Can Kling be used instead of real shots? A: Scene-based: Scenes such as product displays, concept demonstrations, short videos, and advertising storyboards can completely replace real shooting; however, scenes that require real people to appear, real-life interactions, and camera coordination down to the second level, AI video is still not suitable to completely replace real shooting. It is recommended to adopt a hybrid mode of "AI generation as the main source + real shooting/real scene materials as the supplement".

Q: Which one is better, Kling, Runway or Pika? A: Everyone has their own strengths. Kling leads the way in understanding Tusheng videos and Chinese prompt words, and is the most friendly to Chinese creators; Runway has more comprehensive functions in video refinement (green screen, erasure, expansion); Pika is distinctive in stylized output. This solution uses Kling as the core and uses Runway and Pika as supplementary tools.

Advantages and Disadvantages of the Solution

Advantages

  • Chinese Native Friendly: High-quality output can be obtained by using prompt words in Chinese, with the lowest threshold for domestic creators
  • Top Video Quality: In terms of dynamic generation driven by static images, Kling has long been in the first echelon in the world.
  • Rich controls: Motion Brush, Camera Control, Lip Sync, Multi-Image Reference form a complete control toolbox
  • Outstanding value for money: The domestic subscription price is much lower than similar tools priced in US dollars, and the free quota can also meet light usage
  • Kuaishou ecological linkage: can be connected to Kuaishou main website, massive engine advertising system, and short distribution links

Not enough

  • Single segment duration limit: The default is up to 10 seconds. Multi-shot narratives require post-production splicing, which increases workload.
  • Facial long-term consistency still has flaws: In multi-shot projects longer than 30 seconds, character faces occasionally drift
  • Limited understanding of complex scenes: There is still a failure rate in scenes such as multi-person interaction, complex props, and fine hand movements.
  • API Commercial Threshold: Enterprise APIs require price negotiation and are not transparent enough for small and medium-sized teams.
  • Does not support external third-party model plug-ins: Not interoperable with ComfyUI / Stable Diffusion ecosystem

Tool summary

Tools Roles in this solution Priority usage scenarios
Kling AI Core video generation engine Vincent Video, Tusheng Video, Motion Brush, Lip Sync, multi-lens narrative
Midjourney Reference images/keyframe generation High-quality seed images, character design images, scene concept images
CapCut Post-editing and packaging Rough cutting, subtitles, dubbing, background music, color correction, export
HeyGen Digital voice broadcasting and multilingual dubbing Product explanation, multilingual localization, virtual anchor
Runway Video refinement and special effects Green screen snapping, video erasure, video expansion, super resolution
Pika Supplementary generation Style migration, special effects generation, special camera movement supplementation

Best practices and implementation suggestions

  1. Start with small projects: For the first project, choose a simple scene of 15-30 seconds (such as product demonstration, concept preview), and then expand to complex projects after running through the entire process.
  2. Establish a prompt word library: Classify and archive the high-converting prompt words in each project by scene. This is the team's most valuable AI video asset.
  3. Fixed parameter specifications: Unify the naming specifications for exercise intensity, gear, and resolution within the team to reduce communication costs.
  4. Manual review cannot be skipped: AI-generated videos must undergo manual review before delivery, focusing on checking faces, hands, and edge flickers.
  5. Follow Kling version updates: Kling has a fast iteration speed (only 15 months from 1.0 to 3.0), and each new version may bring qualitative changes. Stay tuned to the official Release Notes
  6. Mixed material strategy: Pure AI videos tend to have a "plastic feel". Mixing in 10-20% of real-shot materials or real-life materials can significantly improve the sense of reality.

User Reviews

  • Loading reviews...