Nanny-level tutorial on digital live broadcast video production: from script to multilingual film

🛒 For newcomers to operations and content, we will help you complete one-stop production from topic selection to multi-lingual film production.

Tutorial Objectives

This tutorial takes HeyGen as an example to help you start from a script, complete a publishable digital voice broadcast video, and learn to batch produce multilingual versions. Just follow the steps throughout the process, no editing background required.

Preparation Checklist

  • [ ] Register a HeyGen account (or similar tools, such as Synthesia) and confirm the free quota.
  • [ ] Prepare a spoken script of 100-200 words (can be generated with AI assistance).
  • [ ] Determine the purpose and format of the video: the horizontal version (16:9) is suitable for Bilibili/YouTube, and the vertical version (9:16) is suitable for Douyin/Video Account/Kuaishou.
  • [ ] Prepare brand-related materials (Logo, background image, spoken subtitle copy).
  • [ ] Confirm the source of the image you use: the platform’s own image, or the sound/image material you have authorized.

Step 1: Write a spoken script

The oral broadcast script follows the structure of "hook - body - call to action":

[Hook] Many people think that making short videos is difficult, but in fact there is only one method.
[Text] Today, use 3 steps to turn a piece of copywriting into a spoken video...
[CTA] If you want to learn systematically, reply "tutorial" in the comment area and I will send you the complete list.

Writing points:

  • Each sentence should be limited to 20 words or less to facilitate oral broadcasting and subtitles.
  • Be colloquial and avoid written accents and long clauses.
  • Indicate where "key pauses" or "subtitle emphasis" are needed.

Oral broadcast script template example (can be applied directly)

[Hook] The biggest fear when creating content is not not being able to write it, but that no one will read it once you write it.
[Pain Point] Have you ever encountered: staying up late writing a script and recording it dozens of times, but the number of views is still in double digits?
[Method] Today I will share a method: first use AI to batch produce manuscripts, and then use digital humans to batch produce films. You can produce a week's worth of work in one day.
[Details] First, the topic selection should use "pain points + scenarios"; second, the script should be short and colloquial; third, use a unified image to continuously output and establish memory points.
【CTA】Want to get the complete script template? Reply "Digital Man" in the comment area and I will send you the template.

This structure is suitable for most oral broadcast scenarios such as product introduction, knowledge popularization, event promotion, etc., and can be used directly to replace the content.

Step 2: Generate digital human video

  1. Enter HeyGen’s “Create Video” and select “Avatar (Digital Human Image)”.
  2. Select horizontal or vertical format.
  3. Paste the script into the text box and set the language (default Chinese).
  4. Select image and voice: Give priority to the default image of "natural speech and clear mouth shape".
  5. Click Generate and wait for generation (usually a few minutes).

Check points: Whether the mouth shape is correct, whether the speaking speed is natural, and whether there is an obvious pause at the end.

Step 3: Adjust the voice and speaking speed

  • Speech speed: Default is 1.0, it can be adjusted to 1.1-1.2 for carrying goods/dry goods, but do not exceed 1.3, otherwise the mouth shape will be easily out of sync.
  • Pause: insert a comma/period into the script to break the sentence naturally; start a new sentence where you need to emphasize.
  • Change sounds: If you are not satisfied, regenerate after switching the sound list.

Step 4: Add subtitles and background

  1. Turn on "Automatic subtitles" and select the subtitle style and position.
  2. Add a background: You can use a solid color, a brand image, or a blurred background.
  3. Add brand logo and ending action guide (can be processed in editing software later).

Subtitle suggestion: Use highlighted colors for key numbers and CTA to increase the completion rate.

Step 5: Multilingual batch generation

  1. Copy the original video in the project and select "Regenerate".
  2. Switch the language to the target language (such as English, Japanese).
  3. Use platform translation, or paste the translated text after manual proofreading (the latter is recommended for higher quality).
  4. Export all language versions in batches.

Note: Be sure to check the pronunciation and terminology after translation to avoid "literal translation" that may damage the brand sense.

Step Six: Compliance and Platform Adaptation

Check before publishing:

  • [ ] Turn on the "AI generated content" annotation (as required by the platform).
  • [ ] The oral broadcast script does not contain absolute terms and illegal promises (Advertising Law).
  • [ ] The images and sounds used are from authorized sources.
  • [ ] Export according to the target platform specifications (Douyin vertical version 1080×1920, Site B horizontal version 1920×1080).

Step 7: Release and data review

  1. After publishing, focus on the completion rate, 3-second bounce rate and interaction rate.
  2. Compare the data of different images, speaking speeds, and covers.
  3. Write the data feedback back to the script library and iterate the next batch.

Post tempo suggestions

  • Starting period: 3-5 items per week, testing different hooks and images.
  • Stable period: Fixed 2-3 articles per week, focusing on outperforming topics.
  • Review table: record the video, playback, completion, interaction, and conversion of each video, and check the trend once every two weeks.

Material naming and archiving specifications

{Date}_{Series}_{Version}_{Language}.mp4
Example: 20260815_Product Introduction_v2_zh.mp4

Unified naming makes team collaboration and multilingual management more efficient and avoids the confusion of "final version v3 final version".

Quick check of channel specifications

Platform Recommended frame Recommended duration Remarks
Douyin/Video account 9:16 vertical version 30-60 seconds The first 3 seconds of the hook are the most important
Station B/YouTube 16:9 horizontal version 1-3 minutes High information density, subtitles can be added
Little Red Book 9:16 Vertical 40-90 seconds Cover and title both important

Before exporting, select the corresponding specifications according to the target platform to avoid image quality loss caused by secondary cropping.

Verification method

  • Verification 1: The generated video mouth shape and pronunciation are basically aligned.
  • Verification 2: There are no typos in the subtitles and the C-bit information is readable.
  • Verification 3: The multi-lingual version has been randomly checked and has no mistranslation.
  • Verification 4: The finished film has completed AI annotation and complies with platform rules.

Common mistakes and precautions

  1. The script is too long: a single script should be limited to 60-90 seconds of oral playback (approximately 250-350 words). If it is too long, it will easily end.
  2. Speech speed is too fast: above 1.3x speed, the mouth shape is easy to collapse. It is better to let the user read it by himself at 1.25x speed.
  3. Multilingual literal translation: Overseas versions must be proofread by native speakers to avoid rigid translation.
  4. Unlabeled AI content: Ignoring the labeling risks having your account banned, so be sure to check it when publishing.
  5. Material confusion: Create script, material, and finished film folders for each series, and name them with date and version.

Frequently Asked Questions and Troubleshooting (FAQ)

  1. The generation is too slow or queued?

    Queues are normal during peak periods and can be generated in batches in advance; you need to wait or upgrade after the free quota is used up.

  2. The mouth shape is not correct and the voice is awkward?

    Re-select a more natural image and voice, adjust the speaking speed back to 1.0, and break up the sentences into shorter sentences.

  3. The video has subtitles but the platform automatically adds them. Is this repeated?

    Turn off inline subtitles when exporting, or use the platform's "automatic subtitles" instead.

  4. The digital human image looks like a foreigner, and the Chinese mouth shape is weird?

    Priority is given to images optimized in Chinese; custom images require recording of Chinese material.

  5. Was it judged as being handled or of low quality by the platform?

Ensure original scripts, differentiated images (add real-shot materials/animated transitions), and standardize AI generation.

  1. How to test multiple images at low cost?

    Use the same script to generate 2-3 versions of the image respectively, and use a small budget to compare data before scaling up.

Summary and next steps

At this point, you have been able to independently produce a publishable digital voice broadcast video, and mastered the key points of multilingual mass production and compliance. It is recommended to verify 3-5 items in small batches first, and use real data (completion, interaction, conversion) to compare different images and script styles; after running out a stable model, you can then consider customizing digital humans or programmatic production lines.

Advancement and Expansion

  • Customized digital human: record exclusive image and voice to enhance brand recognition.
  • Live broadcast digital people: From recording and broadcasting to real-time interactive customer service and delivery.
  • Programmed production: use script library + scheduled generation + automatic release to form an assembly line.
  • Content matrix: multiple accounts, multiple languages, and multiple image combinations, and data feed back topic selection.

User Reviews

  • Loading reviews...