I started making two kinds of product videos

Introduction
I wanted one product update to produce two useful videos:
- a wide video that explains the product and the evidence behind it;
- a short vertical video that starts with one problem and gets to the point quickly.
My first attempt was too generic. It used generated slides and a synthetic voice, but it did not show a real product workflow. The narration also sounded worse than the newer voices I was already using for Trend Seeker.
The better approach is to reuse the research, transcript, product footage, and presenter. The final edits can still be different. I show the complete setup in this 10-minute YouTube walkthrough.
The two formats have different jobs
The wide version has room for source evidence, caveats, and a complete product flow. The short version needs one question, one or two supporting signals, and one next step.
| Format | What I use it for | What it needs |
|---|---|---|
| Wide video | A product walkthrough or a complete argument | Chapters, real interface footage, source context, and time to inspect the result |
| Vertical video | One problem or product idea for TikTok, Shorts, or Reels | A direct opening, large captions, a readable crop, and one clear close |
The distinction is editorial, not just a crop. A three-minute explanation does not become a good short when I cut off the sides. I rewrite the short around the strongest supported point.
I start with one source package
Both videos begin with the same small package:
- the product page or feature I want to explain;
- the public sources behind the claim;
- a transcript split into scenes;
- the interface states, screenshots, or charts that belong to each scene;
- the limitations that must stay in the script.
For the Trend Seeker example below, the opening is based on app reviews about failed clock-outs and missing punches. The video says that one reviewer reported losing half an hour of wages. It does not turn one report into a failure rate or proof of demand.
That source package is the useful reusable part. It lets me shorten an explanation without changing what the evidence says.
The presenter is a reusable component
I originally made the product walkthrough without a presenter. It worked, but it felt like an ordinary generated screencast. I wanted a narrator that could add body language while the real product remained on screen.
I built a small illustrated pose library. The current Tõnis presenter has a neutral state and five gestures: offer, point, present, question, and count. Each gesture has a midway pose. The animation moves from neutral to midway to the held gesture, then returns along the same path.
The presenter also has three head positions, speaking and blink states, breathing, small weight shifts, and hair movement. A cue track chooses gestures that match the transcript. A line with two parts can use the count gesture. A question can use a small shrug.
This is illustrated 2.5D animation. It is not a photorealistic talking head. That makes the generated nature clear and fits the paper style of this site.
The same system can produce a vertical UGC-style cut
The short version uses a different composition. The presenter starts close to the camera, moves to the side, and leaves room for the source cards. Captions are burned in because many people will first see the video without sound.
This Trend Seeker cut asks one question: can someone make a better timekeeping app? It then shows the reported wage loss, a narrow product concept, the available search context, and a small employer test. The video is 29 seconds rather than a compressed version of the full three-minute explanation.
The short keeps only five scenes: the question, the complaints, the product concept, the search context, and the test. That structure is easier to follow than a general list of avatar features.
The render pipeline is mostly scripts
The workflow uses a few replaceable parts:
- I write or review the transcript and mark the purpose of each scene.
- Google Chirp 3 HD generates separate voice clips so timing changes do not require rebuilding the entire narration.
- A script maps parts of the transcript to gestures and measures the real audio duration.
- Blender renders the illustrated presenter, including the pose changes, head movement, blinks, and audio-driven mouth movement.
- FFmpeg combines the presenter, product footage, cards, captions, voice, and optional music.
- The same source package can be composed as 16:9 or 9:16 without regenerating the character.
Codex helped generate the character states, write the Blender and composition scripts, and inspect the exported frames. The workflow still needs review. Generated pose images can change a face, sleeve, hand, or neck between frames. The script cannot decide that an almost-correct hand is acceptable.
The mouth movement is still basic
The mouth opens according to audio volume. It does not select mouth shapes for individual speech sounds. At picture-in-picture size this is usable for my current experiments. A close-up makes the limitation obvious.
Character consistency is the harder part. Every new pose needs the same face, clothes, proportions, light, and crop. I keep one approved neutral image as the reference, generate the pose family from it, then compare the images before animation.
The first wide and vertical exports also exposed smaller production problems: weak narration, captions with uneven vertical padding, source cards that were too dense, and presenter art that became soft when enlarged. These are ordinary video bugs. I now inspect the opening, every transition, the complete spoken ending, and representative frames at final size.
I avoided a separate avatar subscription, not all costs
Blender and FFmpeg are free. I already pay for Codex, so I used that existing subscription while building the artwork and scripts. Voice generation is metered separately. When I recorded the walkthrough, Google listed one million free Chirp 3 HD characters per month and then $30 per million characters; the current Text-to-Speech pricing is the source to check before using it.
The tradeoff is maintenance. A hosted avatar service gives you a finished interface and a larger character library. My version gives me the source files, pose library, render scripts, and the ability to make presenters that match each product.
One workflow, two deliberate edits
I am no longer trying to make one video fit every destination. I keep the evidence and production assets shared, then make two edits with different jobs.
The wide video can explain the setup and show the complete reasoning. The vertical version can make one problem understandable in under a minute. The presenter, voice clips, gesture cues, and render scripts are reusable, while the claim in each video still has to come from the source material.
The full YouTube walkthrough shows the avatar library, Blender animation, voice generation, FFmpeg composition, costs, and current limitations.