October 11, 2026 Saigon, Vietnam

Build Log

No. Explainer video Tiếng Việt
Ship it Log it
More writing
on the blog.
Visit mtri.me →
Tiếng Việt

Two character videos read the log

“The machine does most of it. The lines that do not sound like me, I rewrite.”

Log 001 was about building this page. In this issue, two characters read back that work log as five lessons.

The product is two clips, both in Vietnamese. Mai has a cute voice. Kuma has a bold, rousing voice. Below is how I made them.

Video transcript · 5:23

Hi, I'm Tri. This issue is a bit different. I'll talk about the product, and it's not a website. It's two short videos. First, a video about Mai, a girl I drew in the style of Mitsuru Adachi. Kuma is a 3D teddy bear. The two of them read back build log number one, and pull out some quick lessons in just over two minutes. Here's how I made those clips.

First, the script. Claude read about 25 hours of activity in the log, then went through 36 messages, and boiled it down to five lessons. In the first draft, her voice sounded a bit cutesy, and a lot of lines read like slogans. So I opened a page in the browser to review the wording. The page on the left, every line on the right, and I rewrote fifty lines in one go. Then I asked Claude to compare my version with its original, and pull out the rules. One: no attitude on the AI, and every recap line has to give the reason. After this round of scripting, Claude uses those rules for the next videos.

Second, the voice. I used the VieNeu model. VieNeu is a free Vietnamese voice engine, and it runs locally on a Mac. The first voice sounded a bit cutesy, so Claude measured how much each voice goes up and down. The old voice swung more than twelve semitones. The one called Thuc Doan stayed under ten, so I picked Thuc Doan. VieNeu has no speed control, so Kuma's voice was slowed down to about nine tenths, with ffmpeg. Machines often misread English names, so there's a separate respelling list. Not just English names, but other special words too. The captions keep the original spelling.

I also tried out Chatterbox. I used it for English text to speech. When I made the English version of issue one, my Soniox account ran out of money. And the free plan of ElevenLabs had no voice cloning. So I used Chatterbox. It's open source, MIT licensed, and it runs locally on a Mac. It clones a voice from just twelve seconds of audio, then reads each line to match when I speak. My disk had no room for the original model, so I had to use one half the size. The clone worked, but in the end, the English version used my own reading. Then the machine cut each line and put it in the right place.

Last, how I made the characters. I used Codex to draw one base image, then drew it thirteen more times, each with a different face. AI never draws the same thing twice, so every image is off by a little, in a lot of places. Claude had to write code to line up each image with the base one, and only cut out the part of the face that changed. At first, the cut lost the lines on the closed eyes, and left ghost marks in the image. Then I worked to fix it so the seams don't show.

Next, the characters' mouths. In Vietnamese every word is one syllable, so the mouth has to keep opening. I used Whisper to listen to the voice track and find out how the mouth should move, and at which second. Then I added animation, with the eyes blinking every few seconds, in a fixed sequence, so every render comes out the same. On emotional lines, the character changes face or bounces, or gets manga marks like an exclamation point or a sweat drop. In the end it all comes together in one HTML page running GSAP, and HyperFrames renders it as a vertical video, the kind you see on Instagram or Facebook.

The part I spent the most care on was reviewing every word. The machine writes it all, then I go through it. Any line that doesn't sound like how I talk I rewrite. That's it. See you next issue.

Mai, a cute female voice (Vietnamese) · 2:22
Kuma, a lively bear voice (Vietnamese) · 2:11

Mai, a cute female voice

A girl drawn in the style of Mitsuru Adachi. Voice: Thục Đoan, slow speed. Length: 2 minutes 22 seconds. The face changes 33 times.

Kuma, a lively bear voice

A 3D teddy bear with glasses. Voice: Thanh Bình, speed 0.9×. Length: 2 minutes 11 seconds. The same five lessons, told in the first person.

Small details

Five steps. Each step took trial and error.

Three of Claude's lines next to my rewrites, with the count of rewritten lines and rules

A script with a human editor

Claude read back the log of 25 hours of work (I took a while) and 36 messages. Claude pulled out five lessons. The first draft sounded a bit hollow. Many lines read like slogans. I reviewed the draft on a copy-reviewGlossaryA copy review page I built for my agent. The real page is on the left and each line is on the right. The editor rewrites or approves each line, then the agent writes the edits into the code.Full glossary → page. The page is on the left, and each line is on the right. I rewrote 50 lines. Claude compared my version with its version and wrote 19 writing rules. Example: do not give the AI an attitude. A recap line must have a reason.

shipped byTri
Chart of the pitch range of 5 VieNeu female voices: Mỹ Duyên 12.3, Thục Đoan 9.8 semitones

A voice that runs on the Mac

The voices use VieNeuGlossaryA Vietnamese text-to-speech (TTS) model. It is free, runs locally, and has preset voices such as Thục Đoan and Thanh Bình.Full glossary →. VieNeu is free and runs on a Mac. Claude measured the pitch range of each voice. The first voice moved 12.3 semitones and sounded cutesy. The Thục Đoan voice moved 9.8 semitones. VieNeu has no speed control. ffmpegGlossaryAn open-source command-line tool that cuts, joins, converts and changes the speed of audio and video.Full glossary → slows Kuma's voice to 0.9×. Machines read English names incorrectly, so I added a respelling table that spells them the Vietnamese way (example: MIT becomes "em ai ti"). The captions keep the original word (MIT).

shipped byTri
Waveform of the 12-second voice sample and the three steps of the English version with Chatterbox

Chatterbox, for the English issue 001

Issue 001 has an English version. SonioxGlossaryA pay-as-you-go speech AI service: speech to text and translation.Full glossary → ran out of money halfway. The free plan of ElevenLabsGlossaryA speech AI service for text to speech and voice cloning. The free plan has no voice cloning.Full glossary → does not clone voices. ChatterboxGlossaryAn open-source TTS model from Resemble AI, MIT licensed. It clones a voice from a few seconds of audio and runs locally.Full glossary → has an MIT license and runs on the Mac GPU. Chatterbox cloned my voice from a 12-second sample. The machine generated each caption line one time. Then it stretched each line 0.85 to 1.15× to match when I speak. My disk did not have space for the original model, so I used the fp16 version, half the size. For the final version I read the English myself, so it sounds natural. The machine cut each line and put it in the correct position.

shipped byTri
Mai's 13 faces laid over the base image

One drawing, thirteen faces

CodexGlossaryOpenAI's coding agent that runs in the terminal. I use it to generate images through a ChatGPT plan.Full glossary → drew one base image and 13 face images. AI does not draw the same image twice, so each image is off by a few pixels. A program aligns each image to the base image. The program keeps only the face area that changed. In the first pass, a filter removed the eyelash lines on closed eyes. The filter also left ghost marks. Claude removed the filter and repaired the joins. A join passes only when no seam shows.

shipped byTri
Five frames of Mai: mouth open on four words, closed in the pause

A mouth that opens on every word

In Vietnamese, each word is one syllable. The mouth must open on each word. WhisperGlossaryOpenAI's open-source speech recognition model. It turns speech into text, with the time of each word.Full glossary → gives the start time of each word. The mouth stays open for 62% of the word duration. The eyes blink on a fixed sequence. Each render gives the same result. On emotional lines, the character changes face or bounces. The character also shows manga marks: an exclamation point, a sweat drop, stars.

shipped byTri
Frames of the Mai and Kuma vertical clips next to GSAP code and the HyperFrames render command

Put together in HTML

Each clip is one HTML page. The GSAP skillGlossaryGreenSock Animation Platform, a JavaScript library for web animation. The GSAP skill is a set of instructions that helps an agent write GSAP animations.Full glossary → runs the animation. HyperFramesGlossaryA command-line tool that renders an animated HTML page into a video file, one frame at a time.Full glossary → renders the page as a 1080×1920 vertical video. Instagram and Facebook use this format.

shipped byTri
Vietnamese and English waveforms for the first 40 seconds, a dashed line at each caption line

This video, in English

The video in the corner has an English version. Chatterbox learned my voice from an 11-second talking part of this same take. The machine generated each caption line one time and put it where I spoke it. Sadly, the lips do not match the English words.

shipped byTri

Set up the environment

Copy this and paste it into your agent (Claude Code, Codex). The agent installs and checks each step.

Goal: set up a local environment to make Vietnamese explainer videos with an animated character. Target: macOS or Linux.

Rules:
- Do one step at a time. Verify each step before you start the next.

Steps:
1. Check for `uv`, `ffmpeg`, and Node 18 or later. If one is missing, install it.
2. Create a Python 3.12 environment: `uv venv ~/.agentvid/vieneu-env --python 3.12`.
3. Install `vieneu` and `faster-whisper` in that environment.
4. Read one Vietnamese sentence with the presets "Thục Đoan" and "Thanh Bình". Save each voice as a WAV file. The first run downloads a model of approximately 600 MB.
5. Use `faster-whisper` (model small, language vi) to get the time of each word. Write JSON as [{w, s, e}].
6. English voice (if necessary): create a separate Python 3.11 environment and install `chatterbox-tts`. On a Mac, use the GPU (`mps`) and set `PYTORCH_ENABLE_MPS_FALLBACK=1`.
7. I will give you an English voice sample of 10 to 15 seconds. Use it to clone the voice, with the default settings. Read one test sentence and save it as a WAV file.
8. Make a 1080x1920 HTML page with a GSAP timeline and the WAV file. Render 5 seconds with `npx hyperframes render`.
9. Write a Python script (numpy, Pillow). Align each expression image to the base image. Keep only the face pixels that changed. Write PNG files with a transparent background.
10. Mouth: open it on each word, for approximately 60% of the word duration.
11. Blink every 2.6 to 5 seconds. Use a fixed seed.

Script rules:
- Send English names through a pronunciation table (example: read Claude as "Cờ lót"). Captions keep the original word.
- One sentence, one idea. Do not add facts that are not in the log.

Report: what you installed, where the test files are, and which steps failed.

Onelastthing

Kuma read the word build wrong. I tried the respelling "biu", but it came out as "bui". So the respelling changed to "biêu". Both clips had to render again.

on the mind ofTri

Past issues

View all issues +
▶5:23