Your photos
4–8 clear photos showing your face and upper body.
A simple, step-by-step program
Build a digital presenter that looks and sounds like you. Then turn a written script into a talking-head or video-podcast video.
Before you start
You do not need special computer skills or expensive camera equipment.
4–8 clear photos showing your face and upper body.
A quiet place and a phone or microphone to record your voice.
Higgsfield, Pinterest, ElevenLabs and HeyGen.
One useful idea you would like to teach or explain.
Use only your own face and voice, or get written permission.
Never clone another person without permission. Never send your private connection key by email or message.
Make one simple choice
Use AI photos only
Record once, then reuse
Not sure? Choose Avatar IV and continue with this guide.
Create your visual identity
A character sheet is one picture that shows the same person from several angles. It helps the computer keep your face consistent.
Choose 4–8 clear photos of yourself.
Include front, left and right views of your face.
Use photos without sunglasses, hats, filters or other people.
Open Higgsfield and start a new image project.
Select Nano Banana Pro as the image model, then choose 4K quality.
Upload all your reference photos together.
Use your Custom GPT instructions, or copy the optional prompt below, then generate the picture.
Not needed when you use the Custom GPT provided earlier.
Use all uploaded reference photos as the identity reference for the same person. Create a photorealistic 4K horizontal character reference sheet in 16:9 format. Show the same person in six clean half-body portraits: 1. Front view with a neutral expression 2. Front view with a warm, friendly smile 3. Left three-quarter view 4. Right three-quarter view 5. A natural speaking expression 6. A calm, confident expression Keep the face, age, skin tone, hairstyle, body shape and plain dark-blue top exactly consistent in every portrait. Use soft, even studio lighting and a simple light-grey background. Natural skin texture. Clear eyes. Realistic hands. No text, no logos, no extra people, no duplicated features and no change of identity.
Generate it again. Do not use a poor character sheet.
Choose your setting
Find a room, studio or podcast setting you like. Use it for inspiration, then describe its useful parts instead of copying it exactly.
Open Pinterest and search for “vertical talking head studio” or “podcast portrait setup”.
Choose a clear half-body scene with simple lighting and background.
Save the picture to your device.
Upload it to an AI chat that can look at pictures.
Use your Custom GPT, or copy and paste the optional scene-analysis prompt below.
Save the scene recipe that the AI gives you.
Not needed when you use the Custom GPT provided earlier.
Study the uploaded scene reference and turn it into a simple scene recipe for a new, original picture. Do not describe or copy the person in the reference. Describe only: 1. Camera angle and how close the camera is 2. Where the presenter sits or stands 3. Lighting direction, softness and colour 4. Background, furniture and general objects 5. Main colours 6. How sharp or blurred the background is 7. Empty space that can be used for captions Do not copy logos, artwork, brand names or unusual decorations. End with one short scene recipe for a vertical 9:16 talking-head picture.
Do not copy the original person, logo, artwork or unique room exactly. Create your own version.
Create your vertical portrait
Now combine your character sheet with the scene idea. This creates the still picture that HeyGen will bring to life.
Open Higgsfield and start a new image project.
Select Nano Banana Pro as the image model, then choose 4K quality.
Upload your character sheet and the scene-reference picture.
Use your Custom GPT instructions, or paste the optional recreation prompt below.
Generate at least three versions so you can compare them.
Choose the version that looks most like you and has natural eyes and hands.
Download the chosen vertical picture to your device.
Not needed when you use the Custom GPT provided earlier.
Use the uploaded character sheet as the identity reference. Use the uploaded scene reference only for general composition and mood. Create a new, original, photorealistic talking-head scene in vertical 9:16 format at the highest available quality. Show the same person from the character sheet in a half-body portrait. Keep the face, age, skin tone, hairstyle and body shape consistent. The person looks directly at the camera with a warm, calm and confident expression. Recreate the general camera angle, lighting style, colour mood and room type from the scene reference, but do not copy its person, logos, artwork or unique decorations. Keep the face and both hands clear. Leave a little space above the head. Keep the lower part of the picture simple so captions will be easy to read. Natural skin. Realistic eyes and hands. No text, no watermark, no extra fingers, no distorted furniture and no change of identity.
Zoom in and check the eyes, teeth and fingers before downloading the picture.
Create your speaking voice
A voice clone is a computer-made copy of your voice. The quality depends mostly on how clean and natural your recording sounds.
Choose a quiet room. Turn off fans, music, television and phone alerts.
Place your phone or microphone about one hand-length from your mouth.
Record 60–120 minutes in the same room with the same microphone position.
Speak alone, using the friendly energy you normally use on camera.
Include short and long sentences, names, numbers, questions and natural emotion.
In ElevenLabs, open Voices and choose to add a Professional Voice Clone.
Upload your recordings and complete the permission check for your own voice.
After training, test the voice with several different sentences.
Strong noise removal, music or sound effects can make the final voice less natural.
Join the picture and voice
You will create a private connection key in ElevenLabs. Think of it as a special password that lets HeyGen use your voice.
In ElevenLabs, open Developers, then API Keys.
Create a new key named “HeyGen Connection”. Do not use your master key.
Allow: Text to Speech, Voices, Voice Generation, Models, History and User.
Set a monthly credit limit, create the key and copy it immediately.
In HeyGen, open AI Studio and choose or upload your avatar picture.
Open the Voice menu, choose New Voice, then Integrate 3rd Party Voice.
Choose ElevenLabs, paste the private key and select Confirm.
Select your cloned voice from the voice list.
Use Avatar IV for your AI-made still picture. Use Avatar V only if you completed its short setup video.
Check the listed permissions, copy the key again without spaces, and confirm you have credits. If needed, create a fresh key and delete the old one.
Prepare what your avatar will say
A good video script sounds like speech, not an article. Start with one short video and one useful message.
Decide the one topic you want to explain.
In the prompt below, replace [TYPE YOUR TOPIC HERE] with your topic.
Copy the full prompt and paste it into your preferred AI chat.
Read the finished script aloud. Shorten any sentence that feels difficult.
Check every name, number and factual claim before using the script.
Keep your first test video close to 60 seconds.
Not needed when you use the Custom GPT provided earlier.
Write a clear 60-second talking-head video script about [TYPE YOUR TOPIC HERE]. The audience is a complete beginner. Write like a warm, experienced person speaking to one friend. Use this order: 1. Start with one sentence that makes the viewer curious 2. Explain the main idea in plain language 3. Give three useful points or steps 4. End with one simple action the viewer can take Use short sentences and everyday words. Make it natural to say aloud. Do not use jargon, headings, bullet points, stage directions or exaggerated claims. Keep it between 120 and 145 words.
You are responsible for checking the facts before your digital twin says them publicly.
Make and review your video
Make a short test first. Fixing a 20-second sample is faster and cheaper than fixing a complete video.
In HeyGen AI Studio, start a vertical 9:16 video project.
Choose your Avatar IV picture, or your Avatar V digital twin.
Paste your checked script into the script box.
Choose the cloned ElevenLabs voice you connected in Step 5.
Use the simple motion prompt below if motion controls are available.
Generate a 20–30 second test and watch it from beginning to end.
Fix pronunciation, face, hand or timing problems before generating the full video.
Add readable captions, a clear title and only music you are allowed to use.
Download and watch the final file once more before publishing.
Not needed when you use the Custom GPT provided earlier.
Look directly at the camera. Speak in a warm, confident and friendly way. Use small natural nods and gentle open-hand gestures. Keep steady eye contact. Do not overact.
Not needed when you use the Custom GPT provided earlier.
Review this generated talking-head video carefully. Check these seven things: 1. The face stays consistent 2. The voice sounds clear and natural 3. The lips match the words 4. The eyes and blinking look natural 5. The hands and fingers look normal 6. The captions are correct and easy to read 7. There are no false claims, strange objects or sudden visual changes Give me a short list of problems. For each problem, tell me exactly what to fix. If there are no serious problems, say: Ready to publish.
If one part looks wrong, repair and regenerate that short scene instead of starting the whole video again.
You have finished the setup
For your next video, begin at Step 6: write a new script, paste it into HeyGen, check the result and publish.
Keep this guide open
Each link opens a new page, so you can return to this guide easily.