元峰AI笔记 · AI绘画与视频

How to prompt Veo 3 for the best results

🔗 原文链接: https://replicate.com/blog/usingan...

How to prompt Veo 3 for the best results

🔗 原文链接: https://replicate.com/blog/using-an...

⏰ 剪存时间:2025-06-18 16:59:01 (UTC+8)

✂️ 本文档由 飞书剪存 一键生成

Google’s Veo 3 generates videos with audio from text prompts. The audio can be dialogue, voice-overs, sound effects and music.
谷歌的 Veo 3 能根据文本提示生成带音频的视频。音频可以是对话、旁白、音效和音乐。

Let our resident AI podcaster introduce us:
让我们的驻场 AI 播客主持人来为我们介绍:

🎦 点击观看视频

Prompt: A podcast show, a woman in a grey sweater and dark brown tousled hair in an updo, she looks directly at the camera, with strands framing her face. She talks into a mic saying: This is Replicate’s guide to prompting Veo 3…
提示:一个播客节目,一位穿着灰色毛衣、深棕色蓬松盘发的女性,她直视镜头,发丝环绕脸庞。她对着麦克风说:这是 Replicate 的 Veo 3 提示指南……

Write what happens 写下发生的事情

First the basics. A well-crafted prompt is the key to generating good videos. The more you can specify in your prompt, in plain language, the easier it is for Veo 3 to understand and generate the video you want.
首先是最基本的。一个精心制作的提示是生成优质视频的关键。您在提示中用普通语言描述得越详细,Veo 3 就越容易理解和生成您想要的视频。

Try to include these visual elements in your prompt:
尝试在您的提示中包含这些视觉元素:

  1. Subject: Who or what is in the scene — a person, animal, object, or landscape.
    主题:场景中是谁或什么——一个人、动物、物体还是风景。
  2. Context: Where is the subject? Indoors? A city street? A forest?
    背景:主题在哪里?在室内?在市中心街道?在森林里?
  3. Action: Is your subject walking, jumping, turning their head?
    动作:你的主题是在行走、跳跃、转头吗?
  4. Style: The visual aesthetic you’re aiming for (cinematic, animated, stop-motion, etc).
    风格:您追求的视觉美学(电影感、动画、定格动画等)。
  5. Camera motion: Describe how the camera moves: aerial shot, eye-level, top-down, or low-angle.
    摄像机运动:描述摄像机如何移动:空中镜头、平视镜头、俯视镜头或低角度镜头。
  6. Composition: How the shot is framed: wide shot, close-up, etc.
    构图:镜头的取景方式:远景、特写等。
  7. Ambiance: Mood and lighting. You can say things like “warm tones,” “blue light,” or “nighttime.”
    氛围:情绪和光线。你可以这样说,比如“暖色调”、“蓝光”或“夜晚”。

You also need to include audio elements, which we’ll cover in more detail below.
您还需要包含音频元素,我们将在下文中更详细地介绍。

Here’s an example of a basic prompt versus a detailed one:
这里是一个基本提示与详细提示的示例:

A man answers a rotary phone
一个男人在接转盘电话

Versus: 对比:

A shaky dolly zoom goes from a far away blur to a close-up cinematic shot of a desperate man in a weathered green trench coat as he picks up a rotary phone mounted on a gritty brick wall, bathed in the eerie glow of a green neon sign. The zoom reveals the tension and the desperation etched on his face as he struggles to talk on the phone. The shallow depth of field focuses on his furrowed brow and the black rotary phone, blurring the background into a sea of neon colors and indistinct shadows, creating a sense of urgency and isolation.
一个摇晃的推镜头从远处的模糊画面逐渐拉近,展现出一个穿着破旧绿色风衣的绝望男子,他正从布满灰尘的砖墙上拿起一台老式转盘电话,周围笼罩着绿色霓虹灯诡异的微光。镜头拉近后,他脸上的紧张和绝望清晰可见,他艰难地对着电话说话。浅景深聚焦在他紧锁的眉头和黑色转盘电话上,背景被模糊成一片霓虹色彩和模糊的阴影,营造出一种紧迫感和孤立感。

The second prompt contains structural elements to nudge Veo 3 towards the scene we are trying to create.
第二个提示包含结构元素,引导 Veo 3 朝我们想要创造的场景发展。

无法获取视频链接

Basic prompt 基础提示

无法获取视频链接

Detailed prompt 详细提示

Change your prompt each time

每次更改你的提示

If you’re familiar with prompting models like Midjourney or Flux , you’ll know that with those models you’ll get a decent level of variation if you run the same prompt a few times (ie using different seeds).
如果你熟悉像 Midjourney 或 Flux 这样的提示模型,你会知道,使用这些模型时,如果你多次运行相同的提示(即使用不同的种子),你会得到相当程度的多样性。

Veo 3 is different. For the same prompt, even one that’s fairly simple, Veo 3 will output very similar results. You might get the same looking person in the same clothes, in a similar sort of place. This is excellent if you’ve had an output that has a slight error, like a coherency or audio glitch – you can run a different seed and get what you want. But if you’re in discovery mode, when you want to see a range of what’s possible, then running the same prompt multiple times is a waste of your money.
Veo 3 则不同。对于相同的提示,即使是相当简单的提示,Veo 3 输出的结果也非常相似。你可能会得到穿着相同衣服、出现在相似地点的相同样貌的人。如果你已经得到了一个有轻微错误的输出,比如连贯性或音频故障——你可以运行不同的种子来得到你想要的结果。但如果你处于探索模式,当你想要看到各种可能的结果时,多次运行相同的提示就是浪费你的钱。

In the example below we ran the prompt “a woman laughs” twice, with different seeds. Note how she looks the same, she’s wearing the same clothes, she laughs in the same way, the room is the same, she’s even wearing the same earrings. It’s unusual for a model to be this consistent.
在下面的示例中,我们两次运行了提示“一个女人在笑”,使用了不同的种子。注意她看起来一样,穿着相同的衣服,笑的方式相同,房间也相同,甚至戴着的耳环也相同。模型如此一致是不寻常的。

If you’re not sure what you want yet, start with a few broadly different prompts. If you know elements of what you want, then be specific about those.
如果你还不确定想要什么,可以从几个广泛不同的提示开始。如果你知道你想要的内容元素,那就针对这些元素进行具体描述。

In this video the obvious things we could do are begin to play with descriptions for:
在这个视频中,我们可以开始尝试调整描述,包括:

  • how the woman looks (hair color, hair style, skin color)
    女性的外貌(发色、发型、肤色)
  • what she’s wearing 她所穿的衣服
  • where she is 她所在的位置
  • how she is laughing
    她怎么笑的
  • why she is laughing
    她为什么笑

Here are a couple of examples:
这里有两个例子:

🎦 点击观看视频

a woman laughs long and loudly, she’s in an office meeting and she’s embarrassed afterwards
一个女人的笑声又长又响,她在办公室开会,事后感到尴尬

🎦 点击观看视频

a woman laughs quietly, she’s at home watching a tv show
一个女人的笑声很轻,她在家里看电视节目

Character consistency 角色一致性

Typically character consistency is hard when you’re using a video model without a starting frame or scene ingredients. These features are coming to Veo 3 soon.
通常情况下,当你使用没有起始帧或场景元素的视频模型时,保持角色一致性很难。这些功能即将在 Veo 3 中推出。

In the meantime, because similar prompts yield similar characters, if you keep a character’s detailed prompt description consistent across generations, you’ll often get someone who looks the same. This means you can keep a list of character descriptions and repeat them verbatim across different prompts:
与此同时,由于相似的提示会产生相似的字符,如果你在多代生成中保持某个角色的详细提示描述一致,你通常会得到一个看起来相同的人。这意味着你可以保留一份角色描述列表,并在不同的提示中逐字重复它们:

John, a man in his 40s with short brown hair, wearing a blue jacket and glasses, looking thoughtful
约翰,一个 40 多岁的男人,留着棕色短发,穿着蓝色夹克和眼镜,看起来若有所思

The more unique and specific these descriptions, the better Veo 3 maintains visual continuity between separately generated scenes. Create character reference sheets with exact wording to ensure consistency.
这些描述越独特和具体,Veo 3 在分别生成的场景之间保持视觉连续性的效果就越好。创建带有精确措辞的角色参考表以确保一致性。

🎦 点击观看视频

John, a man in his 40s with short brown hair, wearing a blue jacket and glasses, looking thoughtful, he says: Hello, I am also John, and I look kind of the same as that guy over there (no subtitles!). He is in a bright light room.
约翰,一个 40 多岁的男人,留着棕色短发,穿着蓝色夹克和眼镜,看起来若有所思,他说:你好,我也是约翰,我看起来和那边那个人有点像(没有字幕!)。他身处一个明亮的房间里。

🎦 点击观看视频

John, a man in his 40s with short brown hair, wearing a blue jacket and glasses, looking thoughtful, he says: Hello, my name is John, I am a character invented for this blog post (no subtitles!)
约翰,一个四十多岁的男人,留着棕色短发,穿着蓝色夹克和眼镜,看起来若有所思,他说:你好,我叫约翰,我是为这篇博客文章虚构的人物(没有字幕!)

Prompting audio 提示音频

As Veo 3 generates audio with each video, you also need to prompt for the audio you want to hear. Consider these elements:
随着 Veo 3 为每个视频生成音频,你也需要提示你想要听到的音频。考虑以下要素:

  1. What people are saying (dialogue)
    人们在说什么(对话)
  2. The ambient noise of the scene (the sounds of a busy street, a busy office, a busy cafe, etc)
    场景的环境噪音(繁忙街道、忙碌办公室、忙碌咖啡馆等的声音)
  3. Sound effects or noises from outside the scene (like a phone ringing)
    场景外部的音效或噪音(如电话铃声)
  4. Any music the scene might need (a tense cinematic score, a cheerful pop song, etc).
    场景可能需要的任何音乐(紧张的影视配乐、欢快的流行歌曲等)。

Prompting dialogue and avoiding subtitles

提示对话并避免字幕

The characters you can create with Veo 3 are fascinating. They talk, tell jokes, gesticulate, sometimes they can act. But if you want them to talk, you need to prompt for that.
你可以用 Veo 3 创造出非常有趣的角色。它们会说话、讲笑话、做手势,有时甚至能表演。但如果你想让它们说话,你需要为此进行提示。

You can prompt dialogue in two different ways:
你可以用两种不同的方式进行提示对话:

  1. Explicitly: “A guy says: My name is Ben”
    明确地:“一个家伙说:我叫本”
  2. Implicitly: “A guy tells us his name”
    隐含地:“一个家伙告诉我们他的名字”

Both of these will lead to a video of a guy talking, the first will use the exact words you asked for, the second will let the model decide how to say it, in this case the model will decide on a name for you.
这两种方式都会生成一个男士说话的视频,第一种会使用你要求的确切词语,第二种会让模型决定如何表达,在这种情况下,模型会为你决定一个名字。

Writing your own dialogue

撰写自己的对话

If you’re being explicit about what’s being said, try to keep your dialogue short. It should be something that can be said in just about 8 seconds.
如果你明确说明对话内容,尽量保持对话简短。它应该能在大约 8 秒内说完。

If you try to pack too much in, then you can end up with a character that’s speaking way too fast. If you ask them to say too little, you can get either awkward silences or a character saying nonsensical AI gibberish (like the second example below). Without clear guidance the model won’t be able to make up all the words it needs to.
如果你试图塞入太多内容,结果可能会让角色说话速度过快。如果你要求他们说得太少,可能会出现尴尬的沉默,或者角色说出无意义的 AI 胡言乱语(如下面的第二个例子)。如果没有明确的指导,模型将无法编造出它所需要的所有词语。

🎦 点击观看视频

John, a man in his 40s with short brown hair, wearing a blue jacket and glasses, looking thoughtful, he says: You have given me a really long prompt, and I have to speak very quickly and unnaturally to try and fit all these words into just 8 seconds, I’m going to be out of breath at the end of this, phew.
约翰,一个 40 多岁、留着棕色短发、穿着蓝色夹克和眼镜、看起来若有所思的男人说:你给了我一个非常长的提示,我必须很快且不自然地说,才能把所有这些话都塞进 8 秒钟,我会在说完后喘不过气来,呼。

🎦 点击观看视频

Too short (and with AI gibberish): John, a man in his 40s with short brown hair, wearing a blue jacket and glasses, looking thoughtful, he says: Hello, I’m John.
太短(而且有 AI 胡言乱语):约翰,一个 40 多岁、留着棕色短发、穿着蓝色夹克和眼镜、看起来若有所思的男人说:你好,我是约翰。

Letting Veo 3 script the dialogue

让 Veo 3 编写对话

If you aren’t good at writing dialogue, implicit dialogue prompts will help. And you can always transcribe the outputs you liked for use in later prompts.
如果你不擅长写对话,隐式对话提示会很有帮助。而且你可以随时将你喜欢的输出转录下来,用于后续的提示。

Here we ask Veo 3 to create a video of a standup comic telling a joke, first we let Veo 3 decide on the joke. Second video we get Veo 3 to try and deliver the joke we put in the prompt.
在这里我们要求 Veo 3 制作一个脱口秀喜剧演员讲笑话的视频,首先我们让 Veo 3 决定笑话内容。第二个视频我们让 Veo 3 尝试讲我们在提示中给出的笑话。

🎦 点击观看视频

a standup comic tells an awkward joke at a music festival, sounds of distant bands, noisy crowd, ambient background of a busy festival field (no studio audience)
一个脱口秀演员在音乐节上讲了一个尴尬的笑话,远处乐队的声音,嘈杂的人群,繁忙音乐节场地的环境背景(没有现场观众)

🎦 点击观看视频

a standup comic tells an awkward joke at a music festival: You know what’s great about music festivals? Watching 20,000 people pretend they knew this band before today while filming vertical videos they’ll never watch.
一个脱口秀演员在音乐节上讲了一个尴尬的笑话:你知道音乐节有什么好吗?看着 2 万人假装他们今天之前就认识这个乐队,而他们拍的视频永远不会看。

As you can see, given the right prompt and all the appropriate context, Veo 3 can fill in the dialogue for you.
如您所见,只要给出正确的提示和所有适当的背景信息,Veo 3 就能为您填充对话内容。

Some prompts you could try to see how versatile Veo 3 is with dialogue:
您可以尝试一些提示,看看 Veo 3 在对话方面的多功能性如何:

  • a standup comic tells a joke
    一个脱口秀演员讲笑话
  • two people discuss a movie
    两个人讨论一部电影
  • a man is having an argument over the phone
    一个男人在电话里争吵
  • a woman tells us her life story
    一个女人告诉我们她的生活故事

Getting pronunciation right

发音正确

Sometimes you’ll find that the model is pronouncing words incorrectly. The easiest way to handle this is to spell the words phonetically. In the opening example, our podcaster says:
有时你会发现模型发音不正确。处理这个问题的最简单方法是使用音标拼写单词。在开头的示例中,我们的播客主持人说:

Read on to get fofr and Shridar’s guidance on making videos
继续阅读,获取 fofr 和 Shridar 关于制作视频的指导

But to get the correct pronunciation for our names, we had to change the prompt to:
但要获得我们名字的正确发音,我们必须将提示改为:

Read on to get foh-fur’s and Shreedar’s guidance on making videos
继续阅读以获取 foh-fur 和 Shreedar 关于制作视频的建议

Who says what 谁说了什么

When you prompt a conversation between multiple characters you will sometimes find that Veo 3 mixes up who says what. This is common when the characters have similar descriptions, and it’s ambiguous to Veo 3 which character is which.
当你提示多个角色之间的对话时,有时会发现 Veo 3 会搞混谁说了什么。当角色描述相似时,这很常见,对 Veo 3 来说难以区分哪个角色是谁。

Try to be specific in your prompt about who is speaking:
在提示中尽量具体说明谁在说话:

The woman wearing pink says: But I’m the one who’s wearing pink
穿粉色衣服的女人说:但我才是穿粉色衣服的那个人

The man with the glasses replies: No, I’m the one with the glasses
戴眼镜的男人回答:不,我才是戴眼镜的那个人

Avoiding subtitles in outputs

避免输出字幕

Veo 3 must have been trained on plenty of videos with baked-in subtitles, because it’s very common to see poorly spelled and incorrect subtitles in the outputs. They often ruin a generation, but there are a couple of easy ways to avoid them:
Veo 3 一定是在大量带有内置字幕的视频上训练的,因为输出中经常能看到拼写糟糕和错误的字幕。它们常常毁掉一次生成,但有几个简单的方法可以避免这种情况:

  • put the speech you want to hear after a colon, like: “A guy says: My name is Ben” rather than in quotes, like: “A guy says: ‘My name is Ben”
    在冒号后面输入你想要听到的语音,例如:“一个家伙说:我的名字是本”,而不是用引号,例如:“一个家伙说:‘我的名字是本’”
  • put “(no subtitles)” in the prompt, negatives work well in Veo 3 prompts
    在提示中输入“(无字幕)”,否定词在 Veo 3 的提示中效果很好
  • if all else fails, keep saying No subtitles. No subtitles! Multiple times.
    如果其他方法都无效,就反复说“无字幕。无字幕!”多次

The wrong background audio (or the case of the unwanted live studio audience)

错误的背景音频(或是不想要的现场观众)

If you don’t define the background audio you want to hear in your video, then Veo 3 needs to work it out, often that’s ok, but sometimes it gets it wrong. A live studio audience is a common hallucination. Sometimes it’s what you want, like a fake sitcom. But usually the extra laughter doesn’t fit the scene. Veo 3 even did this when making the examples above, here’s an example of an out of place studio audience ruining a generation:
如果你没有定义视频中想要听到的背景音频,那么 Veo 3 就需要自行判断,通常这样是可以的,但有时它会出错。现场观众是一个常见的幻觉。有时这正是你想要的,比如一个虚假的情景喜剧。但通常额外的笑声与场景不符。Veo 3 在制作上述示例时也这样做了,这里有一个不合时宜的现场观众破坏一代人的例子:

🎦 点击观看视频

Example of unwanted studio audience laughter in the background.
背景中不想要的现场观众笑声示例。
Prompt: “a standup comic tells an awkward joke at a music festival”
提示:“一个脱口秀喜剧演员在音乐节上讲了一个尴尬的笑话”

The easiest way to avoid this is to prompt the audio you expect to hear explicitly. In this case we fixed the generation by adding “sounds of distant bands, noisy crowd, ambient background of a busy festival field” to get the right sort of feeling in the output.
避免这种情况最简单的方法是明确提示你期望听到的音频。在这种情况下,我们通过添加“远处乐队的声响、嘈杂的人群、繁忙音乐节场地的环境背景”来调整生成效果,从而在输出中获得正确的氛围。

Prompting music 提示音乐

Just like the rest of the video, if you want music in your scene you need to include it in your prompt.
就像视频的其他部分一样,如果你想在场景中添加音乐,你需要在提示中包含它。

Again, you can be explicit and describe the genre, style and mood of the music you want to hear. Or you can be more vague and let Veo 3 decide.
再次,你可以明确描述你想要听的音乐类型、风格和情绪。或者你也可以更模糊一些,让 Veo 3 来决定。

Styles 风格

Out of the box Veo 3 will typically generate something that looks like a well-produced live action video, think smooth professional demo, a commercial or a music video.
默认情况下,Veo 3 通常会生成类似高质量实拍视频的内容,比如流畅的专业演示、商业广告或音乐视频。

If you want to steer it away from this you need to include a style with your prompt. Here are some examples of styles Veo 3 knows how to generate, the prompt is:
如果你希望引导它偏离这种风格,你需要在提示中包含风格描述。以下是 Veo 3 能够生成的风格示例,提示如下:

In the style of [style name]: A bearded man in a flannel shirt and weathered jeans sits cross-legged beside a flickering campfire, its amber light casting soft, dancing shadows across the pine-needle-strewn ground of a quiet forest clearing. Across from him, just beyond the edge of the firelight, stands a massive grizzly bear, calm and still, its fur catching the warm glow, eyes reflecting the flames with eerie intelligence. The two shake hands, like they’re old friends.
以[风格名称]风格:一位留着胡须、身穿法兰绒衬衫和磨损牛仔裤的男子盘腿坐在摇曳的篝火旁,篝火的琥珀色光芒在松针铺满的宁静森林空地上投下柔和、舞动的阴影。在他对面,就在火光边缘之外,站着一头巨大的灰熊,它平静而静止,毛发沐浴在温暖的光芒中,眼睛映照着火焰,透出诡异的神智。他们握手,就像老朋友一样。

You’ll notice that not only does the look of the video change, but also the way the characters move and interact too.
你会注意到,不仅视频的外观发生了变化,角色的动作和互动方式也改变了。

In each of these, the audio remains very similar, we haven’t prompted the audio differently, and it hasn’t changed much between the different styles.
在这些视频中,音频保持非常相似,我们没有以不同的方式提示音频,而且它在不同的风格之间变化不大。

无法获取视频链接

Original video 原始视频

无法获取视频链接

LEGO 乐高

无法获取视频链接

Claymation 粘土动画

无法获取视频链接

South Park 南方公园

无法获取视频链接

Pixar animation 皮克斯动画

无法获取视频链接

8-bit retro 8 位复古

无法获取视频链接

Graphic novel 漫画

无法获取视频链接

Origami 折纸

无法获取视频链接

Simpsons 辛普森一家

无法获取视频链接

Blueprint 蓝图

无法获取视频链接

Anime 动漫

无法获取视频链接

Marble 大理石

Camera motion 相机动作

As you might expect, just like other video models, Veo 3 responds well to common camera movement prompts. Using terms like these, you can control the action in your video:
正如你所预期的那样,和其他视频模型一样,Veo 3 对常见的摄像机运动提示反应良好。使用这些术语,你可以控制视频中的动作:

  • eye level 视线高度
  • high angle 高角度
  • worms eye 虫眼
  • dolly shot 推镜头
  • zoom shot 变焦镜头
  • pan shot 全景镜头
  • tracking shot 跟拍镜头

无法获取视频链接

Zoom in 放大

无法获取视频链接

Zoom out 缩小

无法获取视频链接

Left to right pan 从左到右的平移

无法获取视频链接

Dolly shot 推镜头

Selfie-style videos 自拍风格的视频

Veo 3 is surprisingly good at making selfie videos that actually look real. We’ve found that certain phrases seem to unlock this behavior consistently.
Veo 3 在制作看起来真实的自拍视频中表现惊人。我们发现某些短语能始终如一地解锁这种效果。

Starting with “A selfie video of…” works much better than just describing a person with a camera.
以“一个自拍视频……”开头效果远好于仅仅描述一个手持相机的人。

Making the arm visible is key for authenticity. The gorilla example shows this well with “holds the camera at arm’s length. His long, powerful arm is clearly visible in the frame.” That’s what makes it look like an actual selfie rather than a close-up shot.
让手臂可见是关键,真实性体现在这一点上。大猩猩的例子很好地展示了这一点,描述为“将相机拿在手臂长度处。他强壮的长臂在画面中清晰可见。”正是这一点让它看起来像真实的自拍,而不是特写镜头。

Natural eye movement also helps a lot. The Tokyo example shows this with “occasionally looking into the camera before turning to point at interesting stalls.” That kind of natural glancing behavior works better than staring directly at the camera.
自然眨眼动作也非常有帮助。东京的例子通过“在转向指向有趣的摊位之前偶尔看向镜头”展示了这一点。这种自然的扫视行为比直视镜头效果更好。

Here are two examples that show how this works:
这里有两个例子展示了这是如何运作的:

A selfie video of a travel blogger exploring a bustling Tokyo street market. She’s wearing a vintage denim jacket and has excitement in her eyes. The afternoon sun creates beautiful shadows between the vendor stalls. She’s sampling different street foods while talking, occasionally looking into the camera before turning to point at interesting stalls. The image is slightly grainy, looks very film-like. She speaks in a British accent and says: “Okay, you have to try this place when you visit Tokyo. The takoyaki here is absolutely incredible, and the vendor just told me it’s been in his family for three generations.” She ends with a thumbs up.
一位旅行博主探索东京繁忙街头市场的自拍视频。她穿着复古牛仔外套,眼中充满兴奋。午后的阳光在摊位之间投下美丽的阴影。她边品尝不同的街头小吃边交谈,在转向指向有趣的摊位之前偶尔看向镜头。画面略带颗粒感,看起来非常像电影。她用英国口音说:“好吧,你们来东京一定要尝试这个地方。这里的章鱼烧绝对不可思议,而且摊主告诉我这已经传到他家族三代了。”她以竖起大拇指结束。

A handheld selfie-style shot, from the point-of-view of a gorilla in a lush jungle. A large silverback gorilla holds the camera at arm’s length. His long, powerful arm is clearly visible in the frame, and his face is perfectly framed. The gorilla says: “I’m just testing out this actually works and I’m going to post it on TikTok later, Essentially it felt cute might delete it later” (lips moving like he’s saying it)
手持自拍风格的镜头,从茂密丛林中一只大猩猩的视角拍摄。一只大型银背大猩猩将相机举到手臂长度。他的长而有力的手臂在画面中清晰可见,脸部也完美地框在镜头中。大猩猩说:“我只是在测试这个是否真的有效,稍后我要把它发到 TikTok 上,基本上感觉很可爱,可能以后会删掉”(嘴唇像在说话一样动)

One more thing the Tokyo example shows: adding “The image is slightly grainy, looks very film-like” seems to push the output away from that too-clean AI look. It ends up feeling more like something that was actually shot on a phone.
东京示例还展示了另一件事:添加“图像略带颗粒感,看起来非常像电影”似乎使输出远离那种过于干净的 AI 外观。最终感觉更像是用手机实际拍摄的内容。

How to make vertical videos with Veo 3

如何使用 Veo 3 制作竖屏视频

At the moment Veo 3 doesn’t natively support vertical videos, only 16:9 horizontal are possible. You can however take a landscape video and outpaint it using a model like Luma’s Reframe Video .
目前 Veo 3 原生不支持竖屏视频,仅支持 16:9 横屏视频。不过你可以使用像 Luma 的 Reframe Video 这样的模型将横屏视频转换为竖屏视频。

Reframe video lets you pass in any video (up to 30 seconds long), and outpaints it as a new video at the specified aspect ratio. All outputs will be 720p.
Reframe Video 允许你输入任何视频(最长 30 秒),并将其以指定的宽高比输出为新的视频。所有输出视频均为 720p。

🎦 点击观看视频

A Veo 3 video that was reframed as a 9:16 vertical video
一个被转换为 9:16 竖屏视频的 Veo 3 视频

Native support for vertical videos in Veo 3 is coming soon.
Veo 3 将很快原生支持竖屏视频。

Physics 物理

Veo 3 excels at simulating realistic physics, maintaining proper motion and interactions while applying different styles. The model preserves the natural movement of objects, ensuring that physics-based animations like falling, bouncing, and fluid motion remain physically accurate even when transformed into different artistic styles.
Veo 3 在模拟逼真物理方面表现出色,能够在应用不同风格的同时保持正确的运动和交互。该模型保留了物体的自然运动,确保基于物理的动画(如下落、弹跳和流体运动)即使在转化为不同的艺术风格时也能保持物理上的准确性。

无法获取视频链接

LEGO

无法获取视频链接

Origami

无法获取视频链接

Chrome

无法获取视频链接

Paint

Upscaling to 4k and 60fps

将分辨率提升至 4k 和 60fps

By default, Veo 3 outputs 1280p x 720p video. We recommend using Topaz Lab’s Video Upscaler to bring your videos up to 4k resolution and 60 frames-per-second.
Veo 3 默认输出 1280p x 720p 视频。我们建议使用 Topaz Lab 的视频提升器将您的视频提升至 4k 分辨率和 60 帧每秒。

Final remarks 最终总结

The difference between a bland video and a great one comes down to your prompt. With Veo 3, you’re not just describing what happens, you’re directing a scene. High quality videos will layer in subject, setting, action, camera work, audio and mood. Think like a filmmaker and Veo 3 will follow your lead.
一部平淡无奇的视频与一部精彩视频的区别在于你的提示。使用 Veo 3,你不仅仅是描述发生了什么,而是在指导一个场景。高质量的视频会叠加主题、场景、动作、摄影、音频和情绪。像电影制作人一样思考,Veo 3 就会跟随你的领导。

🎦 点击观看视频

One final prompt: 最后一个提示:

A podcast show, a woman in a grey sweater and dark brown tousled hair in an updo, with strands framing her face. She is in a room with pink and gold uplighting. No subtitles. She is giving an outro and looks directly at the camera as she talks into a mic saying (no subtitles!): So that’s the end of our guide, we hope you found it useful. Feel free to try Veo 3 on Replicate , and don’t forget to follow us on X .
一个播客节目,一位穿着灰色毛衣、深棕色蓬松盘发的女性,发丝环绕她的脸庞。她身处一个带有粉色和金色顶光的房间。没有字幕。她正在做结束语,并直视镜头,对着麦克风说(没有字幕!):这就是我们的指南的结尾,希望对您有帮助。欢迎在 Replicate 上尝试 Veo 3,别忘了在 X 上关注我们。