Google DeepMind 推出 Gemini Omni 1.1 Flash,支持场景延展、关键帧控制和 4K 放大
要点速览
- •Gemini Omni 1.1 Flash 可将视频场景按 10 秒增量延展,累计最长 40 秒,并可参考最多 10 秒前文上下文,相比早期模型仅参考最后 1 秒有明显提升。
- •开发者可以指定首帧和末帧,模型会在两者之间生成连续视频,以支持镜头环绕、变焦转场和无缝循环。
- •360p 草稿预览的生成速度最高快 60%,成本仅为标准 720p 的三分之一,基于 Google 的系统吞吐量对比;最终项目可渲染为 1080p 或 4K。
- •Adobe 已将 Gemini Omni Flash 集成进 Adobe Firefly,Figma Weave、GMI Cloud 和 Runway 等合作伙伴也表示已在生产工作流中使用该模型。
- •开发者可通过 Google AI Studio 和 Gemini Enterprise Agent Platform 访问该模型,全球 Google AI Plus、Pro 和 Ultra 订阅用户也正在 Google Flow 中陆续获得使用权限,场景延展功能则已在 Gemini 应用中开放。

Google DeepMind 推出 Gemini Omni 1.1 Flash,支持场景延展、关键帧控制和 4K 放大
Google DeepMind 是 Google 面向 Gemini 模型家族的 AI 研究部门,现已推出 Gemini Omni 1.1 Flash,这是一项面向生产环境的更新,可为开发者提供更强的生成式视频控制能力。此次发布于 2026 年 8 月 27 日由 Google DeepMind 产品经理 Anish Nangia 和 Alisa Fortin 宣布,带来影院级视频制作能力,包括延展场景、首帧与末帧插值、清晰的 4K 放大以及更快的原型开发。开发者今天即可通过 Google AI Studio 或 Gemini Enterprise Agent Platform 开始使用该模型。
Google 表示,Gemini Omni 将现实世界推理能力带入生成式创作,而此次更新使 Omni 1.1 通过 Google AI Studio 中的 Gemini API 达到可用于专业用途的生产就绪状态。无论开发者是在构建生成式视频工作流、创意工具还是媒体编辑软件,这次更新都旨在让生成式视频更可控、迭代更快,并且更适合真实环境部署。其功能重点在于解决原始生成与成品制作之间的实际差距:让素材在更长时间内保持连贯,能够有意识地控制镜头,并在最终渲染前以更低成本进行迭代。
主要功能包括:
- 将场景延展至最长 40 秒,并提升视觉一致性和叙事连贯性。
- 设置起始帧和结束帧,以创建平滑、专业的镜头运动和转场。
- 使用 360p 预览更快迭代,并在创作过程中节省成本。
- 将最终项目放大至 4K 分辨率,以获得更精致、专业的观感。
用于更长叙事的场景延展
场景延展可让你接续现有视频,从停止的位置无缝继续生成画面。借助 Omni 1.1,模型现在可以分析最多 10 秒的前文上下文——相比之前只参考最后 1 秒的模型有明显提升。Google 表示,这带来了更好的视觉一致性和叙事贴合度,使开发者能够构建更长的故事,或转向新的创意方向。视频可按 10 秒一段进行延展,累计最长可达 40 秒。对于生成式视频而言,长期连贯性一直是更难解决的问题之一,因为模型通常只生成较短片段,一旦素材被拉长,画面就会出现漂移。
公告中发布的示例提示词链展示了这一功能:
- Prompt 1: Camera slightly pans and we now see she is talking to a man with curly hair, we see man's back, he says "I see it too" dramatic music score
- Prompt 2: Camera slowly pulls out, forgotten dusty catacombs, dramatic music score.
- Prompt 3: Camera slowly pulls out, a vast library where shelves and books float weightlessly in a dusty void, dramatic music score.
另一组序列展示了电影级镜头延续:
- Prompt 1: Continue the video. Execute a cinematic optical dolly-zoom shot. The camera dollies forward while simultaneously zooming out, keeping the character's frozen shocked face locked at the exact same size. The long corridor of stone pillars in the background dramatically stretches and deepens with intense optical perspective distortion. Clean architecture, continuous unbroken shot.
- Prompt 2: Continue the video. The camera executes a fast mechanical snap-zoom directly into the character's wide eyes. Stylized cinematic camera control.
- Prompt 3: Continue the video. Time completely freezes into a static moment: the character, their windblown coat. The camera performs a smooth, high-speed 360-degree orbital rotation around the frozen character, showcasing dramatic 3D depth and parallax across the colonnade. Flawless continuity.
更多示例展示了基于对白的延展:
- Prompt 1: The man in the blue sweater replies: "Did your father go out on the boat too?"
- Prompt 2: The camera pulls back in one continuous movement as he continues his story: "He used to say that this harbour had a soul. And that the boats were a part of us." The music swells.
另一组提示词则覆盖了主持人式延续:
- Prompt 1: The looks directly at the camera and talks about what the final chapter will end with!
- Prompt 2: He then stands up walks around the desk to the camera and says "what would you choose?"
Google 还介绍了如何通过 Gemini API 延展场景。
指定首帧和末帧
开发者可以通过指定镜头的起始帧和结束帧,获得平滑的转场和镜头运动。Omni 1.1 会在两个关键帧之间生成连续视频,因此适合复杂的环绕镜头、变焦转场或无缝循环片段。示例提示词包括:
- Prompt 1: A close-up low-angle shot of a stylish drummer in a beige suit playing a red drum kit in a grand hall transitions as the camera whip-pans to the side, revealing an older saxophonist playing alongside a ballet dancer spinning in a white outfit under soft purple stage lights. One continuous shot, no jump cuts.
- Prompt 2: The camera zooms into the TV screen, where we see the same woman and the same scene from the beginning. Seamless video. One continuous shot, no jump cuts.
以 360p 更高效地生成草稿
Omni 1.1 可生成 360p 分辨率的轻量预览,速度最高快 60%,成本仅为该模型标准 720p 分辨率的三分之一。Google 指出,这有助于快速原型设计、分镜迭代以及开发平台中的快速渲染。公司还补充脚注说明:最高快 60% 的生成速度基于 360p 与 720p 分辨率的系统吞吐量对比。
一个示例提示词为:A microscopic view of iridescent marine diatoms, displaying intricate, glass-like silica shells with breathtaking natural symmetry. The colors range from deep volcanic amber and warm copper to vibrant turquoise and violet, mimicking the rich palette of earth and ocean. Tiny, delicate structures glow softly against a clean dark field background. High-fidelity scientific imaging, sharp details, organic textures, micro-photography. Maintain the microscope lens effect throughout the entire video.
放大到 4K 分辨率
该模型可生成精致的高分辨率 1080p 或 4K 输出,适合专业制作。示例提示词包括:
- Prompt 1: Fish swimming, tracking shot
- Prompt 2: A little chipmunk darting out of the woods from the left side of the screen and sniffing the air inquisitively before darting out of frame on the right side
- Prompt 3: Cinematic macro close-up of vibrant golden-orange Japanese maple leaves on a delicate branch, gently rustling and swaying in a soft, rhythmic autumn breeze. Sunlight filters through the translucent foliage, creating a warm, glowing effect. Shallow depth of field, dreamy bokeh background, hyper-detailed textures, photorealistic, 4k.
在多模态输入中添加视频参考
在构建场景时,开发者可以参考最长三秒的视频,以便基于视频参考保持视觉上下文和角色一致性,使生成结果锚定于所提供的素材,而不仅仅依赖文本提示。示例提示词如下:
Use the three uploaded videos of dancers and replace them with the provided characters. Have them perform their individual dances from the reference videos, all together in the large, open space from the provided image. The dog character dog.png should do the classical dance from dance3.mp4. The octopus octo.png should do the hip hop dance from dance1.mp4, and the bear bear.png should do the breakdance from dance2.mp4. The final result should be one continuous shot with no scene cuts.
开发者可构建的应用场景
Google 还分享了若干概念,展示这些新能力如何在自定义工具和创意工作流中落地:
- 一个关键帧转场应用,用户放入首帧和末帧,并通过预设或提示框生成中间转场,借助 Omni 的全上下文推理生成看起来真实的镜头运动。
- 一个房间导览应用,镜头穿行于各个房间之间——环绕、推进、拉回——在一天中最合适的时刻展示理想化状态下的住宅,同时不添加不存在的家具或细节。
- Draft Room,它让创意探索变得低成本且结构化:生成 3-4 个 360p 草稿变体,每次只改变一个因素,并并排比较。Google 指出,创作者通常会先生成多个视频,再选出最终版本。
客户如何将 Omni Flash 投入生产
Google 表示,客户已经通过 Agent Platform API 使用 Gemini Omni Flash 推动真实生产。Adobe 已将 Gemini Omni Flash 集成进 Adobe Firefly——Adobe 的生成式 AI 创意工具系列——公司还引用了一个展示其视频编辑能力的视频。合作伙伴阵容覆盖设计工具、创意平台和云基础设施,说明生成式视频正被嵌入成熟的专业工作流中,而不仅仅作为独立模型提供。
合作伙伴说明了他们如何使用该模型:
“Gemini Omni Flash 是 Figma Weave 中最强大的视频模型之一,画布帮助创意团队在每一次生成之上继续构建——附加参考资料、分支出不同版本,并塑造出独特内容。借助延展、更丰富的参考素材和 4K 分辨率,Gemini Omni Flash 让团队不再只是生成视频,而是真正对其进行导演。” — Itay Schiff,Figma Weave 创意总监。
“At GMI Cloud, we give creators centralized access to the world's most capable models. What stands out about Gemini Omni Flash is its accuracy: the details hold up under scrutiny. For customers creating educational and explanatory content, where getting things right is essential, that reliability matters more than any single feature. Omni has made AI video viable for a segment that previously couldn't rely on it.” — Louisa Guo,GMI Cloud 市场营销副总裁。
“Omni Flash fits naturally into how people already use Runway: start with a prompt, an image or a video, then generate or edit from there. It's another way for our users to move quickly between ideas.” — Jamie Umpherson,Runway 首席创意官。
现在即可使用 Gemini Omni 1.1 Flash 构建
Gemini Omni 1.1 Flash 的定价表。
Omni 1.1 正在 Google 开发者生态系统中逐步开放:
- 在 Google AI Studio 中开始构建:开发者可直接在 Google AI Studio 中试用 Omni 1.1。
- 部署到 Gemini Enterprise Agent Platform:企业开发者可以在该平台上部署 Omni 1.1。
- 查看开发者文档:官方文档、cookbook 和 提示词指南 说明了如何将场景延展、视频参考和放大功能集成到应用中。
从今天起,全球所有 Google AI Plus、Pro 和 Ultra 订阅用户也可在 Google Flow 中使用 Omni 1.1。场景延展功能同时也向全球所有 Google AI Plus、Pro 和 Ultra 订阅用户在 Gemini 应用中开放。
在收件箱中获取 Google 的最新消息
订阅我们的新闻简报,获取产品更新、活动信息、特别优惠等更多内容。
完成。只差一步。
请检查你的收件箱以确认订阅。
你也可以使用其他电子邮件地址订阅。
你的信息将按照 Google 的隐私政策使用。你可以随时退订。