Veo

2 posts

figma2 min readCurated summary

Creativity meets precision with Gemini 3 Pro Image Pro | Figma Blog

Google’s Nano Banana Pro, part of Gemini 3 Pro, brings more precise and context-aware image editing to creative workflows. The model can generate variations while preserving a design’s visual identity, including its palette, typography, composition, and subject likeness. Figma presents it as a tool for refining and extending ideas across products rather than simply regenerating images. ## Design Coherence Across Variations - Nano Banana Pro retains a design’s “visual DNA,” including color, texture, type, composition, and imagery. - In Figma Buzz, it creates branded social-asset variations while preserving logos, simple illustrations, and overall composition. - It can adapt illustrations for different contexts, such as converting a set of winter-themed images into dark-mode versions with minimal prompting. ## Extending Existing Work - The model can place illustrations into new environments while matching lighting and mood-board references. - In Figma Slides, it embedded an astrology-app illustration into a cozy reading scene and added complementary details such as a star-shaped light. - It can reframe portraits, change camera angles, and update backgrounds while maintaining a person’s likeness. - This makes it useful for keeping employee headshots and other brand imagery visually consistent. ## Building Composite Scenes - Figma Weave enables users to combine graphics, copy, photography, and other visual elements into unified scenes. - Nano Banana Pro helps maintain coherence when disparate assets are composited together. - These scenes can be extended with Google’s Veo for motion, as well as 3D models, upscalers, background-removal tools, and prompt-refinement models. ## Editing in Context Across Figma - The model is integrated into Figma’s products, including Buzz, Design, Slides, and Weave. - Users can refine existing images instead of starting over, making targeted changes to text, spot colors, and other details. - Typography can be localized, faces remain natural, and the surrounding visual context is preserved. Nano Banana Pro is most valuable when creative teams need fast iteration without sacrificing brand consistency or image integrity. Figma recommends experimenting with it directly as a flexible tool for refining, remixing, and expanding visual concepts.

Read original(opens in new tab)
googleOriginal article

Bringing 3D shoppable products online with generative AI (opens in new tab)

Google has developed a series of generative AI techniques to transform standard 2D product images into immersive, interactive 3D visualizations for online shopping. By evolving from early neural reconstruction methods to state-of-the-art video generation models like Veo, Google can now produce high-quality 360-degree spins from as few as three images. This progression significantly reduces the cost and complexity for businesses to create shoppable 3D experiences at scale across diverse product categories. ## First Generation: Neural Radiance Fields (NeRFs) * Launched in 2022, this initial approach utilized NeRF technology to synthesize novel views and 360° spins, specifically for footwear on Google Search. * The system required five or more images and relied on complex sub-processes, including background removal, XYZ prediction (NOCS), and camera position estimation. * While a breakthrough, the technology struggled with "noisy" signals and complex geometries, such as the thin structures found in sandals or high heels. ## Second Generation: View-Conditioned Diffusion * Introduced in 2023, this version addressed previous limitations by using a diffusion-based architecture to predict unseen viewpoints from limited data. * The model utilized Score Distillation Sampling (SDS), which compares rendered 3D models against generated targets to iteratively refine parameters for better realism. * This approach allowed Google to scale 3D visualizations to the majority of shoes viewed on Google Shopping, handling more diverse and difficult footwear styles. ## Third Generation: Generalizing with Veo * The current advancement leverages Google’s Veo video generation model to transform product images into consistent, high-fidelity 360° videos. * By training on millions of synthetic 3D assets, Veo captures complex interactions between light, texture, and geometry, making it effective for shiny surfaces and diverse categories like electronics and furniture. * This method removes the need for precise camera pose estimation, increasing reliability across different environments. * While the model can generate a 3D representation from a single image by "hallucinating" missing details, using three images significantly reduces errors and ensures high-fidelity accuracy. These technological milestones mark a shift from specialized 3D reconstruction toward generalized AI models that make digital products feel tangible and interactive for consumers.