Most AI image tools ask you to describe what you want in words. Google Whisk flips that around. It is a Google Labs experiment that lets you generate images by feeding it other images, so instead of writing a long paragraph of instructions you hand the tool a picture for the subject, a picture for the scene, and a picture for the style, and it produces something new that borrows from all three. The point is speed of visual thinking rather than precision.
There is one important thing to know up front, because it changes how you should read the rest of this piece. Whisk started life as a standalone tool at labs.google, but in 2026 Google folded its capabilities into a larger product called Flow. This article explains what Google Whisk is, how the remixing approach works, what it was good for, and exactly where those features live now.
Images as the prompt, not text
The core idea behind Whisk is simple to state and genuinely different in practice. Traditional generators are text-to-image: you type a description, the model reads it, and it draws. Whisk is closer to image-to-image, or more precisely image-as-prompt. You do not have to find the right words for a mood, a texture, or a character. You show the tool an example and let it infer the rest.
Whisk breaks a picture down into three separate inputs, and you can supply an image for any or all of them:
- Subject is the main thing you want in the frame: an object, a character, a product mockup, or a logo.
- Scene is the setting or background where that subject should sit.
- Style is the look and feel: an oil painting, a soft plush toy, an enamel pin, a retro vector, a cinematic photo.
You mix and match. Keep a subject but change the style five times. Keep a style but drop the subject into a new scene. Because you are pointing at reference images rather than describing them, the loop from idea to result is very short, which is the whole selling point.
How Google Whisk works under the hood
Whisk is a front end stitched over two of Google’s AI models working in sequence. It does not read your images directly into a paintbrush. Instead it goes through a translation step first.
When you upload your subject, scene, and style images, Gemini looks at each one and writes a detailed caption describing it. Google calls this "essence capture," the idea being that the caption records the core character of an image rather than every pixel. Those written descriptions, not the original files, are then passed to Imagen, Google’s text-to-image model, which generates a brand new picture that weaves the three descriptions together.
That design has a consequence worth understanding. Because the pipeline runs through Gemini’s interpretation before Imagen draws anything, the output captures the vibe of your inputs but will not be a faithful copy of any of them. A face you upload will not come back as the same face. A logo will not be reproduced exactly. Whisk is built for remixing and exploration, not for pixel-perfect edits or brand-accurate reproduction. If you have used Google’s Gemini image models or run Imagen inside Google AI Studio, this is the same underlying generation engine wrapped in a much more visual interface.
What Whisk is good for
Whisk was never pitched as a replacement for a designer or for a controllable generator like Midjourney. Google positioned it as a toy in the good sense: a fast way to riff on an idea when you cannot yet put it into words.
That makes it a strong fit for early ideation. If you are exploring a character design, a mascot, a sticker set, or a mood for a campaign, Whisk lets you generate a dozen directions in the time it would take to write one careful prompt. It also has quick preset styles, including playful ones like turning a subject into a sticker or a plush toy, which are useful for social content and quick concept art.
Where it is a poor fit is any job that needs exactness. Product photography that must match a real item, text rendered accurately inside an image, or a specific person reproduced reliably are all outside its comfort zone.
Whisk Animate: from a still image to a short clip
Google later added Whisk Animate, which takes a still image you generated and turns it into a short video clip. That feature is powered by Veo, Google’s video generation model, and it extended Whisk from a two-model image pipeline into something that could produce motion as well. In practice you would generate a still with the subject, scene, and style approach, then animate the frame you liked into a few seconds of video.
Whisk Animate sat behind a paid Google AI subscription rather than the free tier, which is typical for video generation because it is far more compute-heavy than making a still image.
Availability, limits, and the move to Flow
Here is the part that matters most in 2026. Whisk launched as a Google Labs experiment and was rolled out to a growing list of countries over time, but it was never available everywhere at once, and Labs experiments come with no guarantee of permanence. In February 2026 Google announced that it was consolidating its creative Labs tools, and on April 30, 2026 the best of Whisk moved into Flow, Google’s unified platform for AI image and video creation. ImageFX, Google’s older text-to-image tool, was folded into the same product.
The remixing approach did not disappear. The subject, scene, and style system, the Imagen-powered generation, and the quick styles all live inside Flow now, alongside Veo-based video. What changed is the front door. Migration was opt-in rather than automatic, so anyone who had projects in Whisk needed to import them into Flow before the deadline, and media that was not migrated or downloaded by April 30 was deleted. Existing AI credits carried over because both tools share the same credits system. Flow is still not available in every country, so users in unsupported regions lost access without a migration path.
The practical takeaway: if you go looking for a standalone Whisk today, you will be redirected toward Flow. The capabilities described in this article are the ones to expect there.
How Whisk differs from standard text-to-image tools
It helps to place Whisk against the tools most people already know. With a text-to-image generator such as Midjourney or DALL-E, the burden is on your words. You get precise control over what you ask for, but you have to know how to ask, and describing a subtle style in text is genuinely hard. Our guide to Midjourney covers that prompt-driven workflow in depth.
Whisk moves the burden to reference images. You trade fine-grained control for speed and for the ability to communicate a look you cannot name. That is a real advantage when you are brainstorming and a real limitation when you need a specific outcome.
Frequently Asked Questions
Is Google Whisk still available in 2026?
Not as a standalone tool. On April 30, 2026, Google moved Whisk’s capabilities into Flow, its unified image and video platform. The subject, scene, and style remixing still exists, but you now access it through Flow rather than a separate Whisk app.
Is Google Whisk free to use?
Whisk’s core image remixing was free during its Labs run, subject to usage credits. Video generation through Whisk Animate required a paid Google AI subscription. Those same paid and free tiers broadly carry over inside Flow, where the heavier video features sit behind a subscription.
What models power Google Whisk?
Two work in sequence. Gemini reads your uploaded images and writes detailed captions of each one, and Imagen then generates a new image from those captions. Whisk Animate adds Veo, Google’s video model, to turn a still into a short clip.
Does Google Whisk copy my uploaded images exactly?
No. Because the pipeline turns your images into text descriptions before generating anything, the output captures the essence of your inputs rather than reproducing them precisely. A face or logo you upload will inform the result but will not come back identical.
What is the difference between Whisk and a text-to-image tool?
A text-to-image tool asks you to describe what you want in words, which gives precise control but demands good prompting. Google Whisk asks for reference images instead, which is faster and better for communicating a look you cannot easily put into words, at the cost of exact control.
Can Google Whisk make videos?
Yes, through Whisk Animate, which uses Google’s Veo model to turn a generated still image into a short video clip. That video capability now lives inside Flow, where Veo powers image-to-video generation on paid plans.
What happened to my Whisk projects after the Flow migration?
Migration was opt-in. Users were prompted to import their Whisk and ImageFX projects into Flow before April 30, 2026. Any media not migrated or downloaded by that date was permanently deleted, and existing AI credits transferred automatically because both tools share one credits system.
Is Whisk good for professional design work?
It is best for the early, exploratory stage of a project: quickly generating directions, moods, and concepts. For final assets that need exact detail, accurate text, or brand-precise reproduction, most people move to a more controllable tool once Whisk has helped them find a direction.