An infinite visual browser generated entirely on demand in real time
https://flipbook.page/n/8252156e2a0a45fc97ac90cffc5bb1c3Reader - page text saved at the timeSearch Sketchapedia
Anatomy and Evolution of the Horse
/
Clear
Connect to Live Video Stream
Tap anywhere on the page to expand
Flipbook is an infinite visual browser generated entirely on demand in real time.
Zain Shah, Eddie Jiao, Drew Carr
Every "page" you land on is an image. Click on anything in the image and you will get a new image exploring that thing in more depth. What you see contains no HTML, no code, no specific links or fields. The entire web is just generated pixels on your screen.
Is the text rendered by the image model too?
All text on the screen is rendered as pixels by the image model. There are no text overlays applied to the images. Occasionally, the image model might render text imperfectly, or in the wrong place. That will get better as the models improve.
Where does the information come from?
The information in the images comes from a combination of an agentic web search and the image model's own world knowledge. There may be occasional inaccuracies, but it's a useful starting point and usually grounded in real data online. You can expect a similar level of factual accuracy to what you might get using ChatGPT/Gemini/Claude.
Why are you doing this?
A picture is worth a thousand words, yet we fill our screens with mostly text and colored rectangles.
We built this because we saw another way. The walls of text and generative UIs being sold as the future felt like sipping an ocean of wisdom through a tiny straw.
We wanted a computing experience full of rich beautiful visuals made just for us, generated just in time. The screen you're reading this on is already presenting you an image, it's just generated with rigid code and rules that makes it difficult to communicate complex and detailed ideas. Flipbook does away with this, and the benefit is that you can expect it will truly find the most effective way to communicate anything to you, independent of how hard it may be to write code to do it. If the most effective way to communicate something were a single word, an illustration, or a photorealistic rendering, that's what you'd see.
What is the "live video stream" feature?
The live video stream is an experimental feature that turns the static images generated by our system into a more continuous video stream. It animates each image you explore and creates seamless transitions between them. The behavior is a bit unpredictable at the moment, and it is very resource-intensive, so we are keeping it behind a toggle that you can turn on and off as you wish.
Currently, it brings together two separate systems: a custom, highly optimized video generation model and our image generation system. In the future, we expect to integrate these into a single system.
Where does the inference come from?
Our compute is sponsored by Modal and we are generously backed by South Park Commons
What comes next?
Right now, Flipbook is an experiment. It's designed for open-ended exploration and learning, but the overall experience is flexible. As image and video models become more accurate and performant, Flipbook pages could include more real data, be more interactive, and even take actions and store their own data.
Practically speaking, that's the difference between using Flipbook to research your next trip while booking everything elsewhere, and doing the whole thing inside Flipbook.
With these capabilities, more of what you currently expect to require separate apps and websites can happen in something that looks and works like Flipbook. We've prepared some examples of what that might look like below.
We imagine a world where all of the tools you use are as rich and visual as the world we live in.
This demonstration uses pre-generated video and has been cut for speed