AI images in your video
Most of a good product video can be filmed from the page itself — that is the point of the tool. But some beats have nothing to point a camera at: the frustration before someone finds you, an abstract idea, a person, a place, a before. For those, PageToVid draws the shot instead of capturing it. This is the feature that separates a screen recording from a film, and it is also the one most easily overused.
What it is
Any scene can be an image scene with a generate description instead of a captured screenshot. You describe the shot in a sentence; the model draws a photorealistic still at the film's aspect ratio, and it is composited in with the same motion and transition vocabulary as every other scene.
The generation runs during the render, in the same pass as the screen capture, so there is no separate step and nothing to upload. The result is stored in your asset bank, so the same visual can be reused in later films without paying for it again.
When a generated image is the right answer
The test is simple: is there something on the page that shows this? If yes, film it — a real screenshot of your real product is more persuasive than any illustration. If no, draw it.
Beats that usually have nothing to film:
- The problem, before your product exists in the story. A cluttered desk, a spreadsheet at midnight, a queue.
- A feeling — relief, focus, momentum — that a UI screenshot cannot carry.
- A person: a customer, a persona, a face for a quote.
- A physical context your software only implies: a warehouse, a clinic, a kitchen.
- A metaphor you are deliberately using to explain something abstract.
Beats that almost never need one: your features, your pricing, your dashboard, your numbers. Those are on the page, and drawing them invents a product that does not exist.
Writing a description that works
The descriptions that come back well are concrete and short on adjectives.
- Say the subject, the framing and the light. "A freelance designer at a cluttered desk late at night, lit by one monitor, shot over the shoulder."
- Name the mood through the scene, not the adjective. "Rain on a café window" beats "a moody atmosphere".
- Say the format when it matters. A 9:16 film wants vertical compositions; a face in the middle third survives the caption band at the bottom.
- Do not ask for text in the image. Generated lettering is unreliable, and PageToVid draws real typography over the image anyway.
Avoid asking for logos, brands or recognisable public figures. The generator will usually refuse, and the render tells you it refused rather than substituting something else.
What it costs
One AI still is 40 credits, charged on top of the 40 credits for the render itself. The house video clip is 160 — see the clips guide for when that is worth it.
Two mechanisms stop a film surprising you:
ai_planis returned before anything is generated, with the full bill itemised.ai_budgetcaps the spend in advance. Scenes beyond the cap fall back to a drawn card rather than failing the render.
Identical generations are de-duplicated by content hash, so re-rendering the same film does not pay for the same image twice. Changing the description is what triggers a new generation — which is also how you iterate deliberately.
What happens on a free plan
AI generation needs a paid plan, from Starter upwards. This is enforced, not merely advertised.
On the free plan a render still produces real retina screen recordings of your page, an AI script, an AI voice-over, subtitles and the full motion-graphics catalogue. What it does not do is spend generation budget. Scenes that asked for a drawn visual are downgraded to a text card, and the render's warnings say exactly which ones and why — so a free render is a complete film with known substitutions, not a broken one.
Generating outside a render
generate_asset makes an image straight into your bank with no render at all. That is the cheap way to iterate on a description: try three, look at them, then reference the one you kept from a storyboard scene.
It is also how you build a small library of reusable shots — a persona, an office, a product context — that keeps a series of videos looking like they belong together without paying for the same generation each time.
Turn your website into a video — free
Paste a URL. PageToVid scripts, records, voices and renders it automatically.
Create your first video →Frequently asked questions
How much does one AI image cost?
40 credits, on top of the 40 for the render. The house AI video clip is 160, and a clip from a named model is priced from its real cost. create_video and create_animation both return an ai_plan with the bill before anything is generated.
What happens if the generator refuses my description?
The scene falls back to a drawn card and the render's warnings say the visual was refused, with the reason. You are not charged for a refused generation, and the render still completes.
Can I use my own images instead?
Yes. Your asset bank holds uploads as well as generations, and a scene can reference a banked asset directly. If you already have a photograph of the thing, use it — it will always beat a generated approximation.
Will the same description give the same image twice?
Identical generations are de-duplicated by content hash within your account, so re-rendering an unchanged film reuses the image rather than paying again. Change the description and you get a new generation.
Can I keep the same person across several scenes?
Yes — that is what characters are for. Create a character once and reference it by name, and the same face comes back in every scene and every clip that names it. Without one, two separate generations of "a freelance designer" are two different people.
Does an AI image make the video longer?
No. Scene length comes from the narration, not the visual. A generated still does make the render slower, though — generation runs during the render, and clips considerably more so than stills.