AI Content Generator for Stock Metadata
This is the behind-the-scenes guide. It explains what the vision model reads in a photo, how the Adobe-specialist prompt shapes the output, how to get better results, and where the AI still needs a human. If you just want the tool, start with the AI metadata generator page instead.
What the vision model does
Metadata Studio sends each upload to a vision-capable model running on the provider key you supply. The model does not read your filename for its main answer; it reads the pixels. It identifies the visible subject, the setting, the composition, the dominant colours and the mood, then writes three things back: a title, a description and an ordered keyword list. On the Shutterstock tab, a separate classification call also assigns categories, which is a distinct step from metadata generation.
Because the model works from what is visible, the quality of the output depends heavily on the image itself. A sharp, clearly lit photo with one obvious subject produces the most reliable metadata. A busy collage, a heavy filter, or a frame where the subject is tiny gives the model less to work with, and the result is more generic.
How the Adobe-specialist prompt approaches an image
The prompt is written for stock workflows, not for general chat. It walks the model through the same decisions a human contributor would make, in a fixed order, so the output stays consistent from image to image.
- Visual facts first. The model must state what is visibly present — objects, people, setting, action — before it is allowed to generalise.
- Context and mood. It then reads the scene as a whole: indoor or outdoor, season, time of day, and the feeling the frame conveys.
- Commercial concept. It considers what a buyer would use the image for, so the title and keywords cover the concept and not just the literal subject.
- Hallucination guard. The prompt explicitly forbids inventing objects, attributes or events that are not visible, which is the main defence against confident but wrong keywords.
- People and location policy. People are described generically rather than named, and a location is only stated when it is genuinely identifiable in the frame.
- Title rules. Titles are natural, front-loaded and kept within the 70-character hard cap, without keyword stuffing or hype words.
- Keyword scoring. Keywords are ordered by relevance, so the strongest terms come first, and the list stays inside the 20–49 range.
- Compliance filter. A final pass removes banned terms and normalises the text — smart quotes are cleaned, singular and plural duplicates are merged, and each field is capped to its length limit.
Getting better results
A few habits consistently improve the output. Use a vision-capable model; the model picker labels which ones can read an image, and the live discovery button refreshes the list against your own key. Keep batches modest when you are learning what a provider does with your genre, so you can review a small set before committing a large queue. Run two or three images at a time to reduce the chance of rate limits, and use the retry-failed control instead of restarting a whole batch.
Clear filenames still matter, but for a different reason: if a provider rejects an image entirely, the item falls back to a text-only retry based on the filename. A descriptive filename gives that fallback something to work with, even though the primary path ignores it.
Editing the output
The AI produces a draft; you stay the editor. Every title, description, keyword and category is editable inline, and each result shows which provider and model generated it. That makes it easy to keep the strong fields and regenerate only the image that missed. Whole-batch controls let you stop, retry failed items or clear the queue, and a single item can be regenerated or cancelled without touching the rest.
Honest limitations
This is a metadata generator, not an image generator. It works on raster images only — JPG, JPEG, PNG, WEBP and GIF, up to 25 MB each — and it does not accept PDF, EPS, AI or SVG files. It exports Adobe Stock and Shutterstock CSVs only, and it does not promise ranking or sales. The model can misread an ambiguous image, and it does not know your marketplace's policy details, so a human review before upload is still the last step. Keys are yours: they live in this browser's localStorage and are never stored on the server.
AI content generation FAQ
What does the vision model actually read in an image?
It reads the pixels, not a filename. The model identifies visible subjects, setting, composition, colours and mood, then writes a title, a description and an ordered keyword list from what it can see. A filename is only used when a model cannot process the image and that item retries in text-only mode.
Is the AI content generator the same as an AI image generator?
No. Metadata Studio does not create images. It reads an existing raster upload — JPG, JPEG, PNG, WEBP or GIF up to 25 MB — and generates stock metadata for it. There is no text-to-image feature.
Will it name real people or guess locations?
The prompt instructs the model to avoid inventing identities and to describe people generically, such as a woman in a red coat, rather than naming someone it cannot know. Locations are described only when they are visibly identifiable, and the model is told not to fabricate a place.
How many keywords does it write, and how are they ordered?
Each image gets between 20 and 49 keywords, ordered by relevance, with 20 to 40 treated as the practical range. Singular and plural duplicates are merged and banned words are removed before the list reaches you, and you can edit or reorder every entry.
Why do some results say text-only?
If a vision model rejects an image, that item retries once in text-only mode using the filename and is flagged text-only. Switch to a vision-capable model and regenerate the item to get metadata based on the actual image.
Can I trust the output without editing it?
Treat the first pass as a strong draft, not a final answer. Ambiguous images can be misread, and marketplace-specific rules such as Editorial capture details remain your responsibility. Every field is editable inline, and each result shows the provider and model that produced it.
See the AI describe your own images
Add a provider key and watch a photo turn into a title, description and keywords.