Media editor: build vs buy in the age of AI
Polotno

Design files are becoming text

By Misha Zhuro · Oct 7, 2026

Marketing design at scale has a new face – JSON.

Photoshop, InDesign and Illustrator were built long before LLMs existed, and their formats were designed for one app to open. Design work sits in blackbox files (.psd, .ai, etc.) that other software can't read. That's where JSON comes into play.

banner.psdopened in a text editorA text editor can't make sense of a .psd – reading it takes Photoshop or a PSD parser.
8BPS�������������F���8������������8BIM���������Transparency�8BIM���������`�������`������8BIM�������H���HLino����mntrRGB XYZ ���������1��acspMSFT����IEC sRGB�����������������������-HP ������������������������������������������������cprt���P���3desc�������lwtpt��������bkpt��������rXYZ��������gXYZ���,����bXYZ���@����dmnd���T���pdmdd��������vued���L����view�������$lumi��������meas�������$tech���0����rTRC���<����gTRC���<����bTRC���<����text����Copyright (c) 1998 Hewlett-Packard Company��desc��������sRGB IEC61966-2.1������������sRGB IEC61966-2.1��������������������������������������������������XYZ �������Q��������XYZ ����������������XYZ ������o���8�����XYZ ������b���������XYZ ������$���������desc��������IEC http://www.iec.ch������������IEC http://www.iec.ch����������������������������������������������desc�������.IEC 61966-2.1 Default RGB colour space - sRGB�����������.IEC 61966-2.1 Default RGB colour space - sRGB����������������������desc�������,Reference Viewing Condition in IEC61966-2.1�����������,Reference Viewing Condition in IEC61966-2.1��������������������������view����������_.��������������\�����XYZ �����L�V�P���W��meas��������������������������������sig ����CRT curv�����������������������#�(�-�2�7�;�@�E�J�O�T�Y�^�c�h�m�r�w�|���������������������������������������������������������������%�+�2�8�>�E�L�R�Y�`�g�n�u�|�����������������������������������������&�/�8�A�K�T�]�g�q�z��
banner.jsonopened in a text editorA text editor shows everything: sizes, text, font, position without any 3rd party tools.
{
"width": 1080,
"height": 1080,
"pages": [
{
"children": [
{
"type": "text",
"x": 80,
"y": 120,
"text": "Summer sale – 40% off",
"fontSize": 72,
"fontFamily": "Inter",
"fill": "#111111"
}
]
}
]
}

A design is just text

You don’t need to be technical to read JSON – it’s structured text describing your design as objects and their properties. No specialized software to open it, and any LLM will easily read it.

Here's a headline on a banner:

json
{
  "width": 1080,
  "height": 1080,
  "pages": [
    {
      "children": [
        {
          "type": "text",
          "x": 80,
          "y": 120,
          "text": "Summer sale – 40% off",
          "fontSize": 72,
          "fontFamily": "Inter",
          "fill": "#111111"
        }
      ]
    }
  ]
}

That's the whole design – position, type, copy, color – all of it in plain sight. Change "40% off" to "50% off", bump fontSize, swap fill for your brand token, and you've edited a design without opening a design tool. A human can read it and a machine can rewrite it, without the format getting in the way.

What this unlocks

  1. Designs live inside workflows.
    A JSON design is a living object your systems continuously interact with: plug it into any pipeline and it gets updated, re-rendered and published on a schedule or on a trigger. A price changes or a campaign goes live, and the creative updates itself without anyone opening an editor.

  2. Generate designs from a prompt.
    An LLM can write valid JSON directly, so "make 20 banner variants for this campaign" is a single API. It’s early days, but it’s getting better fast.

  3. Bulk personalization at scale.
    Swap text, images and brand tokens programmatically to render thousands of localized or per-customer creatives from one template.

  4. Version control for design.
    JSON diffs cleanly in Git, so you can review changes and roll them back exactly like code.

  5. Automated brand compliance.
    Because every object and property is readable, you can validate designs against brand rules and flag issues before publishing.

  6. Works everywhere.
    Designs move freely between your editor, backend, renderer and third-party tools.

Designers still work on a canvas

JSON is the file format. The editor is still the interface, and nobody positions text by typing coordinates. Instead:

  1. The editor reads the JSON and draws it on a canvas.
  2. You drag, type, resize, restyle as usual.
  3. The editor writes the result back to JSON.

Every canvas action maps to a property change. Move an element and x and y change, resize it and width and height change – and so on for everything else the toolbar exposes. The editing experience is the same as any design tool; only the saved file is different.

{
"type": "text",
"x": 180,
"y": 380,
"text": "Marketing design at scale has a new face",
"width": 910,
"fontSize": 89,
"fontFamily": "IBM Plex Sans",
"fill": "rgba(255,255,255,1)"
}
JSON FILE FORMAT
Marketing design at scale has a new face
Generate designs from a prompt
Bulk personalization at scale
Automated brand compliance
READ THE GUIDE
Priya

Humans and machines can edit the same file at once

Because a design is data, it can be stored like any other data: a row in your database, an object in S3 or a file in Git. Anything with access can read and write it:

  • the editor writes when a designer makes a change
  • a script writes when it fills a template with new data
  • an LLM writes when it generates or rewrites elements
  • a validator reads to check fonts, colors and sizes

None of these need a copy. They all read and write the same record, so there is no export step, and no production version that quietly goes out of sync from the original. Normal concurrency rules apply: last write wins by default, and you add locking or versioning if two writers can collide.

In practice a designer builds and maintains the template, and scripts fill it in and re-render it. The template is the only thing a human has to touch.

Rendering: JSON in, file out

A readable file still has to become an image. That is a separate service: it takes JSON plus the referenced assets and returns a statics like PNG, JPEG, PDF, animated GIFs or videos like MP4 or WEBM. Usually it runs a headless browser or a server-side rendering library.

A full automated pipeline is three steps:

  1. Load a template JSON.
  2. Replace values – text, image URLs, colors – with data from your source (CSV, database, product feed, API).
  3. Send it to the renderer and store the returned file.
javascript
const design = JSON.parse(template)
design.pages[0].children[0].text = product.title
design.pages[0].children[1].src = product.image

const png = await render({ design, format: "png" })

You can run that in a loop for a thousand of creatives, or on a webhook when a parameter gets updated.

products.csvone row per ad
ProductPriceWas
banner.json5 values replaced
[
{ "id": "name", "text": "Trail Runner 2" },
{ "id": "price", "text": "$89" },
{ "id": "was", "text": "$129" },
{ "id": "badge", "text": "-31%" },
{ "id": "photo", "src": "runner.jpg" }
]
banner.pngrendered
photorunner.jpg
-31%
Trail Runner 2
$89$129

Where this is used

  • Web-to-print – print-ready PDFs per SKU or per customer order.
  • Real estate – flyers, social posts and window cards generated from listing data.
  • E-commerce and print-on-demand – creative that follows the product feed.
  • Ad ops – the same design output in every placement size, locale and offer.
  • Internal tools – templates non-designers fill in, with brand rules enforced in code.

Compared to other design formats

FormatPlain textDiffable in GitEditable by script or LLM
.psd, .ai, .inddNoNoOnly via the vendor's app or SDK
.pdfNoNoText and images can be pulled out, the design can't be edited
.docx, .xlsxXML inside a zipAfter unzippingWith tooling
.svgYesYesYes, but no pages, locks or fillable fields
Design JSONYesYesYes

HTML, CSS, YAML and Terraform all became automatable for the same reason: they are structured text. Design files are one of the last formats where that is not the default.

Why not just HTML and a screenshot?

Many teams go that way – render HTML templates in a headless browser and screenshot them. But it quickly falls apart when designers get involved. HTML is made for browsers to display. It has no built-in way to select an element, drag it, rotate it or manage layers, so letting a designer edit it means building all of that yourself. But even then CSS layout is a poor fit for the absolute positioning that design work assumes. There's no font-size-to-fit and no snapping. Print is also weak: browser print output gives you RGB PDFs with limited control over bleed, color profiles and PDF/X conformance.

HTML-to-image may work well when only developers ever touch the template. A design document format is the answer when someone needs to open it on a canvas.

The other alternatives, briefly:

  • SVG. Text, fonts and reflow are painful. It's single-page, with no editor semantics like locked layers or editable-field markers. Fine as an output format, awkward as a source of truth.
  • Figma's REST API. Genuinely useful, and read-oriented. You can pull a document tree, but you can't render arbitrary variants server-side at volume, and the .fig file itself stays opaque behind rate limits. It's an integration point rather than a pipeline.
  • Hosted template APIs. Fastest possible start. The tradeoff is that the template lives in someone else's dashboard, the schema is theirs, and you pay per render. Good until you need custom element types, your own editing UI or cost control.
  • Vendor SDKs on closed formats. Strong tooling, but you're automating through their app rather than owning the document.

Migration: getting your existing designs in

You import most of them, and rebuild/fix some by hand.

JSON, Figma, SVG, PDF and PSD all convert into the same design JSON, in the browser or on the server. Your layer structure carries over, so text arrives as editable text and vector paths stay vector instead of being flattened into a picture of themselves. In practice you should expect ~90% of designs to come through correctly; some cleanup may be necessary, but you don't have to rebuild the design from scratch.

What each path looks like:

  • PSD – parsed directly, layer by layer. Text, vector shapes and raster layers arrive as separate editable elements. The PSD's flattened composite is deliberately ignored. Effects the schema can express (gradients, strokes, simple shadows) come through live; the ones it can't (gradient maps, hue/saturation, adjustment layers) are baked into that one layer's pixels, and everything else stays editable.
  • Figma – export the frame as SVG with Outline Text off and Include "id" attribute on, then run it through SVG import. Text stays editable and your Figma layer names survive as ids, which is exactly what you want for targeting fields later.
  • PDF – layouts come through with embedded fonts, vector paths and images preserved. Multi-page works. This is the strongest path for anything print, since print artwork is usually already sitting in a press-ready PDF.
  • SVG – the most predictable of the lot. The one trap is source tools that outline text on export (Illustrator and Canva do this by default) – outlined text can't be turned back into editable text, so fix it at the export step.

Every import produces schema-validated JSON, so bad input fails at the boundary instead of surfacing on row 400 of a batch.

Scale, speed and cost

A single-page raster render is typically in the hundreds of milliseconds to low seconds, dominated by font and image fetching rather than by drawing. Multi-page PDFs and video are slower and largely depend on complexity.

What was architectural weakness is now strength: every design is independent, so 10,000 renders is a queue for a bunch of workers. Cache fonts and remote images aggressively, because the same logo fetched 10,000 times is the most common source of slowness. If you self-host, a browser-based renderer costs roughly what any headless-browser workload costs; if you use a render API, you're paying per asset and should model it against your worst-case campaign volume.

Guardrails: stopping templates from getting wrecked

Once non-designers and LLMs can write to the file, you need constraints in the file:

  • Lock what shouldn't move. Mark logo position, safe areas and background elements as locked or non-selectable.
  • Expose editable fields explicitly. Give the elements that are meant to change stable ids and treat everything else as structure. Scripts and prompts target ids rather than array positions.
  • Validate before rendering. Check the JSON against your schema, then against brand rules – approved fonts, colors from your token list, minimum sizes, required disclaimers.
  • Check for overflow as part of validation. Measure text after substitution and display errors prominently, or even block the next step in a workflow.
  • Diff and approve. Because designs are text, a generated batch can go through the same review as code, with a human approving the template change rather than each of the 4,000 outputs.

Build or buy

You'd need 5 pieces here: a schema, an editor, a renderer that handles fonts correctly, asset storage, and template management. The schema and the storage are easy. The editor and the renderer are where the years go – text layout, font fallback, undo, snapping, and getting the canvas and the renderer to agree pixel for pixel.

Build it if creative editing is your product. Buy or embed it if creative editing is a feature inside a product that's about something else.

The hard parts in production

JSON solves readability. These are the problems that remain, and they're the ones that actually break pipelines.

Text that doesn't fit

This is the number one failure mode of templated design. A product name might be three words or nineteen, and the German translation of it runs about 30% longer than the English. Add Arabic and the text runs the other way entirely. The layout that looked perfect with sample copy breaks on row 400 of the feed.

The format helps only because it makes the fix programmable. You still have to choose a strategy per text element and store it in the design:

  • Shrink to fit – reduce font size until the text fits a fixed box, with a floor below which it's unreadable.
  • Wrap with a line cap – fixed width, maximum number of lines, then truncate with an ellipsis.
  • Grow the container – let the box expand and push elements below it, which requires the surrounding layout to tolerate movement.
  • Fail the render – sometimes the correct answer is to reject the row and flag it for a human.

When the copy doesn't fit

This text box was designed around the English headline, which takes two lines and leaves room for a button underneath. Switch the language to see how each strategy deals with what comes out.

The German version runs to three lines in a box that holds two, so this is where the four strategies start to behave differently.

Shrink to fitchange-font-sizeIt fits at the full size of 50, so nothing changed.
Spring sale
Kostenloser Versand bei jeder Bestellung
SHOP NOW
Wrap with a line capellipsisIt fits in 2 lines, so nothing was cut.
Spring sale
Kostenloser Versand bei jeder Bestellung
SHOP NOW
Grow the containerresizeIt fits, so the box stays the size it was designed at.
Spring sale
Kostenloser Versand bei jeder Bestellung
SHOP NOW
Fail the renderyour own validationIt fits, so the render goes ahead.
Spring sale
Kostenloser Versand bei jeder Bestellung
SHOP NOW

Test with your longest locale and your longest real product name, not with lorem ipsum. Then make overflow a validation error before you discover it in a customer's ad.

Fonts

Fonts are the most common cause of "the render doesn't match the preview":

  • Licensing. Server-side rendering and webfont use are separately licensed in many foundry agreements. Check before you render at volume.
  • Availability. The renderer needs the actual font file. A font that exists on the designer's machine but not on the render worker falls back and every asset comes out wrong.
  • Deterministic fallback. Define the fallback chain explicitly so a missing font produces a predictable result.
  • Embedding. Print PDFs need embedded font subsets, which interacts with licensing again.

Print output

If your designs end up in print, find out early whether your renderer can produce what a print shop asks for: CMYK color, a bleed area, trim marks, spot colors, overprint and PDF/X files. Most browser-based rendering outputs RGB and gives you very little control over any of these. There are workarounds, such as rendering at a high DPI and converting with an ICC profile and Ghostscript, or handing the job to a print-specific renderer, but they have to be planned from the start. Print shops run preflight on every file they receive, and a missing bleed box or an RGB image is enough to send it straight back, so it is far cheaper to find that out on your first test PDF than on your first real order.

Assets live outside the file

Fonts, images and video are referenced by URL. You need to host them, keep those URLs valid, and accept that a design from two years ago renders differently if an image behind it moved. Snapshot assets alongside templates if you need reproducibility.

The schema is an API

Change the shape of your JSON and existing templates break. Version it, keep changes additive, and write migrations for anything else.

LLMs are unreliable at layout

LLMs are good with copy, variants and property edits, and a decent template gives them a lot to work with. In my experience, roughly 50–60% of the variations an LLM produces from a good template come out usable, so the practical approach is to generate several and filter. Keep the set of editable fields small and validate every output before it renders. Overlapping elements and text that drifts off the canvas are the usual failures, and a check catches them before anyone sees them.

The renderer has to match the editor

If the renderer uses different font handling or text layout logic than the canvas, output won't match the preview, and the mismatch will show up in the one asset a customer complains about. Use the same engine on both sides.

FAQ

Do designers have to write JSON?

No. Designers work on a canvas like they would in any design tool, and the editor writes the JSON for them. Lock the template so only the intended fields can be edited. A developer is only needed once, to wire the pipeline that fills the template with data for bulk generation.

After a script edits a design, is it still editable on the canvas?

Yes. There's one format. Programmatic output loads back into the editor like anything else, which is what makes "generate, then have a human fix it" possible.

Where do designs get stored?

Wherever you keep other data: a JSON column in your database, or an object in S3. A simple design is a few KB and a complex multi-page template can reach hundreds of KB. A design grows into megabytes when images or SVGs are embedded in the file as data instead of linked by URL. Uploading media to storage and referencing it keeps the file small, and either size is easy to store.

Is Git diffing genuinely useful, or just a nice line in a pitch?

Useful for template changes, where you can see that a font size or a color token changed. The diff of a wholesale layout rewrite is large and hard to read, so pair it with pictures: because a design renders straight from its JSON, a CI step can render the before and after versions and attach both images to the pull request. Designers and reviewers get a visual review next to the line-by-line one.

Does LLM generation actually work today?

For editing and variants, yes, reliably, especially with a schema in the prompt and validation on the output. For composing an original layout from a blank page, results are inconsistent. Treat it as a fast junior that needs a template and a reviewer.

Do I need the visual editor at all?

No. You can construct and render JSON entirely in code. The editor exists for the moments a human wants to see and adjust the thing.

Can it do video and animation?

Yes. A design can hold video clips, GIFs and audio, elements can carry animations, and each page has a duration, all in the same JSON. The same render step that produces a PNG can produce an MP4 or a GIF. Video takes longer to render than a still image, so allow for that in bulk runs.

What about accessibility?

Rendered output is a flat image or PDF, so accessibility lives in what surrounds it – alt text, captions, tagged PDFs. Because the source is text, you can generate alt text from the same data that filled the template.

Summary

  • A design can be fully described as JSON: objects and their properties.
  • Designers keep the canvas; the canvas just reads and writes that JSON.
  • Scripts, LLMs and validators work on the same file, at the same time, without exports.
  • A render call turns JSON into PNG, PDF or MP4.
  • Anything you can do to a text file – generate, diff, review, template, automate – you can now do to a design.

The Polotno SDK is built this way: one JSON schema, a visual editor on top of it, and a render API that turns the same JSON into files.

Skip the build, cut dev costs, launch faster

TRUSTED BY

100,000+

CREATORS

300+

BUSINESSES

ExpediaUnbounceLovePopPostGridPredis.ai