Latest posts Visit blog

An online shop lives on images, yet very few shops treat images as master data. The median mobile page contains 13 image elements (HTTP Archive Web Almanac 2024), and 45 percent of all image elements carry no alternative text (HTTP Archive Web Almanac 2024). Both are symptoms of the same gap: images are stored, not managed. This article describes a media structure that protects the master file, carries the usage rights with the image and generates every rendition from a single source – for the shop, the PIM and the catalogue. It continues what we described about ownership of your own product data, only for the files instead of the fields.

Images are master data, not attachments

In many shops the image store grows on the side. A photographer delivers an archive, somebody uploads the files into the shop media area, names them after the product and considers the job done. Six months later, every second image lacks an answer to three questions: who took it, how long may it be used, and where is the version a new size can be generated from. The images are present, the information about them is missing. That is exactly what separates a store from a management system: a store knows files, a management system knows files and their state.

The comparison with other master data helps. A product record carries a number, a name, a price, a tax rate and a supplier, and hardly anybody would keep those details only in the buyer's head. An image likewise carries an author, a type of usage right, an expiry date, releases for depicted people or objects, alternative text and a rendition profile. Anyone who reviews the inventory once a year already knows the pattern from stocktaking in an online shop: what is not recorded appears in no report and is missing all the same as soon as somebody asks for it.

An image without a documented usage right is not an asset but an open item.

XICTRON, e-commerce consulting

How small the individual file ends up being surprises many people: the median is 12 KB per image (HTTP Archive Web Almanac 2024). That is not an argument against care but one in favour of it. Small files do not appear by themselves; they exist because a pipeline generates several small versions from one large master. Without the master only the small file remains, and every new requirement – a larger zoom image, a marketplace format with a fixed edge length, an extract for the printed catalogue – ends in a reshoot.

Two questions before the first upload

First: which file is the master, and does it sit outside the delivery path? Second: which alternative text belongs to this subject? Answering both questions at creation time saves the rework that otherwise lands in the accessibility project – where it is considerably more expensive.

One master, many renditions

At the start there is the master, exactly one per subject. It sits at the highest available resolution, lossless or at very high quality, outside the shop system and outside the directory that is delivered from. Everything else is a rendition: the listing view, the detail image, the zoom image, the square version for a marketplace, the small image in the order confirmation. Renditions are the output of a recipe, not files with a history of their own. Anyone who keeps that distinction in mind and in the directory tree can regenerate any rendition at any time without having to ask anybody.

Which formats are produced depends on what browsers can read, not on habit. Across the open web the image formats on mobile still lean clearly towards the older methods: JPEG 32.4 percent, PNG 28.4 percent, WebP 12 percent and AVIF 1.0 percent (HTTP Archive Web Almanac 2024). The difference in efficiency is measurable: the median is 1.3 bits per pixel for WebP and 1.4 for AVIF, against 2.0 for JPEG, 3.8 for PNG and 6.7 for GIF (HTTP Archive Web Almanac 2024). Which format carries which job is covered in more depth by the article on format choice for shop images.

  • Master – highest resolution, lossless, with complete IPTC and XMP fields, outside the delivery path
  • Detail – long edge around 1600 pixels, for product page and zoom, in AVIF and WebP with JPEG as the fallback
  • Listing – long edge around 640 pixels, for overviews, search results and recommendations
  • Icon – long edge around 160 pixels, for basket, wish list and order confirmation
  • Catalogue – version to the recipient's specification, usually with a fixed edge length and a defined background

The renditions reach the browser through srcset and sizes. Neither attribute is decoration: a quarter of all desktop pages that use width descriptors in srcset load 180 KB or more of wasted image data according to the HTTP Archive estimate, because the sizes value does not match the actual layout width (HTTP Archive Web Almanac 2024). The rendition is therefore half the work and the selection rule in the markup is the other half. How both connect with loading behaviour and priority is shown in the article on adaptive image loading.

Renditions belong in the build, not in someone's hands

As soon as a size is produced by hand, a file without a recipe comes into existence. Generate renditions in the pipeline, log the profile and the source file, and let the run abort when a master is missing. That is a piece of development work which pays for itself several times over the lifetime of a shop.

Rights on the image: what Section 31 UrhG requires

An image is a work under German copyright law, and anybody who uses it needs a right of use for it. Section 31 paragraph 1 UrhG states that such a right may be granted as a non-exclusive or an exclusive right and that it may be limited in territory, time or content. For media management that means every image carries at least four pieces of information: who it came from, which type of right was granted, for which territory and period it applies, and for which types of use.

The difference between a non-exclusive and an exclusive right is tangible in daily work. Under Section 31 paragraph 2 UrhG the non-exclusive right entitles the holder to use the work in the permitted manner without excluding use by others – so the same subject may sit in another shop at the same time. Under Section 31 paragraph 3 UrhG the exclusive right excludes all other persons and allows the holder to grant rights of use themselves. The same care that applies to self-hosted web fonts applies to images: the licence is part of the inventory, not part of somebody's memory.

AspectNon-exclusive right of useExclusive right of use
Use by the holderin the permitted mannerin the permitted manner
Use by othersnot excluded (Sec. 31 para. 2 UrhG)excluded (Sec. 31 para. 3 UrhG)
Granting rights onwardas a rule not provided forexpressly provided for (Sec. 31 para. 3 UrhG)
Possible limitationterritory, time, content (Sec. 31 para. 1 UrhG)territory, time, content (Sec. 31 para. 1 UrhG)
What applies when unclearpurpose of the contract (Sec. 31 para. 5 UrhG)purpose of the contract (Sec. 31 para. 5 UrhG)
Field in the media archiveRights Usage TermsRights Usage Terms and Licensor

When it stays open which types of use were meant, the purpose-of-transfer rule applies. Section 31 paragraph 5 UrhG provides that the scope follows the purpose of the contract as taken as a basis by both partners if the types of use are not expressly designated individually. In practice that is not a free pass but a risk: whoever commissioned an image for the product catalogue has as a rule not thereby ordered its use in a paid advertising campaign. The consequence for the archive is plain – the agreed purpose belongs in a field on the image, not in an email.

Not legal advice, but a filing question

This section does not replace legal advice in an individual case. It describes which fields a media management system should hold so that the legal review is possible at all. In our e-commerce projects we define these fields together with purchasing before the first image inventory is taken over.

IPTC fields that carry weight in the shop

For the technical side there is an established framework: the IPTC Photo Metadata Standard. Version 2025.1 adds several properties and raises the Extension schema to version 1.9 (IPTC Photo Metadata Standard 2025.1). The fields are written into the file as XMP and travel with the image – from the place it was taken through the media archive into the shop. The standard itself points out that rights-related values may be affected by the laws and other regulations of the region in which the image is used. The fields are therefore a place for statements and not a decision about their effect.

The practical advantage of embedded fields lies in their independence from any system. A database column disappears with the database, an XMP block stays in the file. If an inventory is later moved into another system, author, rights notice and alternative text are still there. The only requirement is that the rendition pipeline carries the fields along: stripping all metadata while resizing in order to save bytes removes exactly the statements that will be needed later. The rule is therefore: complete on the master, reduced in the rendition to rights notice, author and alternative text.

  • Copyright Notice (dc:rights) – the rights notice naming the current owner of the copyright
  • Rights Usage Terms (xmpRights:UsageTerms) – the licensing conditions in free text, the place for territory, time and purpose
  • Web Statement of Rights (xmpRights:WebStatement) – an address where the rights situation is documented for reading
  • Licensor (plus:Licensor) – who is to be contacted for a licence; the standard provides for up to 3 entries (IPTC Photo Metadata Standard 2025.1)
  • Alt Text (Accessibility) (Iptc4xmpCore:AltTextAccessibility) – the short description of the purpose and meaning of the subject
  • Digital Source Type (Iptc4xmpExt:DigitalSourceType) – the origin of the file, for example a capture, a reproduction or created graphics

The alternative text deserves attention of its own. The standard describes the field as a short description of the purpose and meaning of an image, read or displayed by assistive technology when images are switched off in the browser, and limits it to 250 characters (IPTC Photo Metadata Standard 2025.1). That defines the task clearly: no file name, no keyword list, no place for search terms. How the text can be generated and reviewed in a structured way for large inventories is described in the article on alternative text for product images.

What reaches the visitor: format, size, alternative text

Between the media archive and the visitor there are three decisions: which format is delivered, which size, and which text describes the image when it does not arrive or is not seen. All three can be derived from the media inventory as long as it holds the matching values. If one of them is missing, manual work appears exactly where it is particularly expensive – in day-to-day operation, under time pressure, usually done by somebody who does not know the subject.

Format by capability

Modern formats deliver the same visual result with fewer bytes. The median is 1.3 bits per pixel for WebP and 1.4 for AVIF against 2.0 for JPEG (HTTP Archive Web Almanac 2024). The pipeline produces both versions and leaves the choice to the browser.

Size from one source

Every size comes from the master and carries its profile in the file name. That keeps it traceable which file belongs to which recipe, and a changed specification regenerates the inventory instead of adding to it.

Alternative text on the image

The text sits in the media archive on the master, not in the page template. That way it applies in every view, in every language and in every channel the image is handed over to.

For cases where the browser is not supposed to choose on its own there is the picture element. According to the HTTP Archive survey it is used far less than srcset (HTTP Archive Web Almanac 2024). For shops it makes sense wherever a subject should be cropped differently in a narrow viewport than in a wide one – a landscape campaign image, say, that appears square on the phone. The crop then belongs in the pipeline as a profile of its own and not in the hands of the editorial team.

Handover to catalogue and marketplace

The handover to catalogues shows whether the media structure holds up. The BMEcat exchange format provides the MIME_INFO block for this, describing product images, data sheets and further documents. What is transferred is not the file itself but the relative path or the address in MIME_SOURCE, relative to a base directory MIME_ROOT from the document header (BMEcat 2005.2). Anyone setting up the exchange as a whole will find the classification side in the article on BMEcat and ETIM.

Handing images to a catalogue does not hand over the file but the route to the file and the meaning it is supposed to have there.

XICTRON, integration development

Two fields deserve particular attention here. MIME_PURPOSE describes the intended use of the MIME document in the target system and knows defined values such as normal, detail, icon and logo (BMEcat 2005.2). MIME_ALT takes the alternative text in case the file cannot be displayed in the target system and is limited to 80 characters (BMEcat 2005.2). Both fields can only be filled if the values exist on the master – and both are the reason why a well-kept alternative text in the media archive pays off twice.

In business with corporate customers there is the added point that the same product needs different images per recipient: a dealer wants the cut-out subject, a marketplace the square one, a key account the view with dimensions. As long as each of these versions is a rendition with a profile of its own, the effort stays manageable. Once they become individual files, the inventory grows faster than the overview. How such requirements can be modelled cleanly in a B2B shop depends less on the shop system than on whether the media structure knows profiles.

Naming, storage, pipeline

Naming is inconspicuous and at the same time effective. A file name should carry the product number, the subject number and the profile, in that order, in lower case, without spaces and without special characters. The name is an identifier, not prose – ASCII is right here, while every text a visitor reads carries proper characters. An example: 12345-01-detail-1600.avif says at a glance which product the image belongs to, which subject it shows and which profile it came from.

catalogue-images.xml
<HEADER>
  <CATALOG>
    <MIME_ROOT>https://medien.example.com/katalog/</MIME_ROOT>
  </CATALOG>
</HEADER>

<MIME_INFO>
  <MIME>
    <MIME_TYPE>image/jpeg</MIME_TYPE>
    <MIME_SOURCE>12345/12345-01-detail-1600.jpg</MIME_SOURCE>
    <MIME_DESCR>Solid oak chair, front view</MIME_DESCR>
    <MIME_ALT>Oak chair with an upholstered seat</MIME_ALT>
    <MIME_PURPOSE>detail</MIME_PURPOSE>
    <MIME_ORDER>1</MIME_ORDER>
  </MIME>
  <MIME>
    <MIME_TYPE>image/jpeg</MIME_TYPE>
    <MIME_SOURCE>12345/12345-01-liste-640.jpg</MIME_SOURCE>
    <MIME_ALT>Oak chair with an upholstered seat</MIME_ALT>
    <MIME_PURPOSE>normal</MIME_PURPOSE>
    <MIME_ORDER>2</MIME_ORDER>
  </MIME>
</MIME_INFO>

The base directory sits in the catalogue header, the references to it sit in the product. The recipient puts both together and fetches the files by the agreed route. The separation matters: the path describes the storage, MIME_PURPOSE describes the role in the target system. Deriving both from the same profile is the difference between a handover that can be repeated and one that has to be assembled again every time.

Terminal
$ bin/medien-pruefen --quelle medien/original
alternative text missing: 12345/12345-02-master.tif
rights usage terms missing: 67890/67890-01-master.tif
$ bin/medien-ableiten --profil shop --quelle medien/original --ziel medien/ableitungen
detail (AVIF, WebP, JPEG), listing (AVIF, WebP), icon (JPEG)
$ bin/medien-uebergabe --katalog bmecat --pruefen
MIME_SOURCE and MIME_ALT set for all entries

The output shows the principle: check first, then generate, then hand over. Every step can abort, and an abort is cheaper than a catalogue with empty image fields. It matters that the check runs before the generation: a missing alternative text on the master otherwise propagates into every rendition and every handover, and the correction then touches dozens of files instead of one.

  • Masters sit outside the delivery path and are backed up
  • Every master carries a rights notice, licensing conditions and alternative text
  • Every profile is stored as a recipe, not as a collection of individual files
  • The pipeline aborts when a mandatory field is missing
  • The rendition keeps rights notice, author and alternative text
  • The file name contains product number, subject and profile

Operations: roles, checks, records

For the structure to hold, it needs clear responsibilities. Who may replace a master, who may only regenerate a rendition, who may change the rights fields? In practice a small number of clearly cut roles works better than general access for everybody involved – the same consideration we described for roles and permissions in the shop backend. A replaced master without a log entry is a frequent reason why a rendition suddenly shows a different subject than yesterday's order confirmation.

  • Expiring licences: monthly report on images whose usage period is ending
  • Completeness: share of masters with a filled rights notice and alternative text
  • Orphans: renditions without a matching master
  • Duplicates: identical subjects under several product numbers
  • Handover: share of catalogue entries with MIME_SOURCE and MIME_ALT set
  • Log: who replaced a master or changed a rights field, and when

How we approach it

We start by taking stock: how many files exist, how many of them are masters, which fields are filled, which renditions exist without a recipe. From that comes a profile plan defining which sizes and formats the shop, the catalogue and the marketplaces need. We then connect the generation to the existing system – through the PIM interface, through the shop build or through both, depending on where the data is maintained anyway.

We introduce the rights part together with purchasing, because that is where the contracts sit. We define which fields are mandatory, how an expiry date is recorded and what happens when an image passes the end of its usage period. The effort as a rule lies in the first pass of data capture; after that it is a routine of a few minutes per new subject. What comes out of it is not a tool but a state: for every image in the shop it can be said where it came from, how long it may stay and how it can be regenerated.

Sources and legal basis

The figures on image use across the web come from the Media chapter of the HTTP Archive Web Almanac 2024. The metadata fields follow the IPTC Photo Metadata Standard 2025.1. The statements on rights of use refer to Section 31 UrhG in the version published on the Gesetze im Internet portal. The catalogue fields come from the BMEcat 2005.2 specification. Measured values from crawls shift with every survey; the figures quoted reflect the state of the respective publication.

A folder knows files but no state. As soon as somebody asks which file is the master, until when a subject may be used or which alternative text applies, the answer has to come from the inventory and not from memory. A media management system therefore keeps the same mandatory fields per image as a product record: author, rights notice, licensing conditions, alternative text and rendition profile.

At least four values: who the image came from, whether a non-exclusive or an exclusive right of use was granted, for which territory and period it applies and for which types of use. Section 31 paragraph 1 UrhG names exactly these limitations, and Section 31 paragraph 5 UrhG makes clear that without an express designation of the types of use the purpose of the contract decides. Keeping those fields lets you answer a query in minutes.

Copyright Notice for the rights notice, Rights Usage Terms for the conditions, Web Statement of Rights for the documented version, Licensor for the contact and Alt Text for the description of the subject. All five are described in the IPTC Photo Metadata Standard 2025.1 and are written into the file as XMP, so they survive a change of system.

As a rule three to five profiles are enough: detail, listing, icon and, where needed, a catalogue or marketplace version. What matters is less the number than the origin: every rendition comes from the master following a stored recipe. If a specification changes, the inventory is regenerated instead of individual files being patched.

The purpose and meaning of the subject in a few words. The IPTC Photo Metadata Standard 2025.1 limits the field to 250 characters. File names, product numbers and keyword lists do not belong there; they help neither assistive technology nor comprehension. For catalogue delivery there is the added point that BMEcat 2005.2 limits the alternative text to 80 characters – so a good wording is short anyway.

Through the MIME_INFO block: every document gets a type, a source, a description, alternative text, an intended use and an order. The source is a relative path or an address, relative to the base directory in the catalogue header. If the media structure knows profiles, the handover is a mapping from profile to intended use – a route we set up as part of a project discussion.