The industry pattern
We keep seeing the same problem at growing D2C brands: the catalog grows faster than the copy team. In our audits and PoCs, around 35 % of products typically have incomplete attributes. Copy varies in tone and length, and in many categories the filters only partly work.
Management and marketing watch conversion stall while traffic and assortment keep growing.
Architecture
We turn the process around: data first, copy second. This is the pipeline we test in the PoC:
- Extraction: the pipeline pulls structured data from manufacturer PDFs, web shops and product photos. An image recognition model adds tags for material, color and style.
- Normalization: it maps the data to the central attribute catalog and shows where information is missing.
- Generation: the structured data becomes copy in German and English, in the brand's voice. The brand manual serves as context for the model.
- Approval: a person approves copy for top SKUs. The long tail is published automatically and spot-checked.
Model calculation
The projection builds on the PoC A/B tests against existing copy and on the assumptions further down:
- +18 % conversion rate on product detail pages (PoC A/B test against the previous copy)
- Goal: complete attributes for all active SKUs (the pipeline flags every gap)
- −72 % time spent by the content team (projected from typical editing times per SKU)
The PoC A/B tests show an interesting side effect: AI copy does especially well in categories where the brand traditionally has little in-house expertise. The reason: it comes strictly from the data, not from gut feeling.
Methodology in detail
To keep the +18 % traceable, here is the math on a generic D2C example. It is not a client measurement, but the logic we use to work through use cases like this.
Assumptions
- Starting conversion on product detail pages: we assume 2.0–2.5 %. This is an assumption for this example, not a measured industry figure.
- Assortment distribution: typically Pareto. The top 10 % of SKUs bring in 60–75 % of revenue, the long tail the rest. PIM upkeep tends to fall behind precisely in the long tail.
- Manual editing time per SKU: 10–15 minutes for writing, review and PIM upkeep by an in-house copywriter. Generating copy automatically with no review at all isn't comparable: the effect on conversion disappears.
- Attribute completeness before the audit: typically 60–70 % in the long tail. In more than half of the categories, the data structure holds back conversion through filters.
- Systems: with a PIM such as Akeneo or Pimcore, the pipeline talks to its API; with a feed tool like Productsup, it uses the feed. Without a PIM, it writes directly to the shop, for example Shopify, JTL or Magento.
The math
- Time spent: 12 min/SKU × 50 SKUs/day = 600 min/day. With the AI pipeline and spot-check review: 2 min/SKU × 50 = 100 min/day. That's −83 % gross. Once we add the extra four-eyes review for top SKUs, the projection lands conservatively at −72 % per editing task.
- Conversion: In the prototype A/B test against the previous copy, we measure +14 % to +22 % per category, depending on the share of hero range and long tail. Weighted across a typical mix of product detail pages, that comes to +18 % on overall conversion. The effect is strongest in long-tail categories whose copy was weak in tone before, and weakest in the hero range, where established copy is already in place.
What moves the numbers
- If revenue leans more heavily on the hero range (top 5 % = 80 % of revenue), the conversion effect drops to +10 % to +12 %. Hero copy is usually already strong, and the long tail carries little weight in the revenue model.
- If attribute completeness in the PIM is below 50 %, data extraction takes longer: implementation needs 1–2 weeks more. The effect itself holds; it just shows up later.
The key takeaway
AI copy without clean data ends up generic. What makes the difference is extracting and normalizing the data; generating the copy is the easy part. Every content automation PoC we have built so far has shown this.