Business

How Zalando Used Data Quality Management To Scale Fashion Ecommerce

A

Admin

Content Team

November 21, 2025•
12 min read
How Zalando Used Data Quality Management To Scale Fashion Ecommerce

How Zalando Used Data Quality Management To Scale Fashion Ecommerce

In fashion ecommerce, design and branding get the spotlight.
Behind the scenes, data quality management decides whether the business actually works.

Zalando is one of Europe’s largest online fashion platforms, with more than 50 million active customers, over 15 billion euros in GMV, and around 10 billion euros in annual revenue. (E-commerce Germany News) At that scale, every small data issue, from a wrong size attribute to a delayed stock feed, multiplies across millions of shoppers and thousands of brands.

This article looks at how Zalando used data quality management to fix real problems in their business. We will walk through their move away from a single central data lake, their shift to data mesh, their AI powered product content pipeline, and how they share clean data with partners. At the end, we will pull out lessons for other ecommerce brands and add a light touch mention of Lestar.


Why Data Quality Management Matters In Ecommerce

Ecommerce data is messy by default.

A typical fashion marketplace deals with:

  • Millions of SKUs that change every season
  • Traffic from web, app, email, ads, and marketplaces
  • Systems for catalog, inventory, pricing, logistics, returns, and support

If data quality management is weak, you see classic symptoms:

  • Revenue numbers do not match between finance, marketing, and product
  • Filters on the website behave strangely because attributes are missing or mislabelled
  • Stock levels are wrong, which creates oversells and stockouts
  • Partners receive reports that they cannot reconcile with their own numbers

Data quality management is the set of policies, processes, and tools that keeps this data accurate, consistent, and timely. In your earlier pillar article, you already covered the foundations, such as dimensions like accuracy, completeness, timeliness, and uniqueness, and the idea of treating quality as an ongoing lifecycle.

Zalando’s story shows what happens when you apply those ideas at very large scale.


Zalando’s Scale And Data Challenges

Zalando is a public company based in Berlin that sells fashion and lifestyle products across more than 20 European markets. It serves over 51 million active customers and works with thousands of brands and partners. (E-commerce Germany News)

That scale creates tough data quality problems:

  • Many internal domains, for example assortment, pricing, logistics, marketing, customer experience
  • A large partner ecosystem that needs reliable performance, inventory, and marketing data
  • Rapid product and feature launches that constantly add new data sources and pipelines

In interviews, Zalando’s data leaders have described how their early centralised data lake and warehouse became a bottleneck. A single central team tried to serve many domains, which made it hard to keep ownership clear and to maintain high data quality as demand grew. (Hyperight)

This set the stage for a strategic change.


The Turning Point, From Central Data Lake To Data Mesh

Limits Of A Central Data Team

With a central lake and a central data team, many organisations hit similar problems:

  • Business teams wait in a long queue for new datasets and changes
  • The central team knows a bit about every domain, but does not deeply own any of them
  • When data is wrong, it is not obvious who is responsible for fixing it

Zalando faced this pattern as well. As the platform grew, it became harder for one team to own quality for everything, from fashion product data to logistics events. (Hyperight)

Adopting Data Mesh

Zalando is often cited as an early adopter of data mesh in real production. (Hyperight)

Data mesh is a way to decentralise data while keeping shared standards. In simple language:

  • Data is treated as a product
  • Domain teams own their own “data products” and are responsible for their quality
  • A central self service platform and governance layer provide common tooling and guardrails (Atlan)

For Zalando, this meant that instead of one central group owning all analytics tables, logistics teams own logistics data products, fashion buying teams own assortment data products, and so on, while a central data foundation team supports them with platform and standards. (Hyperight)

What This Means For Data Quality Management

This shift is very important for data quality management:

  • Quality becomes a shared responsibility, not an afterthought in a central lake
  • The people closest to the business logic define and maintain quality rules for their domain
  • The platform team provides tooling for validation, monitoring, cataloging, and access control

In other words, data mesh operationalises your framework for data quality management inside the organisation chart, not only inside documents.


Building The Foundation, Governance, Catalog, And Quality Platforms

A clever architecture is not enough. Zalando also invested in governance and tooling.

Strong Governance As An Enabler

Zalando’s data leaders describe governance as something that should enable AI and analytics rather than block it. (Zalando Engineering)

For data quality management, governance provides:

  • Shared policies for access, privacy, and retention
  • Standard ways to describe datasets and fields
  • A process for approving critical data products and changes

This gives domains freedom to build, while still protecting customers and partners.

Central Catalog And Metadata

A key element in Zalando’s platform is a catalog that tracks datasets, their owners, schemas, and lineage. (Hyperight)

For data quality management, a good catalog helps because:

  • Analysts can see what data products exist and how to use them
  • Lineage shows which upstream changes might break a downstream report
  • Ownership metadata makes it clear who should fix issues

It is hard to improve quality if nobody even knows what data exists.

A Dedicated Data Quality And Governance Platform

Zalando also invests in a dedicated data governance and data quality platform, and hires people specifically for running self service quality processes across many domains. (Hyperight)

This platform typically offers:

  • Standard checks that teams can apply, such as freshness, null rates, and range checks
  • Dashboards that summarise quality metrics across the organisation
  • Integrations with experimentation and monitoring so that quality issues are visible in context

So far, we have looked at structure. Next, we will zoom in on two very concrete data quality management stories at Zalando: product content and partner data sharing.


Fixing Product Data Quality With AI

The Pain, Manual Attribute Enrichment

Fashion ecommerce lives and dies on product content. Shoppers need to know size, fit, material, colour, and many more attributes to make a decision.

Zalando’s content team used to enrich these attributes manually during onboarding. According to their engineering blog, manual enrichment could take roughly a quarter of the overall production timeline and still produced error rates that needed improvement. (ZenML)

Poor quality attributes hurt:

  • Search and filters on the website feel wrong
  • Recommendations do not match taste or size
  • Returns increase because expectations do not match the product

The Content Creation Copilot

To improve this, Zalando built what they call a Content Creation Copilot.

The copilot:

  • Uses machine learning to read product descriptions and images
  • Suggests structured attributes, for example material or style, for human review
  • Is integrated into existing workflows so content creators stay in control (Zalando Engineering)

The aim is not to fully automate content, but to help humans move faster and make fewer errors.

Impact On Data Quality Management

This is a beautiful example of upstream data quality management:

  • Attributes become more complete and consistent across similar products
  • Manual workload drops, which frees time for reviewing edge cases instead of basic tags
  • Downstream systems that depend on these attributes, such as search and recommendation, get cleaner input data

Instead of just cleaning catalogs after the fact, Zalando improved the way data is created in the first place. That is one of the main best practices you already highlight in your general data quality management guide.


Sharing Trusted Data With Partners Using Delta Sharing

Zalando’s platform is not only about internal analytics. Its partner brands also need high quality data.

Before, Fragmented Exports And Heavy Manual Work

In a recent engineering post, Zalando’s team described how partner data used to flow through many channels, such as CSV exports, custom reports, SFTP transfers, and APIs, each with different schemas and limits. (Zalando Engineering)

Partners who wanted serious analysis ended up downloading and merging many files every month. Zalando reports that some partners were spending around 1.5 full time equivalents per month just to extract and consolidate data coming from Zalando before they could even start their analysis. (Zalando Engineering)

Even if the source data is decent, that much manual handling invites errors, version drift, and misunderstandings.

After, Live Data Sharing With Delta Sharing

To solve this, Zalando implemented Delta Sharing, an open protocol created by Databricks for secure data sharing across organisations. (Zalando Engineering)

In simple terms:

  • Data remains in Zalando’s platform
  • Partners receive live, read only access to well defined tables
  • Access is governed centrally using catalog and permission rules

Partners can then plug these tables into their own analytics tools without building complex import pipelines.

Why This Is A Data Quality Management Story

This change improves data quality management in several ways:

  • Partners see the same “source of truth” tables that Zalando uses internally, instead of exported copies that might drift
  • Schema, definitions, and lineage are documented in the catalog, which reduces confusion
  • Since data is not constantly copied and re shaped, there are fewer chances to introduce silent errors

It is still important to maintain internal quality rules for those shared tables, but the whole system is now designed around a single, governed version instead of hundreds of ad hoc exports.


Treating Data Quality As A First Class Metric In Experimentation

Data quality for experimentation is another area where Zalando has been very explicit.

In a series of posts about their experimentation platform, Zalando engineers explained how they automate and monitor data quality indicators, especially something called sample ratio mismatch, or SRM. (Zalando Engineering)

SRM happens when an A/B test that was supposed to split traffic evenly does not actually receive a fifty fifty split in the data. This is a strong signal that something is wrong in the experiment setup or in the tracking. (Website)

Zalando treats SRM and other indicators as core metrics in their platform, not as optional add ons. They automatically monitor experiments and surface quality problems so that teams can pause or rerun tests when the underlying data is not trustworthy. (Zalando Engineering)

This is another important lesson in data quality management:

  • Quality should be checked at the point where decisions are made, not only in isolated data platforms
  • For experimentation, that means treating quality indicators as seriously as business metrics like conversion or revenue uplift

Results And Lessons For Other Ecommerce Brands

Zalando does not publish every internal metric. However, their public stories show a consistent pattern.

From data mesh and governance side, they describe improved scalability and agility because domain teams can own their data products and move faster, while central platform teams provide reliable infrastructure. (Hyperight)

From the catalog and partner sharing side, they show how structured, governed data sharing removed a large amount of manual work for partners and created a more consistent view of performance. (Zalando Engineering)

From the product content side, they report that AI assisted attribute extraction reduces the time spent on manual enrichment, improves coverage of attributes, and makes catalog data more complete and correct, which feeds into better customer experience. (ZenML)

From experimentation, they highlight how automated SRM checks and other data quality indicators are essential for trustworthy A B tests. (Zalando Engineering)

For any ecommerce brand, even at much smaller scale, the key lessons are clear:

  1. Start from real pain, not just tools
    Zalando focused on concrete problems, such as manual partner data consolidation and slow product onboarding, before designing solutions.
  2. Give domains ownership, keep shared standards
    Putting ownership for data quality into domain teams, while supporting them with a central platform and governance, helps avoid bottlenecks.
  3. Fix data quality as early as possible
    Improving product attributes during onboarding has more impact than trying to clean catalogs later.
  4. Make data quality visible and measurable
    Track simple quality KPIs, such as freshness, completeness, and experiment health, and show them in the same places where people look at business metrics.
  5. Design for partners, not only for internal users
    High quality data that is easy to consume strengthens your brand relationships. Delta style sharing is one model, but the principle is universal.

A Soft Note On Tools

Zalando’s journey required serious investment in in-house engineering, as well as modern data tools for storage, processing, cataloging, and sharing.

If you are a growing ecommerce or digital business, you might not have a full Zalando sized platform team. You still need a place to:

  • Consolidate data from many channels and systems
  • Build and run transformation pipelines
  • Review data analytics and insights
  • Apply quality rules and monitor key datasets
  • Serve trusted data to BI tools, models, and partners

That is where platforms like Lestar come in. They give you the building blocks for modern data quality management, without forcing you to assemble everything from scratch.

If you care about data quality, check out Lestar. It is a platform that helps you consolidate data so that you can manage it easily.

Mandrill Tech, the development agency behind Lestar, also provides consulting and development services for data solution. If you would like help applying these ideas in your own business, get in touch with our team here.

Tags:Business IntelligenceData AnalyticsDashboardsFinance

Table of Contents

Related Articles

Business Intelligence and Analytics: What It Actually Means for SME Leaders in 2026
Business

Business Intelligence and Analytics: What It Actually Means for SME Leaders in 2026

8 min read
Executive Dashboard Software: What CEOs and CFOs Actually Need in 2026
Business

Executive Dashboard Software: What CEOs and CFOs Actually Need in 2026

9 min read
Ad-Hoc Reporting Software: How CEOs Get Answers Without Waiting for the Finance Team
Business

Ad-Hoc Reporting Software: How CEOs Get Answers Without Waiting for the Finance Team

11 min read
Mandrill Tech Logo

Empowering business growth through innovative technology and data-driven solutions. We help companies transform and scale in the digital age.

Solutions

Data Solution

AI/ML Solution

Services

Quick Links

© 2026 Mandrill Tech Sdn. Bhd. (201501005176 | 1130506-W). All rights reserved.