Abstract

Billiger.de Products is an entity matching benchmark requiring matchers to identify product offers that refer to the same real-world product. It contains 13,730 offers describing 2,168 distinct products, collected from the German price-comparison platform billiger.de. The benchmark provides both the original German offers and machine-translated English versions of all training, validation, and test sets, enabling controlled analyses of how language affects entity matching performance. Following the multidimensional design of the WDC Products benchmark, Billiger.de Products provides dataset variants that systematically vary the size of the training set, the proportion of difficult-to-match corner cases, and the proportion of products that are not observed during training. These variants support fine-grained evaluations of data efficiency, robustness to challenging examples, and generalization to previously unseen products. We validate the benchmark using six supervised entity matching methods and GPT-5.2 in a zero-shot setting.

Introduction

Entity matching is the task of determining whether two records refer to the same real-world entity, for example whether two product offers describe the same product. Widely used entity matching benchmarks, such as Abt-Buy, Amazon-Google, or WDC Products, are monolingual: each contains records in only one language, typically English. However, many real-world applications must process and integrate data in multiple languages. For example, a European shopper searching for the lowest price of a product may not care whether the corresponding offer is published by a German, French, or Italian online retailer. A price-comparison platform serving such users must therefore be able to match product offers both within and across languages. Despite the practical importance of this setting, benchmarks for systematically evaluating entity matching systems on multilingual and cross-lingual data remain scarce.

Billiger.de Products is a step in this direction. The benchmark uses product data from the German price comparison platform billiger.de and is, to the best of our knowledge, the largest entity matching benchmark for German-language product data. In addition to the German original, we provide a machine-translated English version of every record. Splits, labels, and identifiers are identical in both versions, so the benchmark can be used as a German benchmark, as an English benchmark, or as a pair of parallel benchmarks for measuring how the language affects matching performance. Since every German record can also be paired with the English translation of its counterpart, the benchmark is also a starting point for cross-lingual matching experiments.

The benchmark follows the multi-dimensional design of WDC Products and controls task difficulty along three dimensions: the size of the training set, the fraction of hard-to-match corner cases, and the fraction of entities unseen during training. The sections below describe the benchmark in detail and report baseline results for six supervised matchers and for GPT-5.2 in a zero-shot setting on both language versions, as well as a cross-language experiment that evaluates matchers trained on German pairs on mixed German-English record pairs.

The Benchmark

The benchmark use product offers originating from the German price comparison portal billiger.de. The benchmark consists of 13,730 product offers describing 2,168 distinct products. Every offer carries the attributes name, description, brand, and price, with the attributes name and price having a value density of 100%, while brand and description have a density above 99 % . Identifiers such as EAN or manufacturer part numbers are deleted from the offers so that matchers need to rely on the textual description of a product rather than identifier lookup.

Following the design of WDC Products, task difficulty is controlled along three dimensions, resulting in 9 training sets, 27 validation sets, and 9 test sets per language:

  • Training set size: before leakage cleaning, the small, medium, and large training sets contained 2,500, 6,000, and 26,888–27,393 pairs. The released sets now contain 2,461–2,475, 5,897–5,913, and 26,571–27,112 pairs, respectively. The smaller sets are subsets of the larger ones, and all three are built from the same 500 product groups.
  • Corner cases: groups of four highly similar but distinct products generate hard positives and hard negatives. Datasets contain 20 %, 50 %, or 80 % corner-case product groups.
  • Unseen entities: in the Seen, Half-Seen, and Unseen test sets, 0 %, 50 %, or 100 % of the to-be-matched products do not occur in the training data. The Half-Seen and Unseen test sets evaluate the ability of a matcher to generalize to new and thus unseen products.
  • Two languages: every record exists in the German original and in an English machine translation. Splits, labels, and identifiers are identical, so performance differences are attributable to the language of the records and the translation process.

Product Categories

The products belong to thirteen different product categories. The figure below shows the amount of product offers that belong to each category. Clothing, furniture, and electronics account for the largest shares.

Clothing & Accessories25.7%
Furniture & Living24.4%
Electronics & Computers15.7%
Tools & DIY9.4%
Sports & Leisure6.4%
Toys & Baby5.6%
Automotive & Motorcycle4.6%
Cosmetics & Drugstore3.6%
Health & Care1.8%
Office Supplies1.0%
Food & Beverages0.9%
Books, Movies & Music0.8%
Pet Supplies0.4%

Record Pairs and Labels

The record pairs in the training, validation, and test sets are either labeled as matches or non-matches. The labels are derived from a grouping (clustering) of offers that refer to the same product which was provided to us by billiger.de. Positive pairs (matches) are directly taken from these groups, negative pairs (non-matches) are generated by combining offers from one group with similar offers from other groups, resulting in hard negatives, with one randomly choosen negative per group. The amount of hard negatives per product grows with the training set size.

The test sets contain roughly 4,440 pairs each with a class imbalance of 1:10 to 1:12 between matches and non-matches. All nine test sets went through a two-round manual verification process involving two human annotators, who checked the pairs using a strict definition of what constitutes a match: product variants that differ in size, color, configuration, capacity, or package quantity were labeled as non-matches. The audit corrected pairs in which similar but distinct product variants had been grouped together.

Training and Validation Sets

Corner cases Size Training Validation
Pos. Neg. Total Pos. Neg. Total
20 % Small 485 1,979 2,464 490 1,984 2,474
Medium 1,457 4,456 5,913 490 2,979 3,469
Large 13,100 14,012 27,112 490 3,972 4,462
50 % Small 482 1,979 2,461 487 1,978 2,465
Medium 1,464 4,441 5,905 487 2,966 3,453
Large 13,024 14,000 27,024 487 3,963 4,450
80 % Small 492 1,983 2,475 490 1,976 2,466
Medium 1,462 4,435 5,897 490 2,958 3,448
Large 12,747 13,824 26,571 490 3,946 4,436

Test Sets

Corner cases Unseen Pos. Neg. Total
20 % 0 % 407 4,047 4,454
50 % 352 4,077 4,429
100 % 360 4,089 4,449
50 % 0 % 364 4,063 4,427
50 % 346 4,072 4,418
100 % 369 4,089 4,458
80 % 0 % 342 4,098 4,440
50 % 361 4,076 4,437
100 % 348 4,099 4,447

Example Pairs

The following matching and non-matching pairs illustrate the variety of product categories and comparison difficulties contained in the benchmark.

A real corner case from the benchmark: two wardrobe offers that agree in brand, model name, color, and almost the entire description, yet refer to different product variants (168 cm vs. 85 cm width). Under the strict identity definition of the benchmark, this pair is a non-match. Matchers that rely on surface similarity are reliably fooled by such pairs.

Record A id 2070627808

brand RAUCH

name rauch Kleiderschrank »Rasa« beige 168 cm x 188 cm x 52 cm

desc …Anzahl Türen, 4 St., Art Türen, Spiegeltüren Drehtüren, Maßangaben Breite, 168 cm, Tiefe, 52 cm, Höhe, 188 cm…

price 409.99

Record B id 2056478070

brand RAUCH

name rauch Kleiderschrank »Rasa« beige 85 cm x 188 cm x 52 cm

desc …Anzahl Türen, 2 St., Art Türen, Drehtüren, Maßangaben Breite, 85 cm, Tiefe, 52 cm, Höhe, 188 cm…

price 249.99

 Non-match: different width and door variants of the same product line


Two offers for the same De'Longhi coffee machine use different spelling, word order, and levels of detail, but share the identifying model number ECAM 650.55 MS. This pair is a match.

Record A id 1497975779

brand De' Longhi

name Delonghi Ecam 650.55 Ms Primadonna Elite Kaffeevollautomat

price 1349.78

Record B id 61480810216

brand De'Longhi

name Superautomatische Kaffeemaschine DeLonghi ECAM65055MS 1450 W Grau 2 L

price 1484.86

 Match: identical ECAM 650.55 MS model


These two METABO power tools share a brand and technical vocabulary, but the model numbers and product types conflict: one is a mitre saw, the other a rotary hammer. The hard-negative pair is therefore a non-match.

Record A id 191526663714

brand METABO

name Metabo KGS 254 M Elektro-Kappsäge

price 323.25

Record B id 59701070745

brand METABO

name Metabo KHE 3250 Elektro-Bohr-/Meißelhammer inkl. Koffer

price 356.23

 Non-match: different METABO models and tool types

Baseline Results

Six supervised matchers spanning four methodological generations (word co-occurrence, feature-based machine learning, fine-tuned transformers, contrastive and graph-based architectures) were evaluated across all 27 configurations of both language versions, plus GPT-5.2 in a strict zero-shot setting. The table shows F1 scores on the German benchmark for all three training sizes. GPT-5.2 uses no training data, and results are shown for both the Simple and the Rule-Guided prompt.

Corner cases Training size Test set WordCooc Magellan RoBERTa R-SupCon HierGAT Ditto GPT-5.2 Simple GPT-5.2 Rule-Guided
20 % Small Seen 57.40 55.85 72.11 63.01 70.69 77.31 85.34 79.94
Half-Seen 45.70 50.71 61.36 48.34 64.50 68.52 81.81 84.91
Unseen 32.59 48.62 55.11 41.68 57.48 59.16 81.58 84.67
Medium Seen 67.36 56.98 81.25 75.12 85.17 87.51 85.34 79.94
Half-Seen 53.51 52.90 65.68 54.26 75.36 75.29 81.81 84.91
Unseen 37.45 49.27 58.01 44.08 63.40 64.26 81.58 84.67
Large Seen 72.12 55.73 84.95 82.33 89.41 91.76 85.34 79.94
Half-Seen 55.98 49.78 69.57 58.28 79.43 78.54 81.81 84.91
Unseen 38.11 47.15 59.24 46.65 65.80 65.24 81.58 84.67
50 % Small Seen 43.42 40.37 34.25 55.85 37.51 59.90 80.62 83.93
Half-Seen 37.98 39.66 29.12 42.63 37.83 55.66 72.44 80.65
Unseen 27.38 38.24 28.83 32.89 35.02 49.83 74.86 77.78
Medium Seen 55.84 43.78 70.83 66.25 70.08 78.08 80.62 83.93
Half-Seen 48.32 44.86 57.37 47.84 64.68 67.77 72.44 80.65
Unseen 29.91 42.11 52.68 36.14 55.96 59.33 74.86 77.78
Large Seen 64.78 42.91 76.31 73.48 79.90 84.91 80.62 83.93
Half-Seen 49.65 42.87 61.87 50.83 71.09 71.75 72.44 80.65
Unseen 31.08 40.95 53.31 38.38 62.16 59.33 74.86 77.78
80 % Small Seen 38.72 33.71 23.99 53.85 30.10 61.86 72.79 79.32
Half-Seen 36.40 34.44 24.56 43.12 31.79 54.36 68.03 79.82
Unseen 24.42 32.79 23.94 30.77 30.33 49.38 66.19 80.63
Medium Seen 49.72 35.44 43.16 63.20 65.10 74.11 72.79 79.32
Half-Seen 44.87 37.39 37.70 47.14 60.66 64.91 68.03 79.82
Unseen 25.68 34.50 32.83 31.69 52.61 54.16 66.19 80.63
Large Seen 51.36 35.98 70.71 65.65 73.39 72.69 72.79 79.32
Half-Seen 43.16 36.42 60.56 47.27 65.40 64.10 68.03 79.82
Unseen 22.01 33.63 48.86 30.86 54.01 51.69 66.19 80.63
GPT-5.2 prompts

The German benchmark used German-language versions with the same instructions and record attributes.

Simple prompt

Do these two product descriptions refer to the same real-world product?
Answer with Yes or No only.
Product 1: Brand: <brand> Name: <name> Description: <description> Price: <price> Euro
Product 2: Brand: <brand> Name: <name> Description: <description> Price: <price> Euro

Rule-Guided prompt

You are an expert product matcher. Your task is to decide whether two product records refer to the EXACT same product (same GTIN/SKU). Analyze the provided records carefully and return your decision as strict Yes or No.

CRITICAL: Product variants are NOT matches. Different sizes, colors, configurations, or package quantities are DIFFERENT products with different GTINs.

Guidelines:
Yes ONLY if both records refer to the exact same product that would have the same GTIN or barcode.
No if they are variants of the same product line such as different size, color, capacity, or configuration.
No if core identifying attributes conflict, including model numbers, dimensions, capacity, color, or configuration.
Missing attributes alone are NOT a conflict.
A match requires positive evidence of equivalence.
Respond ONLY with Yes or No.
Product 1: Brand: <brand> Name: <name> Description: <description> Price: <price> Euro
Product 2: Brand: <brand> Name: <name> Description: <description> Price: <price> Euro

The baseline results underline the difficulty of the benchmark. The traditional feature-based approaches WordCooc and Magellan average only 43.89 and 42.85 F1 across the 27 German variants, which is too low for practical use on this benchmark. Neural matchers such as Ditto reach up to 91.76 F1 in favorable configurations, but, as reported for WDC Products, they remain strongly affected by unseen entities. Averaged across all corner-case ratios and training sizes, Ditto loses 19.53 F1 from Seen to Unseen test sets. Even zero-shot GPT-5.2 does not solve the benchmark: with the stronger Rule-Guided prompt, its scores remain between 77.78 and 84.91 F1. The benchmark therefore remains challenging for traditional matchers, supervised neural models, and the evaluated LLM.

The three difficulty dimensions separate the benchmark variants as intended. F1 generally increases with the training set size, decreases as the corner-case ratio grows from 20 % to 80 %, and decreases from the Seen to the Unseen test sets. Individual medium-to-large results can deviate from the overall training-size trend. Comparing the Large, 20 % corner-case, Seen row with the Small, 80 % corner-case, Unseen row, the losses range from 22.94 F1 for Magellan (55.73 to 32.79) to 61.01 F1 for RoBERTa (84.95 to 23.94). In the Large, 20 % corner-case rows, the Seen-to-Unseen drops for R-SupCon, Ditto, and RoBERTa are 35.68, 26.52, and 25.71 F1, compared with 26.59, 9.60, and 9.16 F1 reported for WDC Products. The broader range of product categories may contribute to the larger drops, but this comparison does not establish their cause.

Ditto and HierGAT are the strongest supervised systems and reach 91.76 and 89.41 F1 on the easiest German variant. Zero-shot GPT-5.2 behaves differently. Under the Rule-Guided prompt its F1 varies by 7.13 points across all difficulty settings (19.15 under the Simple prompt), and it leads on every Unseen test set, where supervised models lose the training signal they depend on. Supervised matchers overtake GPT-5.2 only on Seen test sets with medium or large training data. A caveat applies to interpreting the GPT-5.2 results on the Seen/Unseen dimension. GPT-5.2 was trained on a large, undisclosed corpus, and all products in the benchmark were publicly available online. Some or all individual product records may therefore have appeared in its pre-training data. Here, Seen and Unseen only indicate whether the products occur in the benchmark's supervised training sets, not whether GPT-5.2 encountered them during pre-training.

Performance by Product Category

The difficulty of matching decisions depends on the product category. To avoid selecting a different representative split for each model, the table now averages same-category F1 over all eligible German benchmark configurations. Electronics & Computers ranks among the two easiest categories for all seven models. Clothing & Accessories is the most difficult category for Magellan, RoBERTa, R-SupCon, and GPT-5.2 and the second most difficult for HierGAT and Ditto. Furniture & Living ranks among the two most difficult categories for five models. Distinctive model numbers and technical specifications may contribute to the stronger results for electronics, while variant-heavy, near-duplicate descriptions may contribute to the lower results for clothing and furniture.

Category WordCooc Magellan RoBERTa R-SupCon HierGAT Ditto GPT-5.2
Automotive & Motorcycle 39.34 50.44 53.31 49.10 79.43 85.64 92.74
Electronics & Computers 54.20 57.83 71.53 67.49 78.06 85.05 89.73
Clothing & Accessories 39.43 34.02 46.10 43.73 54.41 61.76 71.33
Cosmetics & Drugstore 34.67 42.62 51.42 46.90 60.01 62.53 76.49
Furniture & Living 39.29 39.76 47.12 45.99 54.16 58.31 78.29
Toys & Baby 51.73 51.86 68.02 71.88 70.62 78.95 88.52
Sports & Leisure 38.37 43.52 51.42 51.45 61.08 67.77 87.59
Tools & DIY 55.31 54.23 62.23 61.74 69.85 73.88 85.98

English Translation

In addition to the German version of the benchmark, we also provide a Englisch version of the benchmark. The Englisch version was created by translating the German product offers to English using GPT-5.2. The English version offers exactly the same variants and the same splits as the German version. The table below compares the F1 results of the eight matchers on the English version to the results on the original German version. The table reports the difference of the F1 score on the English version minus the F1 score on the German version for all 27 benchmark variants.

Corner cases Training size Test set WordCooc Magellan RoBERTa R-SupCon HierGAT Ditto GPT-5.2 Simple GPT-5.2 Rule-Guided
20 % Small Seen −0.52 −3.50 +1.87 +5.11 +1.22 −0.52 −0.10 −2.19
Half-Seen +0.58 −4.17 +3.29 +0.11 −0.14 −3.73 −0.09 +0.63
Unseen +4.72 −3.83 +2.48 −1.33 +0.06 −2.65 −0.82 +0.15
Medium Seen −0.04 −1.67 +0.57 +3.90 −6.88 −3.72 −0.10 −2.19
Half-Seen +1.44 −2.20 +2.84 +0.89 −8.91 −6.99 −0.09 +0.63
Unseen +4.43 −1.69 +4.28 +1.17 −5.38 −5.27 −0.82 +0.15
Large Seen +0.81 −0.79 −0.21 +1.67 −5.23 −5.60 −0.10 −2.19
Half-Seen +0.71 −0.38 −1.06 −0.82 −8.87 −9.18 −0.09 +0.63
Unseen +6.47 +0.38 +2.10 −0.68 −3.83 −4.92 −0.82 +0.15
50 % Small Seen +0.85 −0.33 +5.87 +1.36 +18.86 −22.47 +0.57 −2.08
Half-Seen +0.54 +0.03 +5.67 −2.02 +13.49 −21.29 +0.41 +2.89
Unseen +4.86 −0.65 +4.78 −0.45 +12.63 −18.08 −1.06 +3.34
Medium Seen +0.30 +1.29 +2.78 +4.35 +0.69 −4.91 +0.57 −2.08
Half-Seen −0.57 +0.14 +4.86 −0.33 −0.94 −7.26 +0.41 +2.89
Unseen +2.99 −0.99 +3.02 +0.27 +1.85 −4.92 −1.06 +3.34
Large Seen −1.72 −1.27 +0.56 +1.81 −4.92 −6.91 +0.57 −2.08
Half-Seen +1.18 −2.99 +1.48 −0.48 −5.46 −9.04 +0.41 +2.89
Unseen +3.74 −2.53 +1.49 −1.29 −2.55 −6.06 −1.06 +3.34
80 % Small Seen +3.06 −1.57 −20.99 −0.03 +12.56 −3.60 −0.76 −0.38
Half-Seen +0.39 −1.17 −21.49 −1.17 +9.42 +0.03 +2.99 +2.96
Unseen +1.65 −0.55 −20.91 −0.97 +7.62 −2.15 +1.31 +0.89
Medium Seen −2.66 −2.61 +23.71 +0.51 −1.94 −6.40 −0.76 −0.38
Half-Seen −2.82 −3.02 +22.02 −0.81 −2.72 −0.99 +2.99 +2.96
Unseen +3.07 −1.63 +20.29 +1.24 −1.04 −1.24 +1.31 +0.89
Large Seen +0.93 −2.59 +2.31 −0.59 −5.30 +0.27 −0.76 −0.38
Half-Seen +1.18 −1.85 +5.54 −0.73 −2.54 +0.40 +2.99 +2.96
Unseen +4.74 −1.10 +5.35 +1.16 −0.52 +2.55 +1.31 +0.89
Mean (all 27 variants) +1.49 −1.53 +2.31 +0.44 +0.42 −5.73 +0.27 +0.69
Median (all 27 variants) +0.93 −1.57 +2.78 −0.03 −1.04 −4.92 −0.09 +0.63

The main result is that changing the language rarely changes the practical assessment of a matcher. WordCooc and Magellan remain below 46 mean F1 in both languages, which is too low for practical matching on this benchmark. Their small mean language differences do not alter that conclusion. More generally, the mean language difference remains within ±2.31 F1 for every matcher except Ditto. GPT-5.2 is similarly stable: the mean differences are +0.27 F1 with the Simple prompt and +0.69 F1 with the Rule-Guided prompt, and no individual difference exceeds 3.34 F1. At this scale, the results do not suggest a substantial systematic language effect for GPT-5.2.

Ditto is the notable exception. It scores 5.73 F1 lower on average on the English version, which is surprising because its RoBERTa backbone was pre-trained primarily on English text. The English disadvantage is largest with Small training sets at 8.27 F1 on average and narrows to 4.63 with Medium and 4.28 with Large training sets. Although the disadvantage becomes smaller as the training set grows, the results do not explain why Ditto performs worse on English overall.

The broader training-size pattern provides partial support for the same interpretation. Across the six supervised matchers, the mean absolute language gap decreases from 5.25 F1 with Small training sets to 3.77 with Medium and 2.76 with Large sets. RoBERTa's extreme differences account for part of this trend. After excluding RoBERTa, the corresponding values are 4.36, 2.65, and 2.86 F1. Language sensitivity therefore tends to be greater when little training data is available, but the decrease is not consistent for every matcher. This pattern does not confirm a general English advantage: even with Small training sets, fewer than half of the non-RoBERTa comparisons favor English.

Comparing the results of the matchers on the English version of the billiger.de benchmark to the results of the same matchers on the WDC Products Benchmark, which also features English product data, shows that the F1 results on billiger.de are generally lower, meaning that billiger.de Products (English) is more difficult than WDC Products. A reason for this might that Billiger.de Products spans a broader range of product categories, whereas WDC Products is dominated by offers for electronic products.

Cross-Language Matching

The comparison above changes the language of both records at once. In applications such as European price comparison, matchers also have to compare records across languages. This section evaluates this setting directly: all supervised matchers are trained on the German 80 % corner-case large training set (26,571 pairs) and selected on the corresponding German validation set (4,436 pairs). Each matcher is then evaluated on five variants of the 50 % unseen test set (4,437 pairs each) in which only the language of the records changes. In DE-DE both records are German, in DE-EN and EN-DE one record is German and the other English, in EN-EN both records are English, and Random mixes the four combinations evenly (1,109 or 1,110 pairs each). Pair IDs, labels, and the seen/unseen composition are identical in all five variants, so the differences between the columns are attributable to the language combination alone.

Model DE-DE DE-EN EN-DE EN-EN Random
WordCooc 43.6 36.8 36.6 28.2 33.2
Magellan 36.3 33.9 33.2 35.6 35.6
RoBERTa 63.4 53.3 54.5 58.1 56.7
R-SupCon 48.5 47.0 46.5 42.3 45.2
HierGAT 60.8 49.3 49.9 57.6 54.4
Ditto 65.4 55.3 55.1 62.4 59.2
GPT-5.2 Simple 70.5 71.5 70.8 72.9 70.7
GPT-5.2 Rule-Guided 80.9 83.5 81.7 80.9 82.6

WordCooc and Magellan both lose performance once records leave the training language, but to a different degree. WordCooc drops from 43.6 F1 on DE-DE to 28.2 on EN-EN, since translation removes the exact token overlap it relies on. Magellan stays within 3.1 F1 of its DE-DE score across all variants, as its character-level similarity features depend less on the language. Both matchers operate between 28.2 and 43.6 F1 on this test set, so these differences carry limited practical weight. R-SupCon is stable across the four variants that still contain German records (45.2 to 48.5 F1) and only loses performance on the fully English variant (42.3), though again at a low absolute level.

HierGAT reaches its best score on the German pairs it was trained on (60.8 F1), closely followed by the fully English variant (57.6), while both mixed variants drop to around 49 F1. RoBERTa and Ditto show the same ordering, and for all three models the Random variant, in which half of the pairs are same-language, lies in between. The models therefore do not fail on English data as such. As long as both records switch language together, the cues learned on German pairs, such as overlapping tokens between the two names and descriptions, stay intact, and the English pre-training of the underlying language models likely supports the transfer. In a mixed pair, the translation removes this surface overlap on one side only, and pairs that cross the language boundary never occur in the German training data. For a price comparison portal this distinction matters, because the cross-border case is exactly the mixed-language case. Translating one record into the language of the other before matching may recover much of the loss, although the experiment does not test this.

GPT-5.2 delivers the best scores in every variant: 80.9 to 83.5 F1 with the Rule-Guided prompt and 70.5 to 72.9 with the Simple prompt. Taking run-to-run variation into account, the differences between the language combinations are minimal, so the results do not indicate a language effect for the zero-shot models. For applications that have to match offers across languages, and where the budget allows it, a zero-shot LLM is therefore the most reliable choice among the evaluated systems.

How to Use the Benchmark

The benchmark is distributed as gzipped JSON files with one record pair per line. Each line contains both product records, the binary match label, and a hard-negative indicator. The file name encodes the configuration: products80cc20rnd000un_train_large.json.gz denotes 80 % corner cases, 20 % random pairs, and the large training split. Training files always use 000un. The unseen-product share is defined by validation and test products relative to the training data. The files can be read directly with pandas:

import pandas as pd

pairs = pd.read_json(
    "products80cc20rnd000un_train_large.json.gz",
    compression="gzip",
    lines=True,
)

The complete reproduction code is available in the GitHub repository. From the repository root, install the fixed dependencies and generate all model-specific inputs with:

python3.10 -m venv .venv
source .venv/bin/activate
python -m pip install -r environments/requirements.txt

python src/processing/prepare_pairs.py
python src/processing/prepare_pretraining.py
python src/processing/prepare_ditto_hiergat.py
python src/processing/prepare_wordcooc.py
python src/processing/prepare_magellan.py

The launchers in slurm_runs/ and the model directories run the complete German and English experiment matrix with three fixed seeds. All generated metrics are written below results/generated/ and can be collected into one CSV with:

python src/summarize_results.py

Exact commands for WordCooc, Magellan, RoBERTa, R-SupCon, Ditto, HierGAT, and GPT-5.2 are documented in REPRODUCTION.md. The GPT-5.2 runs require an OPENAI_API_KEY and use the Batch API.

Recommended default: If you only want to report results for a single variant of billiger.de Products, we recommend 80cc20rnd050un: 80 % corner cases and a half seen test set. Use products80cc20rnd050un_gs.json.gz and, for supervised matchers, the corresponding training and validation files. Please always state the language and training-set size.

Report the result as: Billiger.de Products (German or English), 80 % corner cases, half-seen (050un), [training size], F1 = X (mean ± standard deviation over seeds 0, 1, and 2). Cite the benchmark paper once available and link to this repository. For zero-shot methods, state the model version and prompt instead of a training size.

For supervised matchers, pick a training set, select hyperparameters on the corresponding validation set, and report F1 on the nine test sets. The test sets are the same for all training set sizes, so results obtained with different amounts of training data remain comparable.

For zero-shot and in-context learning, demonstration examples should be drawn from the training or validation sets.

For language comparisons, run the same configuration on the German and the English version. Splits, labels, and identifiers are identical, so any difference in the scores is attributable to the language of the records and the translation process. Since the record identifiers line up across the two versions, German records can also be paired with English ones to build cross-lingual matching tasks.

Downloads

Both language versions are available as zip archives containing all training, validation, and test sets described above. The five aligned test set variants of the cross-language experiment are available as a separate archive.

German version (original)

9 training sets, 27 validation sets, 9 manually verified test sets.

Download solute_de.zip
English version (translation)

Structurally identical translation of every record with the same splits, labels, and identifiers.

Download solute_en.zip
Cross-language test sets

The five aligned test set variants of the cross-language experiment (DE-DE, DE-EN, EN-DE, EN-EN, Random), 4,437 pairs each.

Download solute_cross_language.zip

Conclusion

Billiger.de Products extends product matching evaluation beyond English with a large German benchmark and an aligned English translation spanning thirteen product categories. The 27 variants of the benchmark vary the training set size, the share of corner cases, and the fraction of unseen products in the test set.

The validation of the benchmark showed that both the benchmark variant and the product category influence the results. Supervised matchers generally benefit from the larger training sets and score lower as the share of corner-cases or the amount of unseen entities in the test set increases. Tools and electronics are among the easiest product categories, whereas for clothing and furniture matchers generally achieve lower scores.

The comparison of the performance of the matchers on billiger.de Products (English) to the performance of the same matchers on WDC Products showed that Billiger.de Products is more difficult than WDC Products. This feature might make the English version of Billiger.de Products an interesting alternative to WDC Products for matchers that have saturated WDC Products.

The cross-language experiment showed that supervised matchers trained on German pairs keep most of their performance when both records switch to English but lose about 10 F1 points on mixed German-English pairs. Zero-shot GPT-5.2 reaches the best scores on all language combinations and shows no language effect, which makes LLMs the most reliable choice for cross-language matching among the evaluated systems.

We hope that the benchmark will support the evaluation and development of entity matching methods for monolingual German and English settings, as well as for cross-lingual matching scenarios.

Questions and feedback are welcome. Please contact Aaron Steiner or open an issue in the GitHub repository.

Citation

A paper describing the benchmark will be published on arXiv in August 2026. Citation information will be added here as soon as it is available.