Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Short answer: DBP15K is the closest established starting point for cross-lingual entity alignment, but it covers Chinese–English and Japanese–English—not Korean, direct Korean–Japanese or Korean–Chinese pairs, or corporate records. The resources identified here do not establish a ready-made labeled dataset for Korean–Japanese–Chinese corporate-name matching. For that task, use existing benchmarks as limited baselines and build or obtain adjudicated labels from the company records you actually need to resolve.

First, define what “entity resolution” means for your project

For corporate-name matching, the target label is usually a judgment that two records refer to the same company—or an identity cluster that groups records under a defined policy. That is different from several adjacent tasks:

  • Record-to-record resolution: decides whether records from registries or other sources identify the same entity.
  • Knowledge-graph (KG) entity alignment: matches nodes in separate knowledge graphs. DBP15K addresses this task.
  • Entity linking: maps a mention in text to a knowledge-base entity.
  • Named-entity recognition (NER): identifies and classifies entity mentions in text.
  • Entity classification: assigns a category to an entity or page.

Linking, NER, and classification data can help generate candidates, detect names, or add weak supervision. Their labels do not, by themselves, establish that two corporate records identify the same legal entity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which labeled resources are closest to the task?

Resource What its labels cover Fit for Korean–Japanese–Chinese corporate matching Access and caveats
DBP15K Separate Chinese–English, Japanese–English, and French–English DBpedia alignment subsets. The cited IJCAI 2019 paper reports 15,000 reference alignment links per subset. Closest standard alignment benchmark for Chinese and Japanese, but it has no Korean subset and is not shown to contain corporate-registry records. English is a bridge for separate pairwise experiments, not proof of direct Korean–Japanese or Korean–Chinese gold labels. Research benchmark. Check the original release and license before reuse. IJCAI 2019 paper.
EntMatcher DBP15K splits Repository-organized alignment links and graph triples, including zh_en and ja_en folders with support, validation, and reference-link files. Convenient for implementation-oriented general alignment experiments; it retains DBP15K’s language and domain limitations. The repository describes a 70% test, 20% train, and 10% validation split for its listed gold alignment links. Confirm the split and provenance for the specific experiment. EntMatcher repository.
Hansel Human-annotated Chinese entity-linking test data: 10,000 examples, with few-shot and zero-shot slices; Wikidata is the knowledge base. Training and validation examples come from Wikipedia hyperlinks. Useful for Chinese mention-to-KB linking, especially tail or emerging entities. It is not cross-source company-record pair labeling. The project repository states CC BY-SA for Hansel; verify component-data terms and current conditions. Hansel project.
SHINRA2021-ML / SHINRA2020-ML Japanese Wikipedia pages annotated with Extended Named Entity categories, language links, target-language Wikipedia pages, and scripts or data for classification training. Can support multilingual entity-category classification or language-linked examples. It does not directly label whether company records identify the same legal or operating entity. Files are available in multiple formats and sizes; review project terms and data notices. SHINRA project.
Mewsli-9 289,087 entity mentions from 58,717 originally written WikiNews articles in nine languages, linked to Wikidata. Japanese is included; Korean and Chinese are not in its listed language set. Useful for multilingual entity-linking evaluation and domain-shift work, not corporate record matching. The paper uses a WikiNews snapshot dated 2019-01-01; automatically extracted links trade annotation quality for scale and diversity. Mewsli-9 paper.
TAC KBP Chinese Cross-lingual Entity Linking, 2011–2014 LDC collection of English and Chinese documents, queries, entity-type information, knowledge-base links, and NIL equivalence clusters. Potentially useful for Chinese–English entity-linking training or evaluation and realistic text. It does not provide Korean/Japanese coverage or corporate-name record linkage. The LDC catalog lists a release date of November 17, 2017, and an LDC user agreement for non-members. LDC catalog entry.
MELD A standardized collection of NER datasets across languages and domains, with annotations that depend on the source dataset. May help with mention detection or entity-type recognition, but NER labels do not identify same-entity record pairs. Licensing is source-specific; some datasets must be fetched from their original sources because of licensing restrictions. MELD documentation.
KORE 50DYWC An entity-linking evaluation set expanded to DBpedia, YAGO, Wikidata, and Crunchbase. Relevant to entity linking across knowledge-base targets, not a three-language corporate-name pair corpus. Check release terms and label compatibility for your intended use. LREC 2020 paper.

What the reported counts do—and do not—tell you

The counts describe different units, so they should not be compared as if they were equivalent volumes of company-match labels.

  • DBP15K: the IJCAI 2019 paper reports 15,000 reference alignment links in each language-pair subset. Its entity-count table separately lists 66,469 Chinese-side entities and 65,744 Japanese-side entities; those are graph entities, not matched corporate records.
  • Hansel: the project page reports 10,000 Chinese entity-linking test examples. Its dataset-page updates are dated 2022–2023.
  • Mewsli-9: the 2020 paper reports 289,087 mentions in 58,717 WikiNews articles. Its listed languages include Japanese, not Korean or Chinese.
  • About 2,000 reviewed pairs: Tae Kim’s title-matched article search-result page, dated September 23, 2026, reports that the author manually reviewed about 2,000 pairs after not finding a public labeled set for this task. That is an author-reported experience, not a published benchmark statistic or a systematic census of available datasets. Tae Kim’s article.

How to build useful corporate-match labels

When your target is company identity, the most important dataset is one whose labels reflect your own definition of “same company.” A name match alone may confuse a parent with a subsidiary, or two unrelated companies with similar names. Decide in advance how to handle subsidiaries, joint ventures, aliases, legal-entity changes, and changes in ownership over time.

  1. Define the identity policy. Specify which record pairs count as a match, including how the project treats parent/subsidiary relationships, joint ventures, aliases, mergers, and time-varying ownership. Decide whether the desired identity is the legal entity, an operating business, or another unit.
  2. Scope the sources and languages. Identify the Korean, Japanese, and Chinese registries or source systems in scope, the relevant time period, and which fields are available. Do not assume a benchmark’s Wikipedia or knowledge-graph coverage represents those records.
  3. Generate candidates, but preserve provenance. Wikipedia or Wikidata links, transliteration, and name similarity may help find pairs to review. Store how each candidate was generated and its confidence; do not silently promote candidate-generation output to gold labels.
  4. Adjudicate the pairs. Have reviewers apply the written identity policy to positive, negative, and ambiguous cases. Record uncertainty or an “unresolved” outcome rather than forcing a match when the evidence is insufficient.
  5. Split for evaluation and audit leakage. Keep related records or identity clusters from leaking across train, validation, and test sets. Record the source, annotation method, split construction, and license so results can be interpreted and the data reused appropriately.

How to use the existing resources without overstating them

Use DBP15K as a general alignment baseline

DBP15K can help you test a cross-lingual KG-alignment pipeline on Chinese–English or Japanese–English links. Report those as pairwise, English-anchored graph-alignment experiments. Do not describe them as Korean–Japanese–Chinese corporate-name results or as direct Korean-pair gold labels.

Use linking and classification datasets for supporting tasks

Hansel and TAC KBP can inform Chinese entity-linking work; SHINRA can support Japanese-centered entity classification; Mewsli-9 can provide a multilingual linking evaluation with Japanese; MELD can help with NER. These can support components of a larger workflow, but none of their described labels answers the corporate record-identity question directly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check terms and splits before training or sharing data

Language coverage and label type are only part of dataset suitability. Confirm the precise release, source data, split and leakage risks, access agreement, and whether the applicable terms allow your intended training, commercial use, or redistribution. In particular, a repository’s convenience split does not remove the need to check the original benchmark’s provenance and terms.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Is there a public three-language corporate-name dataset?

The resources identified here do not establish a ready-made, labeled Korean–Japanese–Chinese corporate-name matching set. They also do not prove that no specialized dataset exists in a particular industry or jurisdiction. Treat the reported search experience as one author’s account, not a universal claim of nonexistence; check relevant registry, industry, and jurisdiction-specific sources for your use case.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.