Detecting matches
CDF uses string tokenization to break strings into substrings. We refer to these substrings as tokens. The tokens in the source and target entities are compared and used to calculate the similarity between THE two entities. By default, the model uses thename fields to find similarities, but you can configure the model to use any field in CDF.
Suppose you select the simple similarity scoring model to match entities. In that case, the model uses the regular expression \p{L}+|[0-9]+ to remove all punctuation or dash characters and split the string into tokens that consist only of letters or numbers. For instance, the string 11-PDN-26540J-60 splits into the tokens 11, PDN, 26540, J, and 60.
The entity matching model calculates similarity features by comparing strings based on the similarities between tokens and their order.
For medium to large data sets, it’s impossible to compare all source and target combinations. To reduce the number of comparisons, the model uses a blocking step to remove target entities that are not similar to the source entity.
Defining match scores
The entity matching model uses thefeatureType property to define the combination of similarity features created and used in the entity matching model.
Each feature produces one feature score between the source and target. A higher feature score means a more likely match but can’t be interpreted as a probability.
The model calculates all features specified from featureType for each pair of fields in the matchField property. For instance, if featureType entails N individual features and matchField specifies M pairs of fields, the model calculates N x M features for each candidate match.
Be cautious about adding extra matchFields as they may impact performance if they contain little or no similarity information.
Calculating match scores
An unsupervised model or a supervised model calculates the match score using the N x M feature scores and uses the match score to decide whether a source and target entity is a match.Unsupervised model
Use the unsupervised entity matching model if you don’t have any or only a few known matches as training data. The unsupervised model is the default model and creates similarity feature scores between the sources and targets to return a weighted average as the match score. The model uses a simple average, assigning the same weight to each feature. Therefore, it doesn’t make sense to calculate many different features when using the unsupervised model, andmatchFields should only include fields known to contain relevant similarities.
In addition, you should limit featureType to simple features such as Simple or Bigram.
Supervised model
Use the supervised entity matching model when you have some known matches that can be used to decide which features to weigh up or down to calculate the match score from the feature scores. The supervised model often performs better with good training data since it can weigh the different features into fine-tuned criteria to determine a similarity between the source and target entities. It also allows you to use morematchFields by automatically determining which matchFields to weigh as stronger or weaker.