Al
|
a38624ba59
|
[similarity] Adding NameDeduper base class for deduping geographic names using the new Soft TFIDF similarity
|
2015-10-31 00:57:02 -04:00 |
|
Al
|
a5c1296044
|
[similarity] Adding Jaccard similarity with word frequencies instead of simple sets, better for ideographic scripts (Han, Hangul, etc.) in the absence of word segmentation since there may be many high frequency characters
|
2015-10-30 14:35:38 -04:00 |
|
Al
|
cccc3e9cf5
|
[similarity] Using Soft-TFIDF for approximate name matching. Soft-TFIDF is a hybrid string distance metric which balances local token similarities (using Jaro-Winkler similarity by default) allowing for slight spelling errors with global TFIDF statistics so that very frequent words don't affect the score as much
|
2015-10-30 02:48:16 -04:00 |
|