Commit Graph

1866 Commits

Author SHA1 Message Date
Al
d1267145f7 [fix] args to wget 2015-04-13 19:02:50 -04:00
Al
d771da7c78 [i18n] unicode scripts file downloaded and cached locally 2015-04-13 19:02:29 -04:00
Al
cc4d2d08eb [cldr] Adding script to download latest cldr release instead of pulling from the repo 2015-04-13 01:03:15 -04:00
Al
acb575c84c [fix] splitting out methods for unicode scripts 2015-04-12 15:21:23 -04:00
Al
d50d7d182e [fix] geonames import script for admin 1 codes 2015-04-12 12:16:08 -04:00
Al
fdd0c489f3 [fix] refactoring unicode script fetching into more reusable functions 2015-04-09 02:18:13 -04:00
Al
e03c1f21a7 [unicode] generate C headers/data files from unicode.org scripts 2015-03-18 16:59:58 -04:00
Al
6c8e5b45a4 [fix] removing building alias (for OSm it means building category), fix to fetch script 2015-03-18 08:40:07 -04:00
Al
88554c1ef7 [i18n] adding CLDR languages script to this repo 2015-03-18 08:01:36 -04:00
Al
2cf909c01e [utils] script utils 2015-03-17 18:39:08 -04:00
Al
aeac0fe8c0 [geodata] Script to construct OSM training examples for building language dictionaries, disambiguating between abbreviations, classifying venues by type and formatting addresses for use in a sequence model with Lokku's address-formatting repo. 2015-03-17 18:11:07 -04:00
Al
0437271c92 [geodata] OSM planet fetch needs to convert ways/relations to nodes for all data sets 2015-03-17 16:51:17 -04:00
Al
621b25c964 [geodata] script to fetch/transform OSM planet (needs about 100GB of disk free) training language models 2015-03-16 00:45:14 -04:00
Al
26c2823208 [fix] comma 2015-03-14 18:58:18 -04:00
Al
3e20b4f600 [fix] Capturing GeoNames canonical and alternate names with a UNION ALL query, creating C headers with the field orderings for parsing the TSV file downstream 2015-03-14 18:02:14 -04:00
Al
284af74ba4 [geodisambig] Python scripts to prep GeoNames records for trie insertion 2015-03-13 11:56:48 -04:00