Show HN: Open-source text-to-geolocation models | HN Mirror

Y	Hacker News new \| ask \| show \| jobs

	Show HN: Open-source text-to-geolocation models (github.com)
	46 points by yachayai 1348 days ago
	Yachay is an open-source community that works with the most accurate text-to-geolocation models on the market right now

7 comments

neoncontrails 1348 days ago

This is _really_ cool. Early in the pandemic I released a local news aggregation tool that aimed to aggregate COVID-related content and score it for relevance using an ensemble of ML classification models, including one that would attempt to infer an article's geographic coordinates. Accuracy peaked at about ~70-80%, which was just not quite high enough for this use case. With a large enough dataset of geotagged documents I'm pretty sure we could've improved that by another 10-15% which would've likely been "good enough" for our purposes. But one of the surprising things I took away from the project was that there's not a well-defined label for this category of classification problems, and as a result there's few datasets or benchmarks to encourage progress.

yachayai 1340 days ago

Thanks! COVID is a great example of the use case, and we agree problems like this need more attention - we've shared some data already, and will continue to share more with the public to encourage collaboration on this. Hope you will find something useful there for your future projects:)

DOsinga 1348 days ago

This does look interesting but as other comments have pointed out without data or weights it's not clear how well this works. The training notebook seems to suggest it is not actually improving all that much on the training data

yachayai 1340 days ago

As mentioned in the response below, we have posted the challenge with a respective data set of ~ 500k tweets from 100 regions around the world - https://github.com/1712n/yachay-public/tree/master/conf_geot...

We are working on adding more data as well - feel free to create a GitHub issue if there's more you need - we're going to be working on everything there is to do to help the developers here:)

JimDabell 1348 days ago

Depending upon your use-case, you can get pretty good results by using spaCy for named entity recognition then matching on the titles of Wikipedia articles that have coördinates.

yachayai 1340 days ago

Agreed. That said, more often than not, as mentioned in the comment above (COVID use case), we'd look for a higher recall value in the predictions - there, NERs, although helpful, wouldn't be our go-to solution. This is exactly the reason why we open sourced the infrastructure and are rolling out the data

rmbyrro 1348 days ago

Tried this in the past, it's too limited... There are too many ways certain locations can be referred to. Take: New York City, NYC, NY, New York, NYCity, so on...

JimDabell 1348 days ago

Wikipedia handles “New York City” and “NYC” as intended. “NY” and “New York” are ambiguous to both machines and humans (are you referring to the city or the state?) and if you have a resolution strategy for this then Wikipedia gives you the options to disambiguate. I’ve never seen “NYCity” used by anybody.

rmbyrro 1348 days ago

If you start processing web articles on the scale of millions you'll be surprised by how creative people can be. Not talking about tweets, just news and blog articles.

JimDabell 1347 days ago

Not surprised, just not relevant. The criteria here is “you can get pretty good results”, not “you must be able to process millions of articles without failure”.

rmbyrro 1347 days ago

If a method is not generalizable to the entire dataset, it's not that useful.

When processing text at large scale, the usefulness of heuristic approaches like the one we're discussing diminishes rapidly.

tomthe 1348 days ago

There are no weights and no data, only some code to create a pytorch character based network and train it. Will you provide weights or data in the future? Do you have any benchmark over Nominatim or Google maps?

I think something like this (but with more substance) could be helpful for some people, especially in the social sciences.

rmbyrro 1348 days ago

Yea, I was expecting a general-purpose model or dataset to train a model. The idea is great, but - as it currently stands - of no use to most people.

yachayai 1341 days ago

We have posted a challenge with the respective data set of ~ 500k tweets from 100 regions around the world - https://github.com/1712n/yachay-public/tree/master/conf_geot...

rmbyrro 1348 days ago

This would have been tremendously useful in a project I worked at a few years ago.

It's really a difficult task to parse text at large scale with accurate geographical tagging.

yachayai 1340 days ago

What was the goal of the project?

cyanydeez 1347 days ago

Probably should combine this with DELFT.

TuringNYC 1348 days ago

Has anyone got this working? Curious if someone could PR a dependencies file that can be used to run this?

yachayai 1340 days ago

We have updated the wiki and published the dependencies:) Feel free to create a GitHub issue for any further requests, or ask away in our discord - https://discord.gg/msWFtcfmwe