Interactive · the scoring pipeline
Score a post the way the 2023 cluster did
Tweets went through seven steps on the Melbourne Research Cloud: clean, tokenise, lemmatise, score, bucket, geocode and count. Toots took a shorter path through the Mastodon harvester, with no cleaning and no place. Both paths have been ported line for line to TypeScript, including NLTK's tokenisers, the WordNet lemmatiser and VADER, and checked against the original output. Pick a path, type below and watch each step.
Lab
From raw text to a 1-9 score
Runs entirely in your browser. Nothing is sent or stored.
Try an example
Loading the lexicon, WordNet and Punkt model…
1Strip mentions, hashtags and links
twitter/processor.py2Split into sentences (Punkt) and words (NLTKWordTokenizer)
nltk.word_tokenize3Lemmatise every token as a noun (WordNet)
analyzer.py normalize_string4Score with VADER
nltk.sentiment.vader5Bucket the compound score into 1-9
analyzer.py sentiment_analysis6Match the place to a suburb (SAL)
twitter/utils.pyAdd a place name to see the geocoder at work. The 2023 processor skipped tweets without one.
7Would CouchDB's MapReduce views count it?
MapReduce/Income, MapReduce/Crime
Fidelity
How faithful is the port?
The port keeps the original's quirks on purpose. Tokens are lemmatised as nouns, so "was" becomes "wa". VADER looks up a repeated word's first position. The place matcher tries every combination of words, shortest first, and keeps the first known place, so a one-word hit wins over the full name: "Albert Park, Victoria" lands in Albert (NSW), "St Kilda, Victoria" in St Kilda (SA) and "Port Melbourne" in Melbourne. An extra lookup in the original that could never succeed (it passed the function object instead of the string) is simply absent. Language detection (langdetect) is not ported because it never affected the score.