Skip to content
Social Sense

Interactive · the scoring pipeline

Score a post the way the 2023 cluster did

Tweets went through seven steps on the Melbourne Research Cloud: clean, tokenise, lemmatise, score, bucket, geocode and count. Toots took a shorter path through the Mastodon harvester, with no cleaning and no place. Both paths have been ported line for line to TypeScript, including NLTK's tokenisers, the WordNet lemmatiser and VADER, and checked against the original output. Pick a path, type below and watch each step.

Lab

From raw text to a 1-9 score

Score it as

Runs entirely in your browser. Nothing is sent or stored.

Try an example

Loading the lexicon, WordNet and Punkt model…

  1. 1Strip mentions, hashtags and links

    twitter/processor.py
  2. 2Split into sentences (Punkt) and words (NLTKWordTokenizer)

    nltk.word_tokenize
  3. 3Lemmatise every token as a noun (WordNet)

    analyzer.py normalize_string
  4. 4Score with VADER

    nltk.sentiment.vader
  5. 5Bucket the compound score into 1-9

    analyzer.py sentiment_analysis
  6. 6Match the place to a suburb (SAL)

    twitter/utils.py

    Add a place name to see the geocoder at work. The 2023 processor skipped tweets without one.

  7. 7Would CouchDB's MapReduce views count it?

    MapReduce/Income, MapReduce/Crime

Fidelity

How faithful is the port?

The port keeps the original's quirks on purpose. Tokens are lemmatised as nouns, so "was" becomes "wa". VADER looks up a repeated word's first position. The place matcher tries every combination of words, shortest first, and keeps the first known place, so a one-word hit wins over the full name: "Albert Park, Victoria" lands in Albert (NSW), "St Kilda, Victoria" in St Kilda (SA) and "Port Melbourne" in Melbourne. An extra lookup in the original that could never succeed (it passed the function object instead of the string) is simply absent. Language detection (langdetect) is not ported because it never affected the score.