Skip to content
Social Sense

COMP90024 · The University of Melbourne · 2023

Does the mood online match life on the ground?

In 2023 a five-person team built a cloud system that read 37,823,414 tweets and harvested 1,659,690 toots, scored their sentiment, pinned them to Victorian suburbs and set them against official income and crime statistics. This is that project, rebuilt so anyone can explore it.

Short on time? Take the guided tour: three captioned walkthroughs

Melbourne587,004 tweets

1,085 Victorian suburbs · 719,336 geotagged tweets, Feb-Jul 2022

dot size = tweets more negative neutral more positive

The brief

Build a cloud that tells stories about life in Australia

Assignment 2 of COMP90024 asked each team to use the Melbourne Research Cloud to harvest social media, store it in a distributed database, analyse it at scale and present scenarios that compare online talk with official data from the Spatial Urban Data Observatory (SUDO), all deployed automatically.

Team 57 chose two questions: do people in higher- and lower-income areas talk about money differently, and does the tone of crime talk follow where crime is recorded?

57 GB
Twitter corpus provided
2,418,617
tweets matched to a suburb
3 nodes
CouchDB cluster on MRC
8 vCPUs
the whole cloud budget

How it worked

From a 57 GB file to a dashboard, in five moves

Scroll to follow a tweet through the original system. The cloud no longer exists; every number it produced has been recovered and checked.

  1. Twitter corpusMastodon
  2. MPI processorsHarvesters
  3. CouchDB cluster
  4. MapReduce views
  5. SUDOFlask API
  6. React dashboard
  7. All inside the Melbourne Research Cloud, deployed with Ansible and Docker Swarm
  1. 01 · Harvest

    MPI ranks split the 57 GB Twitter file by byte ranges and parse it line by line, while three Mastodon harvesters ask mastodon.social, mastodon.au and tictoc.social for the next 40 toots every ten minutes.

    37.8M tweets read · 1.66M toots

  2. 02 · Score and store

    Each post is cleaned, tokenised, lemmatised and scored by VADER on a 1-9 scale, matched to a suburb, then bulk-loaded 1,000 at a time into a CouchDB cluster replicated across three nodes.

    2.42M geotagged tweets

  3. 03 · Summarise with MapReduce

    CouchDB views emit every score keyed by suburb, once for all tweets and again for posts mentioning income or crime keywords, and a _stats reduce keeps count, sum, min and max.

    3 views · 6,799 suburb summaries

  4. 04 · Serve and compare

    A Flask API joins the view results with SUDO income and crime tables and ABS boundaries, renders Plotly figures, gzips them, and a React dashboard draws them side by side.

    37-46 MB of gzipped Plotly JSON per map

  5. 05 · Automate

    Ansible playbooks create the instances, volumes and security groups on the Melbourne Research Cloud, form the CouchDB cluster and deploy the services onto Docker Swarm, where they can be scaled with one script.

    8 vCPUs · 500 GB · 6 instances

What the data says

Four findings, recomputed

Each card links to an interactive page with the full data and method.

Twitter

Mostly neutral, leaning positive

48% of geotagged tweets score a neutral 5, and the positive side outweighs the negative. Suburb averages show no geographic pattern.

1 · Extremely Negative: 39,726 (1.6%)1.6%12 · Very Strongly Negative: 80,238 (3.3%)3.3%23 · Strongly Negative: 132,745 (5.5%)5.5%34 · Negative: 120,487 (5.0%)5.0%45 · Neutral: 1,158,765 (47.9%)48%56 · Positive: 187,854 (7.8%)7.8%67 · Strongly Positive: 321,908 (13.3%)13%78 · Very Strongly Positive: 241,710 (10.0%)10.0%89 · Extremely Positive: 135,184 (5.6%)5.6%9
Sentiment of all geotagged tweets
ScoreCountShare
1 (Extremely Negative)397261.6%
2 (Very Strongly Negative)802383.3%
3 (Strongly Negative)1327455.5%
4 (Negative)1204875.0%
5 (Neutral)115876547.9%
6 (Positive)1878547.8%
7 (Strongly Positive)32190813.3%
8 (Very Strongly Positive)24171010.0%
9 (Extremely Positive)1351845.6%
Explore

Scenario 1 · Income

Richer areas do not tweet happier about money

Across 167 SA2 areas, median income and the tone of income tweets barely move together (r = 0.14, explaining 1.9% of the variation). Income talk is concentrated in Melbourne's CBD.

Median income (x) vs average tone of income tweets (y), one dot per SA2 · dashed line: least squares

Explore

Scenario 2 · Crime

Crime talk is dark everywhere, not darker where crime is high

The most common score for crime tweets is 1 (extremely negative). But suburbs with more crime talk are not measurably gloomier (r = −0.05).

1 · Extremely Negative: 4,436 (24.9%)25%12 · Very Strongly Negative: 3,205 (18.0%)18%23 · Strongly Negative: 2,382 (13.4%)13%34 · Negative: 1,457 (8.2%)8.2%45 · Neutral: 3,039 (17.1%)17%56 · Positive: 910 (5.1%)5.1%67 · Strongly Positive: 1,058 (5.9%)5.9%78 · Very Strongly Positive: 822 (4.6%)4.6%89 · Extremely Positive: 487 (2.7%)2.7%9
Sentiment of crime-related tweets
ScoreCountShare
1 (Extremely Negative)443624.9%
2 (Very Strongly Negative)320518.0%
3 (Strongly Negative)238213.4%
4 (Negative)14578.2%
5 (Neutral)303917.1%
6 (Positive)9105.1%
7 (Strongly Positive)10585.9%
8 (Very Strongly Positive)8224.6%
9 (Extremely Positive)4872.7%
Explore

Mastodon

Every server has its own mood, and the lexicon speaks English

mastodon.social looked almost uniformly neutral in 2023 while mastodon.au resembled Twitter. Re-scoring a surviving week shows why: non-English toots fall to neutral, and German 'die' reads as negative.

mastodon.social

1 · Extremely Negative: 7,731 (1.4%)12 · Very Strongly Negative: 8,289 (1.5%)23 · Strongly Negative: 5,038 (0.9%)34 · Negative: 15,756 (2.9%)45 · Neutral: 429,414 (78.5%)56 · Positive: 20,428 (3.7%)67 · Strongly Positive: 24,447 (4.5%)78 · Very Strongly Positive: 19,356 (3.5%)89 · Extremely Positive: 16,460 (3.0%)9
mastodon.social sentiment
ScoreCountShare
1 (Extremely Negative)77311.4%
2 (Very Strongly Negative)82891.5%
3 (Strongly Negative)50380.9%
4 (Negative)157562.9%
5 (Neutral)42941478.5%
6 (Positive)204283.7%
7 (Strongly Positive)244474.5%
8 (Very Strongly Positive)193563.5%
9 (Extremely Positive)164603.0%

mastodon.au

1 · Extremely Negative: 11,331 (2.1%)12 · Very Strongly Negative: 27,363 (5.0%)23 · Strongly Negative: 16,426 (3.0%)34 · Negative: 51,299 (9.4%)45 · Neutral: 232,983 (42.6%)56 · Positive: 33,381 (6.1%)67 · Strongly Positive: 75,399 (13.8%)78 · Very Strongly Positive: 46,061 (8.4%)89 · Extremely Positive: 52,277 (9.6%)9
mastodon.au sentiment
ScoreCountShare
1 (Extremely Negative)113312.1%
2 (Very Strongly Negative)273635.0%
3 (Strongly Negative)164263.0%
4 (Negative)512999.4%
5 (Neutral)23298342.6%
6 (Positive)333816.1%
7 (Strongly Positive)7539913.8%
8 (Very Strongly Positive)460618.4%
9 (Extremely Positive)522779.6%
Explore

2026 upgrade

How sure, and who checks the AI?

The revival added the questions a statistician would ask of the 2023 results, and a way to query the data in plain English that shows its working.

Hands on

Run the 2023 pipeline in your browser

The original Python scoring code (NLTK's tokenisers, WordNet and VADER) has been ported to TypeScript and verified against the original on 44,156 real toots. Type a post and watch it become a number.

Open the lab

About this project

Who built it, then and now

Subject
COMP90024 Cluster and Cloud Computing
University
The University of Melbourne
When
Semester 1, 2023, Assignment 2
Team
Team 57
  • Sunchuangyu (Rin) HuangFrontend lead; data processing and API tools
  • Xuan WangCouchDB cluster; Ansible automation
  • Wei ZhaoFlask backend; Ansible automation
  • Zongchao XieScenario analysis; data processing
  • Runqiu FeiData processing; database
rNLKJA/Australia-Social-Media-Analytics-on-the-Cloud
Original stack compared with the revived stack
2023 original2026 revival
ComputeMelbourne Research Cloud, 6 instancesStatic pages + serverless functions (Vercel)
Data storeCouchDB 3.2 cluster, 3 nodes, ~400 GBSQLite analytics.db, 1.7 MB, read-only
Processingmpi4py, MapReduce viewsuv scripts re-running the original code
BackendFlask + Gunicorn, gzipped Plotly JSONNext.js Server Components
FrontendReact 18, MUI, Plotly, TailwindNext.js 16, Tailwind v4, MapLibre, SVG
MapsPlotly mapbox, 37-46 MB per mapOpenFreeMap tiles + 1.2 MB of TopoJSON
SentimentNLTK 3.8.1 in PythonSame algorithm, ported to TypeScript
DeploymentAnsible + Docker Swarmgit push

Academic integrity. The original submission is preserved unchanged in the repository's coursework/ folder for reference. The revival rebuilds the presentation and re-runs the team's own code; it does not alter the methods or the conclusions. The assignment specification and course-provided raw data are not published here.