Data & methods
Where the numbers come from, and how far to trust them
Social Sense was a five-person cloud computing project. This page documents the data, the original processing, what the revival recomputed and the checks that tie every chart back to the 2023 outputs, then how uncertainty is reported, what the optional AI does, and the decisions behind it.
Sources
Data
Twitter corpus
57 GB of tweets supplied by the University via the Australian Data Observatory. The team processed 15.9 GB (37.8 million tweets) before the deadline; 2.42 million carried a place that matched a suburb.
- How it is used here
- Per-suburb counts and score sums from three CouchDB MapReduce views (all, income, crime), February-July 2022.
- Terms
- Licensed for coursework; only aggregates are published here.
Mastodon
Public timelines of mastodon.social, mastodon.au and tictoc.social, harvested every ten minutes (40 toots per request) from 31 Dec 2021; 1.66 million toots by May 2023.
- How it is used here
- The dashboard's histograms, plus one surviving week of raw mastodon.social toots (2-9 May 2023) re-scored for hourly and language aggregates.
- Terms
- Public posts; no text, ids or usernames are stored here.
SUDO: personal income
ABS personal income by SA2 (2015-16), mean, median, sum and median age of earners, on 2016 SA2 boundaries.
- How it is used here
- Scenario 1 and the capital-city summary.
- Terms
- ABS data via the Spatial Urban Data Observatory.
SUDO: jobs and income by industry
ABS Jobs in Australia 2018-19, employee jobs and median income per job by industry and SA2.
- How it is used here
- The industry bar chart on the income page (as summarised in 2023).
- Terms
- ABS data via SUDO.
SUDO: criminal incidents
Crime Statistics Agency Victoria, offences by division per LGA (2011 LGA codes), reference year 2019.
- How it is used here
- Scenario 2.
- Terms
- Crime Statistics Agency Victoria via SUDO.
Boundaries
ABS ASGS suburbs and localities (SAL 2021), SA2 (2016), LGA (2019) and GCCSA (2021, dissolved to states).
- How it is used here
- Choropleths (simplified to 6-8% of vertices with mapshaper) and the suburb-to-area crosswalk.
- Terms
- © Australian Bureau of Statistics, CC BY 4.0.
The assignment specification and the raw course datasets are not reproduced on this site. The team's report is kept in the repository under coursework/.
2023
The original processing
- 1
Split the corpus
mpi4py ranks each read a byte range of the 57 GB file, line by line.
- 2
Extract with regular expressions
Tweet id, author, date, place and text, without parsing the JSON.
- 3
Clean and score
Strip mentions, hashtags and links; tokenise, lemmatise, VADER; bucket the compound score into 1-9.
- 4
Geocode
Normalise the place name and match word combinations against a dictionary of 2021 suburbs.
- 5
Load CouchDB
Bulk uploads of 1,000 documents into a three-node cluster; one database for geotagged tweets.
- 6
MapReduce
Views emit each tweet's score by suburb (all tweets, income keywords, crime keywords) with a _stats reduce.
- 7
Serve
Flask turned the views into gzipped Plotly figures; React drew them beside the SUDO maps.
The cloud is gone, but the steps that produced numbers are preserved in coursework/ and re-run by the scripts in scripts/. You can run steps 3-4 yourself on the pipeline page.
2026
What the revival added
A read-only database. scripts/build_analytics.py re-runs the Flask backend's functions on the saved CouchDB view exports and SUDO files, asserts that the results match the shipped Plotly figures, and writes analytics.db (1.7 MB) plus simplified TopoJSON boundaries.
A crosswalk. Tweets were geocoded to 2021 suburbs (SAL) while income uses 2016 SA2s and crime uses 2019 LGAs. Each suburb is assigned to the SA2 and LGA containing its representative point, and counts and score sums are pooled. All 1,085 Victorian suburbs with tweets fall inside an SA2 and an LGA. This step is new: in 2023 the comparison was visual only.
Correlations. Pearson r, Spearman ρ and a least-squares line at minimum-tweet thresholds of 1, 5, 10 and 30, computed with scipy at build time and again in your browser by a TypeScript port that is unit-tested against the scipy values.
The NLP pipeline in the browser. NLTK 3.8.1's Punkt sentence splitter, word tokeniser, WordNet noun lemmatiser and VADER, plus the BeautifulSoup HTML-to-text step the Mastodon harvester used, ported to TypeScript and verified to reproduce the original Python on every test post and on 44,156 real toots.
Reproduction
Checks against the 2023 outputs
Rows marked exact are asserted by the build script or by the test suite, which CI runs on every push. Rows marked close are reported for transparency.
Notebook example normalises to "Hello ~ What a good weather ! hahahhah , lmao ! ! …" and the raw string scores 8
exactReproduced: Identical string and score in the TypeScript port (the processor scored the normalised text, which gives 9)
Sentimental Analysis.ipynb
Every count and rounded average on the 1,085-suburb Twitter map (3 topics × 2 measures)
exactReproduced: All 6,510 values identical
twitter_vic_sal_2022_02_2022_07.json.gz
2,418,621 tweets contained geographical data
closeReproduced: 2,418,617 in the saved CouchDB view (4 fewer)
Report 4.1
420 Victorian SA2s after outlier removal; Merbein lowest ($28,996), Sydenham highest ($62,029)
exactReproduced: Reproduced with the team's IQR rule in Python and TypeScript
Report 6.2.1
Income tweets: 73% of regions with 1-10, 22% with 10-100, 4.7% over 100; Melbourne 23,281
exactReproduced: 73.2% / 22.0% / 4.7% (bins < 10, 10-99, ≥ 100); 23,281
Report 6.2.2
72 LGAs after outlier removal, Brimbank highest
exactReproduced: Same 72 LGAs as the original map; Brimbank 14,670 offences
Report 6.3.1, crime map
Crime tweets: 79% / 18% / 3% of regions; Melbourne 4,026, Ballarat 485
exactReproduced: 79.4% / 17.6% / 2.9% (bins ≤ 10, 11-100, > 100); 4,026; 485
Report 6.3.2
All six offence divisions for 79 LGAs
exactReproduced: All 474 bar heights identical
crime_vic_lga_bar.json.gz
Rest of WA highest mean (71.5k), ACT highest median (60.2k), Rest of Vic. lowest (49.6k / 41.4k)
exactReproduced: Reproduced from 2,234 SA2s grouped by GCCSA
SUDO summary
mastodon.social scores packed around 5
closeReproduced: Different week re-scored: 61% neutral vs 79% (still the most neutral server)
Report 6.1, dashboard histogram
2026 upgrade · rigour
How uncertainty is reported, and how the statistics are checked
Every quantitative result added in the upgrade carries a sample size and an interval, resampling uses a fixed seed (57, for Team 57) that the page displays, and every TypeScript statistic is compared in CI with the standard Python implementation by scripts/verify_stats.py.
| Quantity | Uncertainty shown | Verified against |
|---|---|---|
| An area's average sentiment | Tweet count and a t-based 95% interval from the CouchDB _stats sums (count, sum, sum of squares); areas under 30 tweets flagged or suppressed (DR-003) | scipy.stats.t |
| Proportions (benchmark accuracy, refusal rate, neutral share by language) | Wilson score 95% interval, with k and n shown | statsmodels proportion_confint |
| Income or offences against sentiment | Spearman's rho with a paired percentile bootstrap (2,000 resamples, seed 57); OLS slope with HC3 robust standard error; residual Moran's I | scipy.stats.bootstrap, statsmodels OLS |
| Spatial clustering | Global Moran's I with analytic moments and a 999-permutation p-value (seed 57); local Moran's I with conditional permutation, two-sided, optional Benjamini-Hochberg | PySAL esda Moran, Moran_Local |
| Neighbours | 6 nearest representative points, or shared borders read from the TopoJSON arcs (islands dropped and counted) | libpysal KNN, Rook |
| Two AI models compared | Paired by question; exact McNemar test on the discordant questions, and the accuracy difference with a Tango score 95% interval | statsmodels mcnemar(exact=True), R PropCIs scoreci.mp |
| Model latency | Median with a percentile bootstrap interval | numpy percentile |
Analytic quantities agree with the Python reference to between 1e-9 and 1e-12; permutation and bootstrap quantities, which use different random streams, agree within Monte Carlo error. Contiguity built from the site's own TopoJSON matches libpysal's rook neighbours on the same files exactly (2,472 SA2 links, 410 LGA links). Ties in permutation tests count against significance.
Comparisons are reported, not ranked: the spatial page shows a sensitivity table across thresholds, neighbour definitions and area units rather than the most striking combination, and nothing is corrected for the number of looks, which is said where it matters.
Areas, not people
Spatial statistics and their caveats
The spatial statistics page tests whether neighbouring areas share a mood (global and local Moran's I). Its main result is negative: the clustering that appears when every area with a tweet counts disappears once averages built on fewer than 30 tweets are suppressed (DR-003).
Caveats
Limitations worth knowing
Different years
Income is 2015-16, offences 2019, tweets 2022. The comparison assumes regional patterns are stable.
Self-reported places
Twitter's place field is often a city, not a suburb: "Melbourne, Victoria" lands in the CBD suburb, inflating its counts.
The shortest match wins
The geocoder tries word combinations shortest first and keeps the first known place, so "Albert Park, Victoria" lands in Albert (NSW), "St Kilda, Victoria" in St Kilda (SA), and Port, North and South Melbourne in Melbourne. 579 of the 3,118 Victorian place keys are shadowed this way and 432 of 2,921 Victorian suburbs can never be matched, which is why St Kilda and Albert Park are missing from the map and why Port Phillip loses tweets to the City of Melbourne. The port keeps this behaviour.
Small samples
Most suburbs have a handful of topic tweets, so their averages are noisy. The threshold slider exists for this reason.
A lexicon, not a reader
VADER scores words, not meaning; sarcasm, slang and non-English text are scored poorly or as neutral.
The outlier rule
The IQR filter removes the busiest LGAs (Melbourne, Casey, Geelong), which are also where most crime tweets are.
Correlation only
Even a clear correlation between areas would say nothing about individuals (the ecological fallacy).
COMP90024 Cluster and Cloud Computing · The University of Melbourne · Semester 1, 2023. Key figures: 37,823,414 tweets processed, 1,659,690 toots harvested. Source code and original submission on GitHub.
Assumptions
What the analysis takes for granted
Stable geography of mood
Income (2015-16), offences (2019) and tweets (2022) are compared as if each area's position had not changed.
Independent tweets
Standard errors treat every tweet as a separate draw. Prolific accounts and repeated posts break this, so intervals are optimistic. No user identifiers were kept, so it cannot be corrected.
Place tags mean place
A tweet's self-reported place is where it is counted, whatever the author's home or the subject of the tweet.
A suburb belongs to one SA2 and one LGA
Each suburb is assigned by its representative point; suburbs that straddle boundaries are not split.
Neighbours
Spatial weights are six nearest areas by default, or shared borders; results are shown under both.
The team's outlier rule
Scenario relationships use the areas the 2023 IQR filter kept, unless a reader brings the outliers back.
AI use statement
What AI does on this site, and what it never does
AI is optional here and runs only on a visitor's own key. This statement is informed by the transparency principles of the Australian Government's policy for the responsible use of AI in government, the EU AI Act's transparency obligations and the NIST AI Risk Management Framework. It does not claim compliance with any of them.
What it does
On /ask, a language model the visitor chooses turns a question into one SQL query and then explains the returned rows in up to three sentences that cite them. On /ask/eval, the same model answers 16 benchmark questions so the visitor can measure it. That is all.
What it never does
It never produced any number elsewhere on the site: every chart, statistic and finding comes from the 2023 pipeline and the revival's tested code. It never touches the database directly, never writes data, never runs on this site's servers and never sees a key belonging to the project (there is none).
Data sent to the provider
From the visitor's browser, with their key: the question exactly as typed, the documented schema, the generated SQL and up to 30 result rows of public aggregates. The database holds no personal data, but anything typed into the question is sent as is, so do not include personal information. This site's server receives only the SQL.
Human in the loop
Every output is labelled AI-generated. The SQL is shown and editable, the validator's verdict and the rows are shown, citations are checked against the rows, and the person records a decision: accepted, edited or rejected.
Records
Each question (two model calls: writing the SQL, then explaining the rows) is logged as one record in the visitor's browser (IndexedDB): id, time, feature, provider, model, question, generated SQL, what the explanation call was sent (the SQL and the number of rows), validator verdict, row count, answer, end-to-end latency, token usage and the decision with its time. Each benchmark question is one record. Re-running edited SQL adds a new record linked to the original; once written, only the decision changes. Viewable and exportable as JSON or CSV at /ai-log. Keys are redacted before every write.
Measurement and limits
The benchmark reports accuracy with Wilson intervals, counts malformed or truncated replies as wrong, and compares models with a paired exact test and a score interval for the difference. The project publishes no AI scores of its own. Known failure modes are listed in the model card.
Model card
Two models, documented
The 2023 sentiment scorer (VADER in NLTK, wrapped in the team's pipeline) and the optional text-to-SQL assistant each get intended use, provenance, evaluation, failure modes and ethical considerations.
Sentiment scorer
Reproduced exactly by the TypeScript port (0 mismatches in 44,156 real toots; Wilson upper bound 0.009%). Accuracy against human judgement was not measured here. Its clearest failure is language: in the re-scored Mastodon week, 45% of English toots score neutral against 98% of Japanese ones.
Ask the data
A third-party model chosen and paid for by the visitor, contained by server-side validation and read-only execution, measured by a 16-question benchmark the visitor runs. No accuracy is claimed.
Decision records
The decisions behind the revival and the upgrade
Each record states the decision first, the options and the reasons, then what actually happened, including the weak numbers, and what I would change. Records are never edited to change a decision; a new record supersedes an old one.
DR-001
Rebuild the analysis from the team's code and use the shipped Plotly figures as the answer key
accepted · October 2026, during the revival (recorded on 6 October 2026)
DR-002
Ship the data as a read-only SQLite file inside the serverless functions
accepted; extended on 6 October 2026 for model-written SQL · October 2026, during the revival (recorded on 6 October 2026)
DR-003
Suppress area averages built on fewer than 30 tweets from the spatial statistics
accepted · 6 October 2026
DR-004
accepted · 6 October 2026
Next
What I'd change
Shrink, don't suppress
Replace the 30-tweet cut with partial pooling (empirical Bayes or a multilevel model) so small areas borrow strength from their neighbours instead of disappearing.
Measure the scorer
Hand-label a stratified sample of a few hundred posts to estimate VADER's agreement with people on this data, with intervals, before reading anything into small differences in tone.
Effective sample size
Keep a count of distinct accounts per area in any future pipeline, so intervals can account for prolific posters.
Pre-register the comparisons
Fix the thresholds and the tests before looking at the data, instead of reporting a sensitivity table afterwards.
A bigger text-to-SQL benchmark
Fifty or more questions with a held-out half, so a prompt cannot be tuned to the test, and a tamper-evident audit log.
Rate-limit the SQL endpoint
The validator bounds each query, but nothing bounds how many arrive.