Open data · analytics.db
Every number on this site, as tables
The app reads a single read-only SQLite file built by scripts/build_analytics.py. It holds 19 tables and 15,828 rows of aggregates: counts and score sums, official statistics and boundary metadata. No tweet or toot text, user names or ids are stored.
Tables
Browse the database
Scenarios
- 457 rows
Scenario 1: income vs sentiment
scenario_income_sa2
Median income joined with suburb tweet sentiment pooled to each Victorian SA2 (revival analysis). avg_* are NULL where an SA2 had no tweets.
10 columns CSV - 79 rows
Scenario 2: crime vs sentiment
scenario_crime_lga
Offence totals joined with suburb tweet sentiment pooled to each Victorian LGA (revival analysis).
9 columns CSV - 19 rows
Scenario correlations
scenario_correlations
Pearson, Spearman and least-squares fits at several minimum-tweet thresholds (computed with scipy).
14 columns CSV
Sentiment
- 81 rows
Sentiment histograms (2023 dashboard)
sentiment_histogram
Counts of 1-9 sentiment scores for Twitter and the three Mastodon servers, as plotted in 2023. Twitter covers geotagged tweets Australia-wide, Feb-Jul 2022.
4 columns CSV - 6,799 rows
Twitter sentiment by suburb (SAL)
twitter_sal_sentiment
CouchDB MapReduce _stats per suburb and topic (all tweets, income keywords, crime keywords), Feb-Jul 2022, every Australian state. Rows are keyed by SAL code: join regions_sal for names.
9 columns CSV - 3 rows
Mastodon servers
mastodon_servers
The three servers the harvesters followed.
4 columns CSV - 27 rows
mastodon.social re-scored histogram
mastodon_rescored_histogram
The surviving week of mastodon.social toots (1-9 May 2023) re-scored with the original pipeline (aggregates only).
4 columns CSV - 144 rows
mastodon.social by hour
mastodon_hourly
Re-scored mastodon.social toots per UTC hour, 1-9 May 2023, with score sums and 1-9 bucket counts.
16 columns CSV - 112 rows
mastodon.social by language
mastodon_language
Re-scored mastodon.social toots (1-9 May 2023) per declared language, with score sums and bucket counts.
14 columns CSV
SUDO
- 2,234 rows
Personal income by SA2
income_sa2
SUDO / ABS personal income 2015-16 (mean, median, sum, median age of earners) for every Australian SA2.
13 columns CSV - 15 rows
Personal income by capital city area
income_gcc
The SA2 rows grouped by Greater Capital City Statistical Area, as in the original summary.
7 columns CSV - 40 rows
Jobs and income by industry
jobs_income_indicators
ABS Jobs in Australia 2018-19 indicators (mean, sd and median across SA2s) from the original bar chart.
6 columns CSV - 79 rows
Recorded offences by LGA
crime_lga
Victorian Crime Statistics Agency offence divisions per LGA (reference year 2019).
11 columns CSV
Geography
- 5,155 rows
Suburbs and localities (SAL 2021)
regions_sal
ABS suburbs and localities that had tweets, every state, with a representative point and (Victoria only) the SA2 and LGA they fall in.
9 columns CSV - 462 rows
Victorian SA2s (2016)
regions_sa2
ABS Statistical Area Level 2 regions in Victoria.
8 columns CSV - 80 rows
Victorian LGAs (2019)
regions_lga
ABS Local Government Areas in Victoria.
5 columns CSV
Context
- 18 rows
Facts
facts
Headline numbers quoted in the team report, each with its source section.
4 columns CSV - 7 rows
Build metadata
meta
Editions of the boundaries and how the database was built.
2 columns CSV - 17 rows
Original summaries
summary_text
The paragraphs the 2023 dashboard showed beside each chart (written by the team).
4 columns CSV
Provenance and licences: Twitter aggregates derive from the University-provided corpus (Australian Data Observatory); SUDO income and offence figures are ABS and Crime Statistics Agency Victoria data; ABS boundaries are CC BY 4.0. The whole file is in the repository at web/data/analytics.db and can be opened with any SQLite browser.