Guided tour
The project in three short walkthroughs
Each video follows one workflow from start to finish, with the step shown on screen and as captions. A Playwright script recorded them from this site and checked every step on the way (the recovered 2023 numbers, the live correlation, Moran's I at two thresholds, the validator verdict, the audit-log record), so a broken feature fails the recording instead of producing a misleading video.
Walkthrough 1 of 3 · the landing page
The story
The landing page from top to bottom: the question, Victoria's geotagged tweets drawn as suburb dots, the 2023 brief, the original cloud system animated step by step as you scroll, the four recomputed findings and the 2026 upgrade.
Steps (transcript)
- 1The question: does the mood online match life on the ground?
- 2Every dot is a Victorian suburb: size is tweet volume, colour is average tone
- 3The 2023 brief: harvest, store, analyse and compare with official data, on 8 vCPUs
- 4How it worked: scroll to follow a tweet through the original cloud system
- 5Harvest: MPI ranks split the 57 GB file; harvesters poll three Mastodon servers
- 6Score and store: VADER on a 1-9 scale, bulk-loaded into a three-node CouchDB cluster
- 7Summarise: MapReduce views keep a count, sum, min and max for every suburb
- 8Serve and compare: a Flask API joins the views with SUDO income and crime data
- 9Automate: Ansible and Docker Swarm on the Melbourne Research Cloud
- 10Four findings, recomputed from the recovered numbers
- 11The 2026 upgrade: intervals, spatial statistics and an optional, audited AI
Walkthrough 2 of 3 · /income → /spatial
Income vs sentiment
Scenario 1 end to end: the SA2 income choropleth linked to a scatter of tweet sentiment, the map recoloured by mood, a selected area with its interval, the correlation recomputed live at higher thresholds, the bootstrap interval, then Moran's I and the LISA cluster map on /spatial, and the caveats that come with them.
Steps (transcript)
- 1Scenario 1: do richer areas tweet more happily about money?
- 2The choropleth: median personal income for 420 Victorian SA2s (ABS, 2015-16)
- 3The linked scatter: one dot per SA2, sized by its income tweets
- 4Recolour the map by the mood of income tweets instead
- 5Select Ballarat: the map, the chart and the detail panel follow, with a 95% interval
- 6Raise the tweet threshold: the correlation is recomputed in the browser
- 7How sure? Spearman's rho 0.10, bootstrap 95% CI -0.06 to 0.26, which includes zero
- 8Spatial statistics: is the mood clustered in space, or is it noise?
- 9Global Moran's I = 0.022, permutation p = 0.24 across 152 SA2s: no evidence of clustering
- 10The LISA cluster map: High-High, Low-Low and outliers, 999 permutations, seed 57
- 11Count every SA2 with a tweet and a 'pattern' appears: I = 0.135, p = 0.001
- 12Back to 30 tweets and zoom to Melbourne, where most analysed SA2s are
- 13The caveats: areas, not people; boundaries change answers; small areas are suppressed
Walkthrough 3 of 3 · /ask → /ai-log → /records
Ask the data
The optional bring-your-own-key feature: the AI settings dialog, a question answered by a mocked model (no real key is used), the generated SQL passing the site's validator, the rows and chart, the human decision, then the audit log and the same rows in the open records.
Mocked AI response for illustration. Steps 3, 5, 7, 9 and 10 show a reply from a mock, not a model: the key is a placeholder and requests to the provider are intercepted in the browser. They show how the feature labels, validates, cites and logs a reply, not what a real model would say.
Steps (transcript)
- 1Ask the data: optional AI with your own key; the rest of the site needs none
- 2AI settings: Anthropic by default, the key stays in this tab and goes only to the provider
- 3For this demo: a placeholder, not a real key; provider calls are interceptedMocked AI response for illustration
- 4Ask an example question: the five Victorian suburbs with the most crime tweets
- 5Step 1: the model writes one SQL query, labelled AI-generatedMocked AI response for illustration
- 6Step 2: the site's validator allows one read-only SELECT and runs it
- 7Step 3: the explanation cites rows, and every citation is checkedMocked AI response for illustration
- 8The rows themselves, with an automatic chart
- 9A person decides: accept or reject, and the decision is recordedMocked AI response for illustration
- 10The AI audit log: question, model, SQL, verdict, latency and decision; JSON or CSVMocked AI response for illustration
- 11Records: Melbourne (SAL 21640) in the open table shows the same 4,026 crime tweets
Screenshots
Every key feature at a glance
Captured by the same script in light mode at 1440 × 900 (the landing page also in dark mode) and on a 390 px phone. Select one to enlarge it; the arrow keys step through the set.
Desktop · 1440 × 900
Mobile · 390 × 844
How these were made
pnpm showcase runs web/e2e/showcase.spec.ts on the system Chrome. It plays each journey at a human pace with an on-screen caption and a visible cursor, asserts what it shows, and records it at 1280 × 800; ffmpeg then encodes the H.264 videos on this page and the GIFs in the README. The captions, the step lists here and the README walkthrough are the same text. The statistics use the site's fixed seed 57, so a re-run shows the same numbers.
No real API key is used anywhere in these recordings. Where the AI feature appears, the key is a placeholder, every request to the provider is intercepted in the browser, and the reply is a labelled mock whose text starts with “Mocked response for illustration.”. The SQL it returns still goes through this site's real validator and runs on the real read-only database.