Skip to content
Social Sense

2026 upgrade · spatial statistics

Is the mood clustered in space, or is it noise?

Maps of regional averages invite the eye to find patterns. This page tests for them: global and local Moran's I with permutation inference, a minimum sample size before an area counts, and intervals on the two scenario relationships.

Moran's I, SA2s with ≥ 30 tweets
0.022
152 SA2s · one-sided permutation p = 0.24
Same, counting every SA2 with a tweet
0.135
296 SA2s · one-sided permutation p = 0.001
Spread of true SA2 means vs single tweets
0.18 vs 1.74
points on the 1-9 scale
Tweets an SA2 needs to be half signal
89
100 of 296 SA2s get there

Explore

Clusters, outliers and how much each area's average can bear

Areas below the tweet threshold are suppressed: hatched on the map and left out of every statistic. Raise it and watch the “pattern” at low thresholds fade. Everything recomputes in your browser with the same seed as the server, so the numbers are reproducible.

Areas
Sentiment of
Minimum tweets (smaller areas suppressed)≥ 30
Neighbours
Find

Global Moran's I

I = 0.022(expected if random: -0.007)

permutation p = 0.24 (one-sided, 999 permutations, seed 57) · normal approximation p = 0.50 (two-sided) · z = 0.67

No evidence that neighbouring areas are more alike in tone than chance would produce. Even where significant, a value this small means weak clustering.

Analysed
152
≥ 30 tweets
Suppressed
144
< 30 tweets
No tweets
166
left out
Islands
0
6 nearest neighbours
Loading map…
Most analysed SA2s are in Melbourne.
  • High-High (3)
  • Low-Low (2)
  • High-Low (0)
  • Low-High (6)
  • Not significant (141)
  • Suppressed, island or no tweets

Local Moran's I with conditional permutation (999 draws, seed 57), two-sided pseudo p < 0.05, not adjusted for multiple testing. With 152 tests at 5%, about 8 “significant” areas are expected by chance alone.

-2.0-1.00.01.02.0-2.0-1.00.01.02.03.0Area's tone (standardised)Neighbours' average (lag)SouthbankGeelongBallarat

The slope of the dashed line is Moran's I. Points top-right and bottom-left are areas that resemble their neighbours; dot size shows how many tweets an area's average rests on.

Select an area on the map or in the plot to see its sample size, interval and local statistic.

Signal or noise?

Individual tweet scores in a SA2 vary with a standard deviation of 1.74 points, while the true SA2 averages differ by about 0.18 (method-of-moments estimate across 269 areas). An average needs about 89 tweets before half of its variation is real difference rather than sampling noise.

Tweets are treated as independent; prolific accounts make the effective sample smaller, so these are optimistic. The suppression threshold (default 30) is explained in decision record DR-003.

Per-area sample sizes and intervals
SA2TweetsMean [95% CI]ReliabilityStatus
587,0885.54 [5.53, 5.54]1.00Low-High
21,7355.49 [5.46, 5.51]1.00analysed
18,9295.19 [5.17, 5.22]1.00analysed
8,3065.65 [5.61, 5.68]0.99analysed
4,9205.77 [5.73, 5.82]0.98analysed
3,4045.72 [5.66, 5.79]0.97analysed
3,3235.66 [5.60, 5.72]0.97analysed
3,2915.55 [5.50, 5.60]0.97analysed
3,1085.47 [5.41, 5.52]0.97analysed
2,9295.56 [5.51, 5.61]0.97analysed
2,8605.72 [5.65, 5.78]0.97analysed
2,6966.07 [5.99, 6.14]0.97analysed

Scenario 1, with uncertainty

Median income and the tone of all tweets

Areas
147
≥ 30 tweets
Spearman ρ
0.20
95% CI 0.04 to 0.36
OLS slope
0.13
HC3 95% CI 0.03 to 0.22
p (HC3)
0.01
robust, two-sided
R²
0.039
SE 0.052 classical vs 0.049 HC3
Residual Moran's I
0.020
p = 0.26 (permutation)

Spearman's rho of 0.20 (0.04 to 0.36) excludes zero, but the association is weak; the slope's robust interval runs from 0.03 to 0.22 points on the 1-9 scale.

Slope: change in average tone (1-9 scale) per $10,000 of median income. Spearman interval: paired percentile bootstrap (2,000 resamples, seed 57). Slope interval: HC3 heteroskedasticity-robust standard error with a normal reference. A residual Moran's I near zero means the areas' spatial arrangement is not inflating the precision. Each combination of settings is a separate, exploratory comparison; none is corrected for the others.

Sensitivity

How much the answer depends on the choices

The same tweets, analysed with different minimum sample sizes, neighbour definitions and area units. A finding that only appears under one combination is not a finding.

Global Moran's I of average tone (all tweets); one-sided permutation p, 999 permutations, seed 57
AreasMin tweetsNeighboursnIp (one-sided)
SA2≥ 16 nearest2960.1350.001
SA2≥ 1border (8 islands)2880.1640.001
SA2≥ 106 nearest2120.0640.04
SA2≥ 10border (9 islands)2030.0910.06
SA2≥ 306 nearest1520.0220.24
SA2≥ 30border (17 islands)1350.0070.43
SA2≥ 1006 nearest97-0.0620.15
SA2≥ 100border (19 islands)78-0.0280.44
LGA≥ 16 nearest790.1360.01
LGA≥ 1border790.1090.05
LGA≥ 106 nearest760.0670.09
LGA≥ 10border760.0830.09
LGA≥ 306 nearest72-0.0190.49
LGA≥ 30border720.0070.37
LGA≥ 1006 nearest61-0.0830.12
LGA≥ 100border (1 island)60-0.0440.38
Scenario relationships (team's IQR outliers removed): Spearman's ρ with a 95% bootstrap interval, and the residual Moran's I of the OLS fit
ComparisonMinnρ [95% CI]Resid. I
Income vs all tweets (SA2)≥ 12790.10 [-0.02, 0.22]0.125
Income vs all tweets (SA2)≥ 102010.22 [0.08, 0.36]0.010
Income vs all tweets (SA2)≥ 301470.20 [0.04, 0.36]0.020
Income vs all tweets (SA2)≥ 100940.15 [-0.06, 0.34]-0.065
Income vs income tweets (SA2)≥ 11670.10 [-0.06, 0.26]-0.021
Income vs income tweets (SA2)≥ 10690.09 [-0.14, 0.31]-0.022
Income vs income tweets (SA2)≥ 30330.00 [-0.35, 0.34]-0.122
Offences vs all tweets (LGA)≥ 1710.26 [0.02, 0.47]0.050
Offences vs all tweets (LGA)≥ 10680.22 [-0.02, 0.45]-0.003
Offences vs all tweets (LGA)≥ 30650.16 [-0.10, 0.40]-0.044
Offences vs all tweets (LGA)≥ 100550.11 [-0.19, 0.39]-0.068
Offences vs crime tweets (LGA)≥ 1490.27 [-0.03, 0.53]-0.017
Offences vs crime tweets (LGA)≥ 5280.05 [-0.34, 0.45]0.150
Offences vs crime tweets (LGA)≥ 1017-0.03 [-0.53, 0.46]-0.071

The Moran's I p-values are one-sided, in the direction of the observed I: they test for clustering when I is above its expectation and for dispersion when it is below (the cluster map uses two-sided local p-values). These are many looks at one dataset and no correction for multiple comparisons is applied, so an interval that excludes zero is a lead to check, not a result. Intervals excluding zero: income vs all tweets (SA2) at ≥ 10 and ≥ 30 tweets; offences vs all tweets (LGA) at ≥ 1 tweet. None holds at every threshold. All are weak (ρ under 0.3). The 2023 conclusions stand: no meaningful link between income and the tone of income tweets, and none between recorded offences and the tone of crime tweets.

Read this first

What these statistics can and cannot say

Small areas

Most SA2s have a handful of geotagged tweets, and an average of five tweets can swing by two points on its own. Areas under 30 tweets are suppressed by default (DR-003); the scenario pages keep the 2023 thresholds but now show each area's sample size and interval.

Verified, not just computed

The TypeScript statistics are checked in CI against PySAL (libpysal, esda), statsmodels and scipy: neighbour sets identical, Moran's I and its moments to 1e-9, local statistics to 1e-10, and permutation p-values within Monte Carlo error. See scripts/verify_stats.py and the methods page.