Scenario 2 · Crime
Where crime is recorded, does the conversation turn darker?
The second question paired Victoria Police offence counts with tweets that mention crime, police, theft, robbery, arrests, murder or violence. Crime talk is rare and overwhelmingly negative, but is it more negative where more offences are recorded?
- LGAs analysed
- 72
- 7 removed as outliers
- Most offences (Brimbank)
- 14,670
- among the 72, reference year 2019
- Crime-related tweets in Victoria
- 5,133
- of 17,796 Australia-wide
- Mean score of crime tweets
- 3.53
- vs 5.54 for all tweets (1-9)
Explore
Offences on the map, sentiment on the chart
The offence axis is logarithmic because counts span two orders of magnitude, from a few hundred in alpine shires to tens of thousands in the inner city. Tick the box to bring back the seven LGAs the team's outlier rule removed, including the City of Melbourne with its 4,026 crime tweets.
- Pearson r
- 0.31
- p = 0.03
- Spearman ρ
- 0.27
- p = 0.06
- R²
- 0.097
- of variance
- Regions
- 49
- in the fit
Computed in your browser; identical to the scipy values stored by the build script.
49 LGAs with at least 1 crime-related tweet: a moderate positive association (the fit line is linear in offences, drawn on a log axis). Most LGAs have only a handful of crime tweets, so their averages swing between extremes; raise the threshold and the association disappears.
Select an LGA on the map or a dot in the chart (or search above) to see its offence mix and tweet sentiment.
2026 upgrade · uncertainty
How sure can we be?
The correlations above are point estimates. Here the same comparison carries a 95% interval, a heteroskedasticity-robust slope and a check for leftover spatial structure, first at the 2023 threshold and then counting only LGAs with at least 30 tweets, where an average can bear some weight.
Crime tweets, 2023 threshold
LGAs kept by the IQR rule with at least one crime tweet (offences on a log scale).
- Areas
- 49
- ≥ 1 tweets
- Spearman ρ
- 0.27
- 95% CI -0.03 to 0.53
- OLS slope
- 1.00
- HC3 95% CI -0.04 to 2.04
- p (HC3)
- 0.06
- robust, two-sided
- R²
- 0.091
- SE 0.461 classical vs 0.530 HC3
- Residual Moran's I
- -0.017
- p = 0.45 (permutation)
The interval for Spearman's rho runs from -0.03 to 0.53 and includes zero: the data are consistent with no association.
All tweets, LGAs with ≥ 30 tweets
Crime tweets are too sparse for this threshold (4 LGAs reach it), so this uses all tweets.
- Areas
- 65
- ≥ 30 tweets
- Spearman ρ
- 0.16
- 95% CI -0.10 to 0.40
- OLS slope
- 0.11
- HC3 95% CI -0.05 to 0.27
- p (HC3)
- 0.18
- robust, two-sided
- R²
- 0.034
- SE 0.074 classical vs 0.081 HC3
- Residual Moran's I
- -0.044
- p = 0.33 (permutation)
The interval for Spearman's rho runs from -0.10 to 0.40 and includes zero: the data are consistent with no association.
Slope: change in average tone (1-9 scale) per tenfold increase in recorded offences. Spearman interval: paired percentile bootstrap (2,000 resamples, seed 57). Slope interval: HC3 heteroskedasticity-robust standard error with a normal reference. A residual Moran's I near zero means the areas' spatial arrangement is not inflating the precision. Each combination of settings is a separate, exploratory comparison; none is corrected for the others.
Spatial statistics: clusters, suppression and the full sensitivity table →
Context
What the 2023 analysis concluded
Revisited with the same data
Two things hold up. Crime tweets are strongly negative: the most common score is 1 (25% of them), against a neutral mode for tweets in general. And the volume ranking matches.
The link between offences and tone is fragile. Across LGAs with any crime tweet, Pearson r is 0.31 (p = 0.03, 49 LGAs), but it vanishes (r = −0.01) once each LGA needs five tweets. At suburb level, more crime talk does not predict a more negative tone (r = −0.05, p = 0.63).
Tone
Crime talk skews hard to the negative
Bars show the share of crime-related tweets at each score; dashed outlines show all geotagged tweets for comparison.
Twitter, Feb-Jul 2022
17,796 crime tweets · 2,418,617 geotagged tweets
| Score | Count | Share | All geotagged tweets share |
|---|---|---|---|
| 1 (Extremely Negative) | 4436 | 24.9% | 1.6% |
| 2 (Very Strongly Negative) | 3205 | 18.0% | 3.3% |
| 3 (Strongly Negative) | 2382 | 13.4% | 5.5% |
| 4 (Negative) | 1457 | 8.2% | 5.0% |
| 5 (Neutral) | 3039 | 17.1% | 47.9% |
| 6 (Positive) | 910 | 5.1% | 7.8% |
| 7 (Strongly Positive) | 1058 | 5.9% | 13.3% |
| 8 (Very Strongly Positive) | 822 | 4.6% | 10.0% |
| 9 (Extremely Positive) | 487 | 2.7% | 5.6% |
Does more crime talk mean a darker tone? (suburbs)
102 Victorian suburbs · slope -0.15 points per tenfold increase · r = −0.05
Offences
The busiest LGAs, including the ones the filter removed
- Melbourne (C)26,694 (removed as outlier)
- Casey (C)16,344 (removed as outlier)
- Greater Geelong (C)15,906 (removed as outlier)
- Hume (C)15,661 (removed as outlier)
- Brimbank (C)14,670
- Greater Dandenong (C)14,234 (removed as outlier)
- Whittlesea (C)11,952
- Wyndham (C)11,413
- Moreland (C)11,376
- Darebin (C)11,267
- Frankston (C)11,165
- Yarra (C)11,154
- Latrobe (C)9,958
- Port Phillip (C)9,691 (removed as outlier)
- Kingston (C)8,646 (removed as outlier)
Grey bars are LGAs the team's IQR rule removed (any of the six offence divisions outside 1.5 times the interquartile range). That rule drops the very places with the most crime and the most crime tweets, which is why the explorer lets you add them back.
Offence divisions follow the Crime Statistics Agency classification: against the person, property and deception, drug, public order and security, justice procedures, and other offences. Missing divisions were filled with zero, as in the original Flask code.
See data and methods for how suburbs were assigned to LGAs.