I tested the 'Scratching the Surface' methodology on 200 Australian family law judgments. Here is what I found.

First published on LinkedIn, 11 June 2026. Read the original.

Last week Right to Equality published Scratching the Surface: Victim-Blaming and Bias in Family Court Judgments (Hayton, Quinlan and Sayer, June 2026). The study applied a taxonomy of victim-blaming language, developed by herEthical AI, to 91 published family law judgments of England and Wales. It found that 72.5 per cent of judgments contained at least one instance of judicial victim-blaming, most of it subtle and embedded in reasoning rather than overt.

I read it and wanted to know whether the same patterns appear in Australian judgments. I maintain a knowledge base that includes 28,000 Australian family law decisions, so I had the corpus to test it at scale. I took the report's published taxonomy, severity scale and coding rules, turned them into a written protocol, and applied them to 200 published judgments of the federal family courts: 160 with a family violence catchword and 40 parenting and property judgments without one as a comparison set, decided mainly between 2020 and 2026.

The results track the England and Wales findings closely. 70 per cent of the 200 judgments contained at least one instance of victim-blaming language attributable to a court professional, against 72.5 per cent in the source study. As in England and Wales, the language was overwhelmingly subtle or moderate: of 724 court-professional instances, only 2 per cent were coded obvious. In both jurisdictions the same mechanism dominated by a wide margin:

Discrediting a person's account of abuse through expected-victim-behaviour reasoning.

She did not report to police or tell her doctor. She stayed in the relationship.

As an educated, capable professional, she would surely have spoken up.

The source study found this reasoning in English judgments; it appears in Australian ones in near-identical form.

There were two differences. In the Australian sample, role reversal (the parent raising abuse allegations recast as the aggressor, often through alienation framing) and mutualising (one-directional violence described as 'high conflict' or a 'volatile relationship') ranked higher than in the English data, which may reflect the prominence of alienation and high-conflict framing in Australian parenting litigation. The second difference is concentration. In about two thirds of judgments the court's overall reasoning was protective of the person claiming victimisation, and 30 per cent of judgments contained no court-professional instance at all. A minority of judgments carry dense patterns, while the median judgment carries three instances. This is not evidence of uniform judicial bias.

The comparison stratum produced a finding I did not expect. Judgments without a family violence catchword still showed court-professional instances in 45 per cent of cases, and some of the densest judgments in the whole sample came from that stratum. Published judgment catchwords under-capture where this language occurs.

The necessary caveats. This was an AI-assisted review of language and framing, and every finding requires verification by a human reviewer against the source judgments before it can be relied on; a human-coded validation subset is the next step and these figures are preliminary until it is done. Classifying language under a published taxonomy is an analytical judgment, not a finding of judicial misconduct or an assessment of whether any decision was correct. Published judgments are a small, non-random fraction of family court decisions. The coders themselves flagged 38 per cent of instances as borderline, which gives a fair indication of how indirect most of this language is.

Beyond the substantive question, the exercise shows that systematic review of judicial language at scale is now practical in Australia. The source study needed a bespoke tool and a research team to read 91 judgments. With a well-built corpus and a disciplined protocol, 200 judgments can be coded, double-coded and aggregated in a day. The bottleneck has moved from reading capacity to validation: the hard work that remains is human review, and that is where it should be.

The full methodology, including the sampling design, reliability measures and limitations, is documented and I am happy to share it with anyone who wants to replicate or critique the approach.