Australian team publishes the 5W-PL protocol: linking 14 Victorian health and human services datasets to trace mental health service pathways for young people aged 12–25
Synopsis
The protocol describes the design of the 5W-PL study: using the Centre for Victorian Data Linkage's Victorian Linkage Map to link 14 datasets covering mortality, mental health, hospital, emergency, ambulance and human services (child protection, disability, sexual assault, homelessness, alcohol and drug) through deterministic and probabilistic methods, building a cross-sector longitudinal cohort of people born between 1970 and 2010 with service-use records from 2015 to 2025, in order to identify subgroups by care pathway, map geographical patterns of service need, and develop predictive models of service type and intensity for young people aged 12–25.
A table listing datasets sourced for a study, detailing their names, abbreviations, descriptions, available data, and data collection start dates.
PubMedInterpretation
The protocol proposes cross-sector administrative data linkage as an alternative to traditional cohorts for characterising real-world pathways through youth mental health care. Existing linkage studies involving young people and mental health service use have largely focused on associations between early mental health and later outcomes rather than explicitly modelling cross-sector service utilisation pathways or continuity of care; 5W-PL explicitly targets pathways, discontinuities and the 'missing middle'. This is a study protocol; the text states that 'no datasets were yet generated or analysed for the purpose of the current study', so there is no result-level evidence.
The protocol integrates 14 datasets spanning mortality, mental health, health, emergency services and human services. The text notes that Australia lacks a universal population identifier across sectors, so linkage studies require bespoke approvals and coordinated infrastructure; this protocol uses the CVDL's Victorian Linkage Map and Integrated Data Resource for cross-sector integration and handles ambulance data separately because it is not held by the CVDL. Each dataset's name, abbreviation, content and start date are listed in a table, which is a verifiable protocol-level description.
The protocol plans to use natural language processing and machine learning to extract features from free-text ambulance clinical notes, and to analyse care pathways with clustering, spatial time-series and causal inference frameworks. The text emphasises that ambulance data contain both structured fields and free text, and that the unstructured text can help understand the underlying causes and trajectories of acute mental health service use, which is otherwise not feasible with structured datasets alone. Specific techniques are named (e.g. BERT, sequential pattern mining, deep learning and multi-view clustering, random forest and XGBoost causal forest), but all are planned rather than implemented.
The protocol embeds youth participation and stakeholder consultation throughout, and sets out confidentialisation, data suppression and secure storage procedures. Orygen's Youth Research Council provided input on the acceptability of the study design, and a dedicated 5W youth advisory group will assist with translation; data undergo confidentialisation with new unique identifiers, and reporting follows ABS data suppression rules. Ethics approval (University of Melbourne HREC ID 21041) and data security arrangements (CVDL cloud platform, multi-factor authentication, isolated Secure Research Environment) are described in the text.
Perspective
The protocol targets young people in Victoria aged 12–25 as of January 2018 who have been in contact with any of the included services, with younger (<12) and older (26–53) groups as comparators; data cover 1 January 2015 to 31 December 2025, with some datasets ending between June 2019 and December 2025, enabling comparison of service use before, during and after COVID-19. The authors state the protocol could be adapted to other Australian states and territories and is relevant to other countries seeking to integrate health and human services data to address fragmentation. The intended setting is state-level, multi-sector administrative data linkage analysis to identify care-pathway subgroups, geographical patterns of service need, and prediction of service type and intensity from clinical, socioeconomic and environmental factors.
As a protocol article, no data have yet been generated or analysed, so everything about subgroups, geographical patterns and predictive models is planned rather than found. Readers should still watch: cross-sector variables (such as gender and marginalisation-related variables) are often incomplete or inconsistently recorded in administrative data, and the authors state they will harmonise systematically and explicitly account for limitations in data completeness and quality; single-state scope and limited linkage to national datasets may affect capture of full service use, which the authors plan to address through sensitivity analyses and clearly bounded interpretation; in addition, the loaded text is an incomplete version missing some figures and reference details, so checking specific dataset fields or the statistical analysis plan would require the original article.
