Adaptive Optics Image Analysis Using Generative AI (GPT-4) Scripting: An Exploratory Study
Synopsis
This exploratory study used GPT-4 to generate and iteratively fine-tune an R script for preprocessing and blob identification/counting in adaptive optics flood illumination ophthalmoscopy (AO-FIO) images, debugging it on images from 4 participants (1 healthy individual and 3 patients with Stargardt disease), having another researcher naive to the prior coding check the code for errors using a different test set of images from 4 other participants (1 healthy individual and 3 patients with Stargardt disease), and comparing cone counts from 5 AO image snippets with counts independently recorded by 2 human graders and those from pre-existing AO analysis software, with the authors positioning the script as functional but nonvalidated.
Interpretation
The study demonstrates the feasibility of using a widely available generative AI (GPT-4) to produce an R script for AO-FIO image analysis, yielding after 54 iterations of instructions a functional script that identified and quantified blobs. Dedicated image analysis scripts for AO-FIO were limited, especially for large-scale measurements and analyses of nonhealthy images; this work brings a large language model into the script development loop as a proof of principle for generating a functional script. It is an exploratory proof of principle: the script was fine-tuned iteratively by trial and error and tested for image preprocessing and analysis on images from 4 participants, and the authors explicitly describe it as functional but nonvalidated.
The script code was checked for errors by another researcher who was naive to the previous coding, using a different test set of AO-FIO images from 4 other participants (1 healthy individual and 3 patients with Stargardt disease). Beyond generating the script, the work adds an independent code-checking step using a separate test set to examine whether the script runs. An independent checker and a separate test set were used, with a sample of 4 participants, making this an initial verification step.
Cone counts from 5 AO image snippets were compared with counts independently recorded by 2 human graders and with counts generated by pre-existing AO analysis software trained on healthy participants. The generated script's output is placed alongside human grading and an existing software's output, providing reference points for later validation. The comparison rests on 5 image snippets, 2 human graders, and one pre-existing software package; the sample is limited, and the authors do not report a quantitative agreement result from this comparison.
Perspective
The script is positioned as a functional but nonvalidated preliminary tool for AO-FIO images in an R environment, suited to exploring the feasibility of generative AI-assisted script development; the authors state that before clinical research use it should undergo more fine-tuning and extensive testing, and that future work should enhance the script's image analysis capabilities and validate its results to assess the potential of AO-based cone counts as biomarkers in clinical trials.
A careful reader may still wonder how the script performs on larger and more varied image sets; what the quantitative agreement results are from the comparison with human graders and pre-existing software; how robust the script is on nonhealthy images such as those from Stargardt disease; and what validation path is needed to move from a preliminary script to a tool usable in clinical research. The loaded text is summary-level and does not include figures or full methodological detail, so some of these questions cannot be answered from the available text.
