Which LLM is Best for Translating Natural Language Goals to PDDL
Synopsis
This paper designs an iteratively refined prompt template that lets six contemporary large language models translate informal video-game testing goals written in natural language into PDDL goals for classical planning, and systematically evaluates them on 90 natural language goals (45 expressible and 45 inexpressible across eight domains) for correctness, speed, and error tendencies, finding that all models exceed 92% correctness, with Gemini 2.5 Flash highest at 96% and fewest false positives, while GPT-4.1 is fastest.
Figure 2: Comparative Prompt Response Time Analysis of Large Language Models (LLMs). Each line represents a distinct LLM, plotting its prompt response time (y-axis) against the problem index (x-axis) after sorting all prompts by increasing runtime for that model.
arXivInterpretation
It presents and publishes a complete prompt template that fills placeholders with the PDDL domain file, the list of objects and their types, and the natural language goal, and instructs the model to start with "Goal cannot be expressed" and give a short explanation when the goal cannot be expressed. Unlike earlier practice of including the full problem.pddl, the template keeps only essential object definitions and adds examples of inexpressible cases plus a directive to answer simply and concisely. The template is shown in full in Figure 1, and its design rests on the authors' reported iterative experiments: removing the full problem file, adding inexpressible examples, and requesting concise answers each addressed observed accuracy drops, frequent false positives, and overly complex and often incorrect generated goals.
On 90 real-world-style natural language testing goals, all six models (OpenAI GPT-4.1, OpenAI o3, Gemini 2.5 Pro, Gemini 2.5 Flash, Claude Opus 4, Claude Sonnet 4) achieve correctness above 92%. Prior work often focuses on using LLMs to generate plans or acquire domain models, whereas this evaluation concentrates on the goal-translation step and covers both expressible and inexpressible goals. The benchmark contains 90 goals, 45 expressible and 45 inexpressible, spanning four IPC domains (Rovers, Barman, Woodworking, Parking) and four video-game-inspired domains; expressible goals were first matched exactly against reference PDDL goals, with unmatched cases manually reviewed by a PDDL expert, while inexpressible goals were judged automatically.
There are actionable differences between models: Gemini 2.5 Flash has the highest correctness (96%) and the fewest false positives (2), GPT-4.1 has the shortest average response time (1.61), and Gemini 2.5 Pro, despite 94% correctness, has the longest average response time (15.83). These differences indicate that which model to choose depends on whether a pipeline prioritizes accuracy or latency, rather than models being interchangeable. Table 1 reports each model's correctness rate, average prompt response time, number of false positives, number of false negatives, and number of responses with invalid syntax; only a single invalid-syntax response occurred across all evaluations.
Remaining failures stem mainly from language ambiguity and limitations in domain representation, and the authors propose next steps of expanding evaluation datasets, refining prompt engineering, and integrating feedback loops for clarification and error recovery. This frames translation as an ongoing process requiring interaction and correction rather than a one-shot mapping. The conclusion explicitly lists these three next steps and notes that some failures arise from inherent language ambiguity while others come from current limitations in parsing nuanced domain logic or handling edge cases where the formalism cannot keep pace with human creativity.
Perspective
This work targets goal translation for classical STRIPS planning with numeric fluents, where the input includes a PDDL domain model, a list of objects and their types, and a natural language goal, and the initial state can be extracted directly from the game, so the domain model is still built and maintained once by an expert. It applies to game-testing pipelines that want to turn test goals into PDDL goals automatically and offers a reference point for other natural-language-to-formal-goal settings; the evaluation covers eight domains and 90 goals, 45 expressible and 45 inexpressible.
A careful reader would still want to know: the judgment of inexpressible goals was automated and the quality or detail of the natural language explanations was not manually checked, so the semantic quality of correct refusals remains unclear; among expressible goals, cases not matching a reference exactly relied on manual review by a single PDDL expert; each benchmark used a single prompt per model with no reported stability across repeated runs, and the paper itself notes that even temperature-zero inference can diverge due to batch-size-dependent numerical effects and non-batch-invariant kernels; and the goals were crowd-sourced from company employees, so whether their distribution matches that of real testing teams remains an open question.
