Skip to main content
Back to timeline
arXivSource publication:

pm4aa mined eight roles from eight years of Commitizen repository logs and generated five executable AI agents, with smoke tests routing correctly but autonomy scoring lowest in the user study

Synopsis

The work presents pm4aa, a pipeline that extracts object-centric event logs from GitHub repositories via PyStack't, maps commits to SE tasks with a Conventional Commits regular expression, partitions 589 users into eight roles with a priority-ordered rule classifier, applies object-centric, imperative (BPMN), and declarative (DECLARE) process mining per role, has an LLM generate process descriptions, and synthesizes a LangGraph multi-agent application through IBM BOB; on the Commitizen project (November 2017 to November 2025, 21,488 events, 4,813 objects) it produced five role agents, three smoke tests routed correctly, and a ten-participant user study rated knowledge schema, operational clarity, and human engagement at a median of 4, accountability at 3, and autonomy at only 2.

Source-provided article image: Using Process Mining to Generate AI Agents from Software Engineering Process Records
Fig. 1

transforms SE repository data into operational, role-specific AI agents. Fig. 1.1 provides an overview of the steps of the pm4aa pipeline, which we elaborate below. A repository containing the full pipeline implementation and the accom- panying application is available at: https://github.com/liorlimonad/pmaa.

· Page 5

Interpretation

It introduces pm4aa, a pipeline that treats software repository event logs as business process logs, uses object-centric process mining to delineate each role's task scope and interacting objects, uses declarative mining (DeclareMiner, minimum support 0.5, maximum constraint cardinality 2) to elicit behavioral constraints as guardrails for agent execution, and then uses an LLM to generate process descriptions and synthesize executable agents. Existing LLM multi-agent SE frameworks such as ChatDev and MetaGPT rely on fixed canonical roles (programmer, reviewer, tester) that do not adapt to a project's actual workflow, while Agent System Mining in the BPM field mines role structures but is retrospective and simulative. pm4aa differs by turning mined artifacts directly into deployable LangGraph agents rather than synthetic traces or digital twins. The paper gives a complete six-step pipeline description and a public implementation repository, and runs it end to end on a real open-source project; however, declarative constraints are informative only for the bot and issue_reporter roles, because the remaining roles performed a single activity type after flattening.

In the Commitizen case, the pipeline identified 1,721 of 2,765 commits (62.2%) with well-formed Conventional Commit messages and created task objects for them, and assigned 589 users to eight roles via priority-ordered rules, with issue_reporter at 425 users and 16,037 events, maintainer at 8 users and 2,149 events, and bot at 6 users and 3,394 events, showing a highly skewed event distribution. It shows that SE task types and a ten-class semantic scheme can be derived automatically from the type(scope): description prefix of commit messages without manual annotation, turning unstructured repository history into minable role behavior data. The data come from a single project spanning roughly eight years, with 21,488 events and 6,534 objects across four types (commits, tasks, issues, users); the authors state that generated specifications for low-volume roles should be read as more exploratory, and that role thresholds such as maintainer requiring at least 20 commits across at least three task classes are pragmatic proof-of-concept heuristics not statistically optimized.

The final implementation did not turn every mined role into an agent but consolidated them into five: issue_reporter_agent, bot_workflow_agent, implementation_agent, quality_agent, and technical_writer_agent, where implementation_agent merges the behavior mined from feature developer, contributor, maintainer, and DevOps engineer. It shows a deliberate design trade-off between mining and executable systems: mined roles reflect organizational reality, whereas an executable application needs roles that are semantically similar and operationally useful, so the two need not correspond one to one. The paper explicitly states the consolidation was deliberate because the mined roles showed similar implementation semantics; three smoke tests routed a documentation request, an implementation request, and a crash report to technical_writer_agent, implementation_agent, and bot_agent respectively, and the authors stress these are sanity checks rather than evidence of general runtime effectiveness.

A ten-participant exploratory user study found that process-mined specifications scored well on knowledge schema (median 4), operational clarity (4), and human engagement (4), with accountability at 3 and autonomy lowest at 2; at the role level feature_developer aligned best (knowledge schema 4.5, accountability 3.5) while technical_writer scored only 1.5 on autonomy. It applies HCI human-AI alignment dimensions (knowledge schema, autonomy, operational, accountability, human engagement) to evaluate process-mined agent specifications and profiles differences across roles, rather than only reporting whether the pipeline runs. The sample was ten master's students, teaching assistants, and PhD researchers from a single university recruited as a convenience sample, and participants evaluated written specifications rather than running agents; the authors state this limits causal claims and that participants, as non-practicing engineers, may not reflect developers deploying hybrid human-AI teams in production.

Perspective

The work targets projects with observable development history; the authors state it is not applicable to greenfield projects with no prior traces, where pm4aa-derived agents can serve only as templates requiring later adaptation once project-specific traces exist. The intended users are SE teams and researchers who want to derive project-specific agent specifications and implementations from repository event logs, under the precondition that commit messages follow a structured convention such as Conventional Commits and that the log is reasonably clean. Evaluation is set in a single open-source project, Commitizen, plus a review of written specifications, so the findings apply to whether specifications are readable, understandable, and assignable, not to runtime behavior in production.

Whether clearly scoped specifications translate into aligned runtime behavior is left open by the paper itself, since the user study assessed written specifications rather than running agents. Autonomy scored a median of only 2, meaning humans struggled to tell when an agent may act autonomously versus when it must wait for approval, a divergence especially visible for technical_writer (1.5) and quality_engineer (2); how to encode autonomy boundaries into specifications remains an open question. Open-text responses also raise the risk of endless loops between feature_developer and quality_engineer due to ambiguous validation criteria, and of high-volume agent output turning the human developer into a vetting bottleneck. The hard single-role partition lets maintainer (8 users, 2,149 events) absorb behavior that could otherwise enrich the technical writer, quality engineer, and feature developer sub-logs, and overlapping role membership is listed as future work. Finally, SE event logs capture coarse-grained activities such as commits and issue updates, whereas coding agents operate at a finer granularity such as writing code or running tests, and the text does not conclude whether that granularity gap affects specification usability.

Sources