MilleMiglia: A realistic instance generator for middle-mile logistics
Synopsis
The work introduces and open-sources MilleMiglia, a C++ instance generator serialized with Protocol Buffers that synthesizes middle-mile logistics networks from statistical distributions over spatial placement, demand, and vehicle rotations, embedding hard constraints such as fixed schedules, distribution-center throughput limits, and cross-vehicle synchronization into a single data format, thereby offering reproducible benchmarks from small academic toy problems to continent-wide industrial scale while preserving corporate privacy.
Interpretation
It formulates middle-mile logistics as a multi-commodity flow problem on a space-time graph, where nodes represent a specific distribution center at a specific time interval and arcs represent vehicle movements over time or a shipment being held at a distribution center for storage and sorting. Compared with prior practice of modeling first- and last-mile stages as vehicle routing problem (VRP) variants, this formulation explicitly captures that a shipment can move between trucks and span a multi-day time horizon. A conceptual and modeling contribution; the text defines nodes and arcs in prose and states that existing VRP solvers cannot apply to the middle mile.
It identifies three hard constraints of middle-mile operations that are difficult to relax: fixed vehicle schedules, distribution-center throughput limits per hour, and synchronization where the arrival of one vehicle is the prerequisite for the departure of shipments on a different vehicle. The text notes that many academic VRPs are defined with few constraints, whereas relaxing these would distort the structure of the operational problem, motivating a dedicated problem formulation. A qualitative argument based on a description of middle-mile operational characteristics, without quantitative experiments.
It implements and open-sources the MilleMiglia generator, which synthesizes networks using statistical distributions: distribution centers are placed via gravity models or spatial clustering to reflect population and industrial density, demand is generated as origin-destination pairs following realistic volume and weight distributions, and rotations are structured vehicle schedules rather than arbitrary connections. Responding to the lack of public high-quality data in this domain and to firms treating network topologies and demand volumes as sensitive proprietary information, it provides a privacy-preserving synthetic substitute. The text states these distributions interpolate between publicly available information from industrial actors and privately disclosed data; the generator is written in C++ and uses Protocol Buffers so each instance is stored in a single file consumable by solvers in different programming languages.
It offers a spectrum of instances across sizes and hardness and supports learning scenarios. Small instances are equivalent to academic toy problems for testing exact algorithms, industrial instances are large-scale continent-wide problems requiring advanced heuristics or metaheuristics, intermediate sizes and hardness are available, and the generator can create huge data sets to train machine learning algorithms. A tool-capability description; the text reports no benchmark results on solution quality or runtime.
Perspective
The work targets researchers and solver developers who need middle-mile network benchmarks, applicable to e-commerce and retail, automotive parts, and time-sensitive movements such as temperature-controlled pharmaceuticals; its synthetic instances are meant to substitute for corporate network topologies and demand data that cannot be disclosed, spanning from academic toy problems to continent-wide industrial problems and usable for generating machine learning training data.
As an introductory article, it reports no quantitative results on solution quality, runtime, or fidelity to real networks, and does not present scale parameters of specific instances; how closely the synthetic distributions match real networks and how applicable the generated instances are to actual operational decisions remain to be examined in later use. The text also notes that the specialized solver and API are still under development, so their effectiveness is yet to be reported.
