Skip to main content
Back to timeline
arXivSource publication:

MiniCorp lets an agent firm run an e-commerce company: price elasticities of 2.9–3.2, 23–32% ad-attributed lift, and delayed organic gains

Synopsis

MiniCorp is a closed-loop office simulator that couples a persistent agent firm with an evolving e-commerce external market: the external world explicitly models search demand, matching and ranking, click and conversion, sponsored-search auctions, and dynamic competitors, while the internal world is a firm of standing roles that read exported business data, communicate, and propose decisions. Experiments show the integrated market reproduces empirical response patterns in price elasticity, the persistent effects of temporary discounts and stockouts, and the delayed organic gains of advertising, and that agents coordinate across roles while generating replayable counterfactual enterprise data.

AI-generated editorial illustration: MiniCorp: The Last Mile of the AI Agent Firm

Interpretation

MiniCorp models business operation as a closed loop between two stateful systems: the external world is a self-advancing economy that fires shocks, delivers in-flight orders, moves competitors, and shifts demand whether or not the firm acts; the internal world is a persistent agent firm whose standing roles read a sanitized weekly export, communicate, and propose decisions that take effect only in the next period. Unlike task-bounded benchmarks that assemble multi-agent teams around a task and discard them when it ends, the organizational structure here exists first and work is routed into it, and a one-period lag prevents the company from observing and rewriting the same outcome within a single period. The paper describes the architecture and states that the export passes through whole-table drops, protected-column removal, and a schema audit that fails the entire export if anything protected survives.

The external e-commerce world explicitly models the response chain from decision to sale: 1,000 search queries (316 common queries from Amazon autocomplete and 684 constructed long-tail queries) with Zipf-distributed traffic, passing through matching, organic ranking, click and conversion, a four-slot second-price sponsored-search auction, and competitor pricing and advertising responses, with an unobserved conversion multiplier per query–category pair. The authors argue against mapping an action directly to sales or profit while omitting intermediate mechanisms, because an experienced firm diagnoses whether performance is constrained by impressions, click-through, conversion, advertising cost, or competitor actions and intervenes at the appropriate stage. Mechanisms draw on Capitalism II economic logic, ESCI few-shot relevance labels, Amazon autocomplete queries, and published evidence on tariff pass-through, supplier lead times, and defect rates; parameters are sampled once at world initialization and held fixed for the run.

End-to-end fidelity tests show the integrated simulator reproduces empirical market response patterns: a 10% price reduction and increase yielded absolute arc elasticities of 2.9 and 3.2; a 20% discount for four weeks followed by restoration raised units 71% during the discount, left them 31% above control for the first four weeks after restoration and 11% above for the following six, with nearly half of incremental units sold after the discount ended; a two-week stockout left 59% of the cumulative unit shortfall occurring after restocking and sales 36% below control in week 26; advertising raised ad-attributed units 23–32%, grew the organic gain from 0.5% to 5.8%, returned $0.53 in contribution per additional advertising dollar within the observation window, and drew 84% of incremental units from competitor displacement. The authors separate outcome-level from pathway-level validation, checking not only whether an intervention produces the expected result but whether intermediate responses and their temporal ordering match empirical evidence, and focusing on response chains that were not direct calibration targets to reduce the risk of getting the right outcome for the wrong reasons. The four tests use controlled interventions across worlds initialized from the same checkpoint and, where applicable, the same random draws; the authors note that robustness across independently initialized worlds remains to be evaluated.

At the organizational and strategic level, of 30,248 work items 58% were carried over from the prior cycle, 21% were message-triggered, 10% interrupt-driven, and only 11% self-planned; the CEO's highest rejection rates were replenishment and logistics (29%) and advertising and pricing (25%), and the lowest was quality incidents (13%); in a paired comparison, the guided firm spent about $1,700 on advertising over 26 weeks, generated about $31,000 in revenue and $1,913 in gross profit after advertising, while the unguided firm spent about $50 in total, generated under $200 in revenue, and recorded a $7 loss. The paper decomposes organizational realism into work provenance, process signatures, and delegation structure, and shows that separating capability from authority routes coordination around the reporting line: quality engineering can diagnose defects but cannot place a lot on hold, producing high-volume two-way exchange with supply chain. Statistics come from run records of work-item provenance, category, message counts, and rejection rates, plus a paired guided/unguided run with the same initial state and market rules; the authors note both gross-profit measures exclude refunds, inventory write-offs, and overheads.

Perspective

The environment targets researchers and engineering teams studying how agents can collectively run a company and generating longitudinal, counterfactual enterprise data, with e-commerce as the demonstration setting. It applies where the external world explicitly models search demand, matching and ranking, click and conversion, advertising auctions, and competitor responses, and where the internal firm consists of standing roles deciding in weekly cycles with approved decisions taking effect one period later. The authors present the framework as extendable to additional firms, organizational structures, and industries, and as usable for evaluating and training models on long-horizon organizational tasks.

Robustness across independently initialized worlds remains to be evaluated, an open question the authors state explicitly. Fidelity tests treat direction, shape, temporal dynamics, distribution, and trade-offs as primary criteria, with numerical agreement used as supporting evidence only when settings are comparable, so specific magnitudes are context-dependent. The RQ4 shock-response evaluation answered 11 of 20 runs fully and 4 partially, and scores whether the expected action was proposed and accepted within eight weeks rather than final profit. The evasive organizational behavior and the self-initiated title rewrite in RQ5 come from run records, and their reproducibility is not quantified in the text. This is a full-text parse without figure images, so readers checking table details should consult the original.

Sources