Moving bidding on device: one tick of sync staleness overspends budgets by 17.77%, and 50 ticks by 1,669.31%
Synopsis
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
Interpretation
The paper formalizes accounting information misalignment, in which a shared budget is replicated across autonomous devices that observe only their own debits, and derives a finite-window expected excess-debit bound. Industrial pacing work already notes that delayed spend information can push fast-spending campaigns over budget; the narrower novelty here is to model budget state as replicated across autonomous devices that cannot see one another's debits, derive a finite-window bound from lag, arrivals, crossing campaigns, and conditional payment caps, and quantify the failure under the product's scoring and pacing logic. In a simulation with 36 campaigns, 50 devices, 3,000 pool-wide Poisson arrivals per tick, and a 100-tick window, the 12 vertical intensities sum to about 309.7 servable arrivals per tick; Table 2 reports positive paired slack confidence intervals in every bounded-value cell, and the authors state this is consistency of estimated expectations rather than a claim that the bound holds on each finite replicate.
Overspend from sync lag rises sharply with lag, does not require severe budget pressure, and retains its direction under a declared bursty, heterogeneous-device demand process. A new pressure-lag sweep shows long-lag excess debit persists at two-times as well as 20-times pressure, and a visible-balance no-sale guard ablation separates local crossing debit from stale cross-device spending. Across 30 paired demand paths, Even overspends 17.77% at one tick and 1,669.31% at 50 ticks; at two-times pressure the 50-tick figure is still 106.95%, with 367.37% and 801.41% at five- and ten-times pressure. In the bursty/heterogeneous sweep the mean curve is strictly increasing, 24 of 30 individual paths are themselves strictly increasing, and every paired lag-minus-zero 95% confidence interval is positive.
When the ML/pacing score transformation used for ranking is also allowed to change payment units, a second incentive misalignment appears: the auction charges the runner-up's scaled score rather than the winner's critical base bid. The paper characterizes a canonical dominant report on rival auctions and gives an executable implementation-level counterexample that isolates the runner-up's multiplier in the winner's price; it also notes that critical-base-bid payment makes the per-auction rule DSIC conditional on current multipliers but neither proves dynamic truthfulness nor changes paced ranking. At deterministic zero lag, 98.23% of rival auctions admit a profitable deviation, or 96.97% of all sales including the outcome-invariant one-bidder branch; in the counterexample truthful reporting scores 320 and loses to 400, while reporting 1100 scores 440, pays 401, and obtains utility 399, with a unit test running that vector through the auction port. Critical-base-bid payment finds no profitable per-auction deviation in any of 420 replicate cells.
The paper audits accounting consistency and payment rules as two distinct design axes and quantifies how pacing affects price and delivery volume. The authors argue that neither a more sophisticated controller nor faster sync alone repairs both problems, and report that score-space pacing multiplies delivery by about 11.5 and divides price by the same factor, which is qualitatively different from probabilistic participation throttling. At deterministic zero lag, Even wins 19,358.1 impressions at an average numeric charge of 7.07 score units while ASAP wins 1,679.0 at 81.35; PI is detectably better only at the two shortest deterministic lags, and all calibrated-synthetic effects are inconclusive. The visible-balance guard produces exactly zero excess debit in all 30 replicates at zero lag but still overspends 11.88% at one tick.
Perspective
This work targets on-device auctions in which ML-mediated eligibility and ranking decisions move onto privacy-preserving clients, under a setting where shared budget state is replicated across devices, each device sees only its own debits, and a central service broadcasts a pacing multiplier and spend snapshot every tick. It offers an accounting and mechanism audit in dimensionless integer score units, useful for comparing faster sync against escrow-style device reservations along communication, overspend, and utilization, and for evaluating payment rules separately from ranking rules. For a reader, it supplies directional evidence on how lag becomes excess debit plus a reusable audit protocol, not an estimate of real advertiser charges.
The context pool is entirely single-teacher labelled, with 5,245 of 50,808 records (10.32%) admitting at least one campaign and no human-gold rows or completed human ratings, so topic, intent, safety, and ad-eligibility accuracy are unaudited; the label-perturbation stress test covers one construction and cannot estimate real precision, recall, or taxonomy confusion. Every draw is a fresh session, so the product's three-impressions-per-campaign-per-session cap never binds and results cannot be extrapolated to repeated-session inventory. The primary curve uses homogeneous independent Poisson streams and independently sampled contexts, while the robustness arm covers one declared two-state Markov process and one fixed high/low device mixture, not diurnal cycles, trace-linked device contexts, network failures, or strategic timing. All delivery runs assume truthful reports even though the deployed payment rule makes truth-telling suboptimal, and the resulting repeated strategic game is not solved. The log-normal value arm is untruncated, has no finite conditional charge cap, and lies outside Theorem 1's support premise. Integer rounding is a chosen accounting granularity, and the authors explicitly withdraw any claim that the visible-budget guard and small-lag crossing residuals transfer to a real currency ledger. In addition, the loaded text leaves the numeric cells of Table 1 and Figure 1 empty, so specific percentages can only be reported from the prose and Table 2.
