Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

ServeTwin couples specification-driven analytical timing with a stateful serving loop, reproducing the throughput-interactivity frontier at 3.6% mean error without target-hardware profiling

The authors present ServeTwin, a closed-loop simulator that couples specification-driven analytical timing with a stateful serving loop, uses iSTAGE to derive per-iteration traces from model and engine specifications while avoiding target-hardware operator profiling, and implements a vLLM-compatible interface that runs unmodified serving benchmarks; against real deployments it reproduces InferenceX's steady-state throughput-interactivity frontier with a 3.6% mean error and predicts LMBenchmark's multi-turn performance with a 9.9% error while tracking KV-cache evolution, and pre-silicon sweeps show that the preferred HBM bandwidth-capacity tradeoff reverses across workload states and that increasing concurrency shifts the bottleneck from memory bandwidth to scheduler and runtime overheads.