Skip to main content
Back to timeline
arXivSource publication:

SGAnalog builds a benchmark from 273 open-source tapeout circuits: top model hits 56.1% exact isomorphism, 91.2 on sizing

Related research and updates

Synopsis

The authors harvested human-designed analog circuits from Tiny Tapeout open-source repositories at pinned submission revisions to build SGAnalog, a benchmark of 273 topologically distinct top-level designs with schematic-to-netlist transcription and device-sizing tasks; across seven models on 66 transcription tasks the strongest reaches 56.1% exact graph isomorphism, and on 17 sizing tasks the leading model converges on all proposals and scores 91.2 out of 100, with the two tasks producing different model rankings.

Source-provided article image: SGAnalog: An End-to-End Circuit Benchmark from Open-Source Silicon Tapeouts
Figure 2 ·

Figure 2: Both tasks, end to end. Top: transcription. The input schematic image (left) and claude-opus-5’s verbatim output netlist (right) for a five-transistor OTA. Its net names differ from the drawing, but its graph is isomorphic to the reference. Bottom: sizing for a fully differential amplifier. The model receives the topology with sizes removed, then its proposal and the human reference run in the same author testbench. Aggregate results appear in Table 3 .

arXiv

Interpretation

Starting from 27 shuttle indexes with blobless clones and license verification, the benchmark yields 933 circuits, 425 top-level designs, and 273 unique device-net graphs, fixing 66 transcription and 17 sizing tasks. Unlike analog benchmarks drawn from textbooks, published figures, or synthetic enumeration, every circuit is fetched at the revision the shuttle index records, and each schematic image and SPICE netlist is exported from the same .sch source, avoiding a separate annotation step. The harvest funnel and composition tables give counts at each stage; licenses are verified from license texts fetched at pinned commits; classification tables are byte-identical across two complete runs, and simulation tables regenerate identically up to per-run timing columns.

Transcription is scored by graph isomorphism rather than text match, treating MOSFET drain and source as interchangeable and ignoring net names, alongside structural F1 and W/L parameter accuracy. This scoring makes render style a controllable nuisance variable and gives a common evaluation across circuits regardless of testbench state, enabling the commit-date analysis. On claude-opus-5, blinding author-chosen labels lowers exact isomorphism from 51.5% to 36.4% while structural F1 is unchanged and parameter accuracy improves; because the scorer ignores names, the authors read this as labels serving as visual anchors for connectivity tracing.

The sizing task fixes the human topology and asks the model for device sizes, scored against the simulated performance of the human-authored sizing under the same PDK, testbench, and simulator via model-to-reference metric ratios. The reference is the submission author's own design evaluated under the same testbench, not a literature amplifier topology; nonconvergent or sizeless proposals score zero, and a convergent task without an eligible scalar metric receives a survival score of one. The 17 tasks come from 412 author-testbench runs, of which 190 pass the 600-second, real-device-card, and hierarchy-resolution gate; the leading model converges on all 17 proposals and scores 91.2, while the two newest Claude models refuse 4 and 11 of the same prompts they transcribe without objection.

The benchmark ships commit dates, blinded renders, release canaries, and a closed-book protocol for probing potential exposure. Closed-book evaluation declares no tools by construction, blinded renders remove author-chosen labels, and canaries mark generated artifacts; the authors separate the training-time channel from inference-time retrieval. The authors state the commit-date comparison is descriptive rather than a contamination estimate, because difficulty differs across the date boundary, graph-hash groups are not separated across it, and an earlier design version may predate the pinned revision; across the two usable older-cutoff comparisons, exact-match accuracy shows no consistent pre-cutoff advantage.

Perspective

The benchmark targets schematic-level reasoning, not layout or subjective design style; transcription measures whether a model recovers circuit structure from an image, and sizing measures whether model-selected dimensions preserve simulated behavior in the author's testbench. The fixed task sets are eligibility-filtered subsets of the broader collection, 66 transcription and 17 sizing tasks, held fixed across models. The release includes a browser interface over verified circuit-testbench pairs and a static companion page with precomputed runs of ten representative pairs, requiring only Docker. The authors list next milestones: per-class measurement templates starting with gain, gain-bandwidth product, and phase margin for the amplifier class, per-circuit memorization probes, and topology generation as a task.

The commit-date comparison is descriptive, and the authors state it is not a contamination estimate because difficulty differs across the date boundary, graph-hash groups are not separated across it, and an earlier design version may predate the pinned revision. The blinded-render result is consistent with labels acting as visual anchors but, as the authors note, does not rule out prior exposure. Human sizes are public in the upstream repositories, so sizing is scored at perturbed operating points with regurgitation diagnostics. The refusal set is stable in bulk but not per prompt: a full repeat run of claude-fable-5 refuses 12, with ten prompts refused both times and three changing sides. The authors also note convergence alone is too weak a bar, since two proposals converge with output swing below one percent of the reference. Topology generation is not claimed in this version because a uniform design-brief representation and automatic pin-matching remain unsolved.

Sources