GROB's multi-agent architecture captured Census identifiers on 9 September from public traces alone, later resolved by public revision records to specific 16–17 June Census requests
Synopsis
GROB is a multi-agent architecture that performs controlled, read-only collection of public Internet traces when privileged telemetry is unavailable and preserves observations for later resolution; in a frozen September 2026 corpus, Census identifiers captured on 9 September were later resolved by public revision records to specific read-only Census requests from 16–17 June, while other traces showed links of varying strength and public traces alone did not establish organizational attribution.
Figure 1: GROB architecture. Sentinel controls detection and gating, Scout performs hypothesis testing, and Librarian maintains validated semantic memory.
arXivInterpretation
On 9 September 2026 at 22:43:48 UTC, GROB collected three Census-labelled identifiers, including AgentOpenAICensusTest and AgentOpenAICensusLink1781645460, and the same capture preserved OpenAICensusGCTBridge. These identifiers were already retained before OpenAI's Census-specific public acknowledgement on 25 September, showing that sparse public traces can be collected before their significance is understood. Evidence comes from specific capture timestamps and identifiers in the frozen September corpus, and the identifiers recur across multiple collected workspaces, indicating persistence of the same source material rather than separate incidents.
Public DSE revision records resolved these identifiers to read-only Census requests dated 16–17 June targeting api.census.gov/data/2020/acs/acs5/pums, performing a PWGTP-weighted tabulation by sex for Wyoming (state=56) and containing a redacted Census developer-key parameter. The same request persists across several public revisions and labels within roughly eleven minutes, establishing an artifact-level DataUSA-to-Census chain whose correspondence extends beyond a target-name match to a compatible request mechanism. Resolution rests on the specific request endpoint and query structure in the public DSE revision archive; however, the redacted credential prevents comparing whether the same key was involved, and no first-party record links the June revisions to the execution later disclosed by OpenAI.
In the SEC branch, GROB collected AgentCustom006PrettyLinks1782006000 in a 9 September workspace, and public revision reconstruction resolved the page to a June event containing the exact resource sec.gov/files/county.json; in the XX91 branch, the pre-22 corpus collected the complete five-page ProbierWiki MemProof/PlaceProof0--3 family and two corresponding DSEWiki identifiers, with public revision records completing the same five-role family in the same role order within a 106-second relative interval. These cases demonstrate artifact-level correspondence and cross-surface structural evidence, respectively, providing stronger cross-record joins than naming similarity alone. The SEC exact resource matches the same county-data resource independently reconstructed later; the XX91 cross-surface structure is stronger than naming similarity, but neither establishes common control or a unique underlying agent, and the rendered timestamps do not specify a timezone.
The pre-22 September corpus contains 1,246 unique strict agent/OpenAI-style identifiers, of which 726 participate in 1,614 similarity edges; the lexical graph is positioned as a candidate-grouping mechanism, not evidence of interaction, common control, or attribution. The report distinguishes lexical, temporal, direct-provenance, artifact, external-resolution, and execution-identity layers as differing in evidentiary strength, making explicit that execution identity requires first-party evidence. Counts come from 25 snapshots collected on 9, 10, and 21 September and are reported separately from the persistence audit using 35 deduplicated workspaces, because these populations do not represent directly comparable measures of real-world activity.
Perspective
The work targets defenders and researchers conducting public-trace investigations when privileged telemetry or known targets are unavailable, in settings that permit read-only collection of publicly accessible material. It enables sparse identifiers collected early to be resolved into specific requests or resources as later public evidence emerges, supporting retrospective identification and reconstruction. Methodologically, the system performs controlled link traversal from allowlisted starting URLs, retains source and capture time, and admits evidence to persistent memory only through deterministic code, so model-generated interpretations cannot enter persistent memory without review. In this setting, the value lies in preservation and prioritization rather than in producing an attribution decision.
Public traces alone do not establish organizational attribution or shared execution; self-identifying labels are not identities, and a public revision IP records only the address associated with an edit. The redacted credential prevents determining whether the same key was involved, and no first-party record links the June revisions to the execution later disclosed. Some branches (NovOne, Thailand, Education) remain contextual leads rather than incident-level joins, and the Wikimedia case is a public-page-level rather than revision-level correspondence. The prototype did not always retain the complete content of every discovered page, so fuller reconstruction depended on public records available elsewhere; temporary services, private scans, and expiring records can leave parts of the activity unavailable, so absence from the corpus is treated as an auditability limitation rather than evidence an event did not occur. A text-provenance screen found no known fixed watermark marker in the core wiki subset, and statistical watermark attribution remains unresolved because no matched keyed detector could be applied reliably.
