Public articles linked to the same research event.
arXiv The authors introduce universal transformers: fixed transformers whose internal parameters stay fixed while a suitable input embedding encodes a description of a target model, letting them simulate any transformer in a given class; they give explicit sparse constructions achieving universality when the embedding dimension is sufficiently large, show that randomly initialized transformers are universal almost surely, and empirically validate the theory on parenthesis balancing and multi-hop reasoning, suggesting much of a transformer's expressive power may reside in its input representation rather than its learned weights.
The authors introduce universal transformers: fixed transformers whose internal parameters stay fixed while a suitable input embedding encodes a description of a target model, letting them simulate any transformer in a given class; they give explicit sparse constructions achieving universality when the embedding dimension is sufficiently large, show that randomly initialized transformers are universal almost surely, and empirically validate the theory on parenthesis balancing and multi-hop reasoning, suggesting much of a transformer's expressive power may reside in its input representation rather than its learned weights.
The authors introduce universal transformers: fixed transformers whose internal parameters stay fixed while a suitable input embedding encodes a description of a target model, letting them simulate any transformer in a given class; they give explicit sparse constructions achieving universality when the embedding dimension is sufficiently large, show that randomly initialized transformers are universal almost surely, and empirically validate the theory on parenthesis balancing and multi-hop reasoning, suggesting much of a transformer's expressive power may reside in its input representation rather than its learned weights.
The authors introduce universal transformers: fixed transformers whose internal parameters stay fixed while a suitable input embedding encodes a description of a target model, letting them simulate any transformer in a given class; they give explicit sparse constructions achieving universality when the embedding dimension is sufficiently large, show that randomly initialized transformers are universal almost surely, and empirically validate the theory on parenthesis balancing and multi-hop reasoning, suggesting much of a transformer's expressive power may reside in its input representation rather than its learned weights.