View Single Post
  #2  
Old 07-21-2026, 18:43
chants chants is offline
VIP
 
Join Date: Jul 2016
Posts: 836
Rept. Given: 47
Rept. Rcvd 52 Times in 32 Posts
Thanks Given: 742
Thanks Rcvd at 1,148 Times in 531 Posts
chants Reputation: 52
I mean, this is precisely how a model trained to predict the next toen with attention layers and thinking modes would be expected to work when the training data is full of such strong patterns. I would be more surprised knowing the transformer architecture that it did not work this way.
Reply With Quote