One quibble: while Wikipedia is currently an excellent set of human generated summaries, it is by definition NOT a “primary source” and is a “secondary source” instead.
This distinction I think is important past pedantry as a well written Wikipedia article ALSO has links to primary sources as well. This is important as it seems like good training of AI to produce summaries would include them reading complete primary sources and then producing Wikipedia like output rather than just memorizing Wikipedia directly.
I think AI output watermarking and AI detection like pangram can also help a lot for curation of AI training data sets.
However this only removes the first order feedback loop. If AI text starts dominating what humans read on the internet then that will start influencing how humans themselves write, which WILL be picked up in future training. However I don’t see this as necessarily a bad thing. Culture evolves over time even with only humans without AI, so if AI-isms like claudish for example are accepted by large numbers of humans such that LLM compression picks it up in the training set, that’s just the new form of cultural evolution.
As a Wikipedia proponent, who will pedantically explain why Wikipedia is one of the best things about the Internet, I agree.
With that said, I think probably part of the strength of LLMs is that sources like Wikipedia exist, as reinforcement and validation of the "knowledge." It's not enough to just go with primary sources. Wikipedia serves as a focusing mechanism of what do humans care about at a top-level. Reddit (another favorite of LLMs) serves a similar purpose, but more in the category of "explain this thing to me." I don't think these tools work for general purpose without these general purpose-focusing mechanisms.
Put another way, as a published scientist, who is also very pedantic about publishing, I think you'd be missing a whole lot of knowledge of what humans care about if you only look at what was published in primary sources. The LLMs can't just produce Wikipedia articles for us because they need Wikipedia to even work in the first place.
Excellent article!
One quibble: while Wikipedia is currently an excellent set of human generated summaries, it is by definition NOT a “primary source” and is a “secondary source” instead.
This distinction I think is important past pedantry as a well written Wikipedia article ALSO has links to primary sources as well. This is important as it seems like good training of AI to produce summaries would include them reading complete primary sources and then producing Wikipedia like output rather than just memorizing Wikipedia directly.
I think AI output watermarking and AI detection like pangram can also help a lot for curation of AI training data sets.
However this only removes the first order feedback loop. If AI text starts dominating what humans read on the internet then that will start influencing how humans themselves write, which WILL be picked up in future training. However I don’t see this as necessarily a bad thing. Culture evolves over time even with only humans without AI, so if AI-isms like claudish for example are accepted by large numbers of humans such that LLM compression picks it up in the training set, that’s just the new form of cultural evolution.
As a Wikipedia proponent, who will pedantically explain why Wikipedia is one of the best things about the Internet, I agree.
With that said, I think probably part of the strength of LLMs is that sources like Wikipedia exist, as reinforcement and validation of the "knowledge." It's not enough to just go with primary sources. Wikipedia serves as a focusing mechanism of what do humans care about at a top-level. Reddit (another favorite of LLMs) serves a similar purpose, but more in the category of "explain this thing to me." I don't think these tools work for general purpose without these general purpose-focusing mechanisms.
Put another way, as a published scientist, who is also very pedantic about publishing, I think you'd be missing a whole lot of knowledge of what humans care about if you only look at what was published in primary sources. The LLMs can't just produce Wikipedia articles for us because they need Wikipedia to even work in the first place.